Apparatus, system and method for detecting physiological motion from audio and multimodal signals

By using speakers and microphones on smartphones or tablets to generate and process inaudible sound signals, the challenge of monitoring breathing and movement during sleep without relying on dedicated equipment has been solved, enabling efficient detection of sleep states and apnea.

CN116035528BActive Publication Date: 2026-08-25RESMED SENSOR TECH LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202310021127.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-09-19
Filing Date
2017-09-19
Publication Date
2026-08-25
Estimated Expiration
2037-09-19

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently monitor human physiological movements without relying on specialized hardware, particularly breathing and body movements during sleep, such as sleep apnea.

Method used

Using the speakers and microphones of mobile devices such as smartphones or tablets, inaudible sound signals are generated and sensed. These signals are then processed to detect breathing and movement, including the use of pitch pairs, frame sequences, and baseband motion signal processing techniques.

Benefits of technology

It enables efficient monitoring and detection of breathing and movement during sleep without relying on specialized equipment, and can identify sleep states and apnea events, providing the ability to automatically manage and assess sleep conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116035528B_ABST
    Figure CN116035528B_ABST
Patent Text Reader

Abstract

Methods and devices provide for detecting physiological motion with active sound generation. In some versions, a processor can detect respiration and / or gross body motion. The processor can control production of a sound signal in the vicinity of a user through a speaker coupled to the processor. The processor can control sensing of a reflected sound signal through a microphone coupled to the processor. The reflected sound signal is a reflection of the sound signal from the user. The processor can process the reflected sound, such as through a demodulation technique. The processor can detect respiration from the processed reflected sound signal. The sound signal can be produced as a series of tone pairs in a time slot frame, or as a phase-continuously repeating waveform with varying frequency (e.g., triangular or ramp sawtooth). Evaluation of the detected motion information can determine sleep state or score, fatigue indication, object identification, chronic disease monitoring / prediction, and other output parameters.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] 1.1 Divisional Application Statement

[0002] This application is a divisional application of Chinese Patent Application No. 2017800692184, which was filed on September 19, 2017 (PCT International Application No. PCT / EP2017 / 073613) and entered the Chinese national phase on May 8, 2019. The invention is entitled "Apparatus, System and Method for Detecting Physiological Motion from Audio and Multimodal Signals".

[0003] 1.2 Cross-references to related applications

[0004] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 396,616, filed September 19, 2016, which is incorporated herein by reference in its entirety. 2 Background Technology 2.1 Technical Field

[0007] This technology relates to the detection of biological motion associated with a living object. More specifically, this technology relates to the use of acoustic sensing to detect physiological motion, such as respiratory motion, cardiac motion, and / or other smaller periodic body movements of a living object.

[0008] 2.2 Description of relevant technologies

[0009] Monitoring a person's breathing and body (including limb) movements (e.g., during sleep) can be useful in many ways. For example, such monitoring can be used to monitor and / or diagnose sleep-disordered breathing conditions, such as sleep apnea. Traditionally, a barrier to entry for active radio positioning or ranging applications has been the need for dedicated hardware circuitry and antennas.

[0010] Smartphones and other portable electronic communication devices are ubiquitous in daily life, even in developing countries where landline telephones are unavailable. There is a need for a method to monitor biological movement (i.e., physiological movement) in an efficient and effective manner, without requiring specialized equipment. The realization of such a system and method would address considerable technical challenges. 3. Summary of the Invention

[0012] This technology relates to systems, methods, and apparatus for detecting motion in an object, for example, when the object is asleep. Based on such motion detection, including, for example, respiratory movements, motion, sleep-related characteristics, sleep state, and / or apnea events can be detected. More specifically, motion applications associated with mobile devices such as smartphones and tablets use mobile device sensors, such as integrated and / or externally connectable speakers and microphones, to detect breathing and motion.

[0013] Some versions of this technology may include a processor-readable medium having processor-executable instructions stored thereon, which, when executed by a processor, cause the processor to detect the user's physiological movements. The processor-executable instructions may include instructions for controlling the generation of sound signals, including those near the user, via a speaker coupled to an electronic processing device. The processor-executable instructions may include instructions for controlling the sensing of sound signals reflected from the user via a microphone coupled to the electronic processing device. The processor-executable instructions may include instructions for processing the sensed sound signals. The processor-executable instructions may include instructions for detecting respiratory signals from the processed sound signals.

[0014] In some versions, the audio signal may be within the inaudible sound range. The audio signal may include pitch pairs forming pulses. The audio signal may include a sequence of frames. Each frame of the sequence may contain a series of pitch pairs, where each pitch pair is associated with a corresponding time slot within the frame. The pitch pair may include a first frequency and a second frequency. The first frequency and the second frequency may be different. The first frequency and the second frequency may be orthogonal to each other. A series of pitch pairs in a frame may include a first pitch pair and a second pitch pair, where the frequency of the first pitch pair may be different from the frequency of the second pitch pair. The pitch pairs of time slots in a frame may have zero amplitude at the beginning and end of the time slot and may have a ramp amplitude with round-trip peak amplitude between the beginning and the end.

[0015] In some versions, the time width of a frame can vary. The time width can be the width of a time slot within a frame. A series of tone pairs within a time slot can form different frequency patterns relative to different time slots within the frame. These different frequency patterns can be repeated across multiple frames. For different time slot frames within multiple time slot frames, these different frequency patterns can vary.

[0016] In some versions, instructions for controlling the generation of sound signals may include a pitch-to-frame modulator. Instructions for controlling the sensing of sound signals reflected from the user may include a frame buffer. Instructions for processing the sensed sound signals reflected from the user may include a demodulator to generate one or more baseband motion signals that include a breathing signal. The demodulator may generate multiple baseband motion signals, which may include quadrature baseband motion signals.

[0017] In some versions, the processor-readable medium may include processor-executable instructions to process multiple baseband motion signals. The processor-executable instructions for processing the multiple baseband motion signals may include an intermediate frequency processing module and an optimization processing module to generate a combined baseband motion signal from the multiple baseband motion signals. The combined baseband motion signal may include a breathing signal. In some versions, instructions for detecting the breathing signal may include determining the breathing rate from the combined baseband motion signal. In some versions of the audio signal, the duration of a corresponding time slot of a frame may be equal to 1 divided by the difference between the frequencies of a tone pair.

[0018] In some versions of this technology, the sound signal may include a repetitive waveform with varying frequencies. The repetitive waveform may be phase-continuous. The repetitive waveform with varying frequencies may include one of a ramp sawtooth, a triangle, and a sine waveform. The processor-readable medium may also include processor-executable instructions, which include instructions for changing the form of one or more parameters of the repetitive waveform. The one or more parameters may include any one or more of the following: (a) the peak position in the repetitive portion of the repetitive waveform, (b) the slope of the ramp in the repetitive portion of the repetitive waveform, and (c) the frequency range of the repetitive portion of the repetitive waveform. In some versions, the repetitive portion of the repetitive waveform may be a linear function or a curve function that changes the frequency of the repetitive portion. In some versions, the repetitive waveform with varying frequencies may include a symmetrical triangular waveform. In some versions, instructions for controlling the generation of the sound signal include instructions for looping sound data representing the waveform of the repetitive waveform. Instructions for controlling the sensing of sound signals reflected from a user may include instructions for storing sound data sampled from a microphone. Instructions for controlling the processing of the sensed sound signal may include instructions for associating the generated sound signal with the sensed sound signal to check synchronization.

[0019] In some versions, instructions for processing the sensed sound signal may include a downconverter to generate data containing the respiratory signal. The downconverter may mix a signal representing the generated sound signal with the sensed sound signal. The downconverter may filter the output, which is the mixed output representing the generated sound signal and the sensed sound signal. The downconverter may window the filtered output, which is the mixed output representing the generated sound signal and the sensed sound signal. The downconverter may generate a frequency domain transform matrix of the windowed, filtered output, which is the mixed output representing the generated sound signal and the sensed sound signal. Processor-executable instructions for detecting the respiratory signal may extract amplitude and phase information from multiple channels of the data matrix generated by the downconverter. Processor-executable instructions for detecting the respiratory signal may also include processor-executable instructions for calculating multiple features from the data matrix. The multiple features may include any one or more of the following: (a) in-band squared divided by a full-band metric, (b) an in-band metric, (c) a kurtosis metric, and (d) a frequency domain analysis metric. Processor-executable instructions for detecting the respiratory signal may generate a respiratory rate based on the multiple features.

[0020] In some versions, the processor executable instructions may also include instructions for calibrating sound-based body motion detection by evaluating one or more characteristics of the electronic processing device. The processor executable instructions may also include instructions for generating sound signals based on the evaluation. The instructions for calibrating sound-based detection may determine at least one hardware, environmental, or user-specific characteristic. The processor executable instructions may also include instructions for operating a pet setup mode, wherein a frequency for generating sound signals may be selected based on user input, and one or more test sound signals may be generated.

[0021] In some versions, the processor executable instructions may also include instructions for interrupting the generation of a sound signal based on detected user interaction with the electronic processing device. Detected user interaction may include any one or more of the following: detecting movement of the electronic processing device using an accelerometer, detecting a pressed button, detecting a touchscreen, or detecting an incoming call. The processor executable instructions may also include instructions for initiating the generation of a sound signal based on the detection of no user interaction with the electronic processing device.

[0022] In some versions, the processor executable instructions may also include instructions for detecting large-amplitude body movements based on the processing of sensed sound signals from the user's reflections. In some versions, the processor executable instructions may also include instructions for processing audio signals sensed via a microphone to evaluate any one or more of ambient sounds, speech, and breathing sounds to detect user movement. In some versions, the processor executable instructions may also include instructions for processing breathing signals to determine any one or more of the following: (a) a sleep state indicating sleep; (b) a sleep state indicating wakefulness; (c) a sleep stage indicating deep sleep; (d) a sleep stage indicating light sleep; (e) a sleep stage indicating REM sleep.

[0023] In some versions, the processor executable instructions may also include the following instructions: for operating the setting mode to detect sound frequencies near the electronic processing device and select a frequency range of sound signals that differ from the detected sound frequencies. In some versions, the instructions for operating the setting mode may select a frequency range that does not overlap with the detected sound frequencies.

[0024] In some versions of this technology, the server can access any processor-readable medium described herein. The server can be configured to receive requests for downloading processor-executable instructions from a processor-readable medium to an electronic processing device over a network.

[0025] In some versions of this technology, a mobile electronic device or electronic processing device may include one or more processors; a speaker coupled to one or more processors; a microphone coupled to one or more processors; and a processor-readable medium of any processor-readable medium described herein.

[0026] Some versions of this technology relate to a server method that has access to any processor-readable medium described herein. The method may include receiving at the server a request to download processor-executable instructions from the processor-readable medium to an electronic processing device over a network. The method may also include sending the processor-executable instructions to the electronic processing device in response to the request.

[0027] Some versions of this technology relate to methods for detecting body motion using a processor in a mobile electronic device. This method may include utilizing the processor to access any processor-readable medium described herein. This method may include executing any processor-executable instructions on the processor-readable medium within the processor.

[0028] Some versions of this technology relate to methods for a processor that detects body movement using a mobile electronic device. The method may include controlling the generation of a nearby sound signal, including a user, via a speaker coupled to the mobile electronic device. The method may include controlling the sensing of a sound signal reflected from the user via a microphone coupled to the mobile electronic device. The method may include processing the sensed reflected sound signal. The method may include detecting a breathing signal from the processed reflected sound signal.

[0029] Some versions of this technology relate to methods for detecting motion and breathing using a mobile electronic device. The method may include transmitting an audio signal to a user via a speaker on the mobile electronic device. The method may include sensing a reflected audio signal, reflected from the user, via a microphone on the mobile electronic device. The method may include detecting breathing and motion signals from the reflected audio signal. The audio signal may be inaudible or audible. In some versions, before transmission, the method may involve modulating the audio signal using one of an FMCW modulation scheme, an FHRG modulation scheme, an AFHRG modulation scheme, a CW modulation scheme, a UWB modulation scheme, or an ACW modulation scheme. Optionally, the audio signal may be a modulated low-frequency ultrasonic audio signal. The signal may include multiple frequency pairs transmitted as frames. In some versions, upon sensing the reflected audio signal, the method may include demodulating the reflected audio signal. Demodulation may include performing a filtering operation on the reflected audio signal; and timing synchronization of the filtered reflected audio signal with the transmitted audio signal. In some versions, generating the audio signal may include performing a calibration function to evaluate one or more characteristics of the mobile electronic device; and generating the audio signal based on the calibration function. The calibration function can be configured to determine at least one hardware, environment, or user-specific characteristic. Filter operations can include high-pass filter operations.

[0030] Some versions of this technology relate to methods for detecting movement and respiration. The method may include generating an acoustic signal directed at a user. The method may include sensing an acoustic signal reflected from the user. The method may include detecting respiration and movement signals from the sensed reflected acoustic signal. In some versions, the generation, transmission, sensing, and detection may be performed at a bedside device. Optionally, the bedside device may be a therapeutic device, such as a CPAP device.

[0031] The methods, systems, devices, and apparatuses described herein can provide functional improvements in processors (such as processors in general-purpose or special-purpose computers, portable computer processing devices (e.g., mobile phones, tablets, etc.), respiratory monitors, and / or other respiratory devices utilizing microphones and speakers). Furthermore, the described methods, systems, devices, and apparatuses can provide improvements in the technical field of automating the management, monitoring, and / or prevention and / or assessment of respiratory and sleep conditions (including, for example, sleep apnea).

[0032] Of course, the parts of each aspect can form sub-aspects of this technology. In addition, sub-aspects and / or aspects of an aspect can be combined in any way and also constitute other aspects or sub-aspects of this technology.

[0033] Other features of the present technology will become apparent from consideration of the information contained in the following detailed description, abstract, drawings and claims. 4. Attached Figure Descriptions

[0035] This technology is illustrated in the accompanying drawings by way of example rather than limitation, wherein the same reference numerals denote similar elements, including:

[0036] Figure 1 An example processing device for receiving audio information from a sleeper is shown, which can be adapted to implement the process of this technology;

[0037] Figure 2 This is a schematic diagram of a system based on this technical example.

[0038] Figure 2A A high-level architecture block diagram of a system used to implement various aspects of this technology is shown.

[0039] Figure 3 This is a conceptual diagram of a mobile device configured according to some forms of this technology.

[0040] Figure 4 An exemplary FHRG sonar frame is shown.

[0041] Figure 5 An example of an AFHRG transceiver frame is shown.

[0042] Figure 6A , 6B Figures 6 and 6C show an example of a pulsed AFHRG frame.

[0043] Figure 7 This is a conceptual diagram of the (A)FHRG architecture.

[0044] Figure 8 Examples of isotropic omnidirectional antennas and directional antennas are shown.

[0045] Figure 9 An example of a sine filter response is shown.

[0046] Figure 10 An example of the attenuation characteristics of a sine filter is shown.

[0047] Figure 11 This is a graph illustrating an example of a baseband breathing signal.

[0048] Figure 12 An example of initial synchronization with the training tone frame is shown.

[0049] Figure 13 An example of an AFHRG frame with intermediate frequency is shown.

[0050] Figure 14 An example of an AToF frame is shown.

[0051] Figure 15A , 15B Figures 15C and 15C show the signal characteristics of an instance of an audible version of the FMCW ramp sequence.

[0052] Figure 16A , 16B Figure 16C shows the signal characteristics of an example of an FMCW sine curve.

[0053] Figure 17 The signal characteristics of an example sound signal (e.g., an inaudible sound) in the form of a triangular wave are shown.

[0054] Figure 18A and 18B This shows the effect of the triangle that cannot be heard (which is composed of) Figure 17 The detected breathing waveform (emitted by the speaker of a smart device) is demodulated, which has... Figure 18A The treatment of the upper slope shown and Figure 18B The downslope treatment shown in the diagram.

[0055] Figure 19 An exemplary method for processing a stream operator / module in an FMCW is shown, including “2D” (two-dimensional) signal processing.

[0056] Figure 20 An exemplary method for a downconversion operator / module is shown, which is Figure 19 It is part of the FMCW processing flow.

[0057] Figure 21 An exemplary method for a 2D analysis manipulator / module is shown, which can be Figure 19 It is part of the FMCW processing flow.

[0058] Figure 22A method for performing absence / presence detection via absence / presence operator / module is shown.

[0059] Figure 23A The graph shows several signals, illustrating the sensing operation over time in the “2D” data segment for various sensing ranges as a person “moves out” and enters the device 100.

[0060] Figure 23B yes Figure 23A A portion of the signal diagram shows the signal from Figure 23A Area BB.

[0061] Figure 23C yes Figure 23A A portion of the signal diagram shows the signal from Figure 23A The region CC. 5. Detailed Implementation

[0063] Before describing this technology in further detail, it should be understood that this technology is not limited to the specific instances described herein, and the specific instances described herein may be modified. It should also be understood that the terminology used in this disclosure is for the purpose of describing the specific instances described herein only and is not intended to be limiting.

[0064] The following description is provided in relation to various instances of shared common characteristics and / or features. It should be understood that one or more features of any example form can be combined with one or more features of another form. Furthermore, any single feature or combination of features of any form described herein can constitute further example forms.

[0065] 5.1 Screening, monitoring and diagnosis

[0066] This technology relates to systems, methods, and apparatus for detecting motion in a subject, including, for example, respiratory movements and / or cardiac-related chest movements, such as when the subject is asleep. Based on such respiratory and / or other motion detection, the subject's sleep state and apnea events can be detected. More specifically, mobile applications associated with mobile devices such as smartphones and tablets use mobile device sensors (such as speakers and microphones) to detect such motion.

[0067] Now for reference Figures 1 to 3 Describe an example system suitable for implementing this technology. For example... Figure 2The mobile device 100 or mobile electronic device is configured with an application 200 for detecting the movement of the object 110 and may be placed on a bedside table near the object 110. The mobile device 100 may be, for example, a smartphone or tablet with one or more processors. The processor may be configured, among other things, to perform the functions of the application 200, including generating and transmitting an audio signal, typically through the air (such as near the room where the device is located), the air being a typically open or unrestricted medium, receiving the reflection of the transmitted signal by sensing the reflection of the transmitted signal with, for example, a transducer (such as a microphone), and processing the sensed signal to determine body movement and respiratory parameters. Among other components, the mobile device 100 may include a speaker and a microphone. The speaker may be used to transmit the generated audio signal and the microphone may be used to receive the reflected signal. Alternatively, the sound-based sensing method of the mobile device may be implemented in or by other types of devices, such as bedside devices (e.g., respiratory therapy devices such as continuous positive airway pressure (e.g., "CPAP") devices) or high-flow therapy devices. Examples of such devices include pressure devices or blowers (e.g., motors and impellers in a volute), one or more sensors, and a central controller for the pressure device or blower. Reference can be made to the devices described in International Patent Publication No. WO / 2015 / 061848 (application number PCT / AU2014 / 050315), filed October 28, 2014, and International Patent Publication No. WO / 2016 / 145483 (application number PCT / AU2016 / 050117), filed March 14, 2016, which are incorporated herein by reference in their entirety.

[0068] Figure 2A A high-level architectural block diagram of a system for implementing various aspects of this technology is shown. This system can be implemented, for example, in a mobile device 100. A signal generation component 210 can be configured to generate an audio signal for transmission to an object. As described herein, the signal can be audible or inaudible. For example, as discussed in more detail herein, in some versions, the sound signal can be generated in the low-frequency ultrasonic range, such as about seventeen (17) kHz to twenty-three (23) kHz, or in the range of about eighteen (18) kHz to twenty-two (2) kHz. Such frequency ranges are generally considered herein to be the range of inaudible sounds. The generated signal can be transmitted via a transmitter 212. According to the invention, the transmitter 212 can be a mobile phone speaker. Once the generated signal has been transmitted, a receiver 214 can sense the sound waves, including those reflected from the object; the receiver 214 can be a mobile phone microphone.

[0069] As shown at 216, signal recovery and analog-to-digital signal conversion stages can occur. The recovered signal can be demodulated to retrieve signals representing respiratory parameters, as shown at 218, such as in a demodulation processing module. The audio components of the signal can also be processed, as shown at 220, such as in an audio processing module. Such processing can include, for example, passive audio analysis to extract sounds of breathing or other movements (including different types of activities such as rolling in bed, PLM (periodic leg movements), RLS (restless legs syndrome), etc.) such as snoring, panting, and wheezing. Audio component processing can also extract sounds from interfering sources such as speech, television, other media playback, and other ambient / environmental audio / noise sources. The output of the demodulation stage at 218 can be processed in-band at 222, such as by an in-band processing module, and out-of-band at 224, such as by an out-of-band processing module. For example, the in-band processing at 222 can target the portion of the signal band containing respiratory and cardiac signals. The out-of-band processing at 224 can be directed to those portions of the signal containing other components, such as large or fine movements (e.g., rolling, kicking, gesturing, moving around a room, etc.). The outputs of the in-band processing at 222 and the out-of-band processing at 224 can then be provided to the signal post-processing at 230, such as in one or more signal post-processing modules. The signal post-processing at 230 may include, for example, respiratory / cardiac signal processing at 232 (such as in a respiratory / cardiac signal processing module), signal quality processing at 234 (such as in a signal quality processing module), large-amplitude motion processing at 236 (such as in a body motion processing module), and absence / presence processing at 238 (such as in an absence / presence processing module).

[0070] Although not in Figure 2A As shown, however, the output of the signal post-processing at 230 can undergo a second post-processing stage, which provides sleep state, sleep score, fatigue indication, object recognition, chronic disease monitoring and / or prediction, sleep disorder breathing event detection, and other output parameters, such as assessments of any motion characteristics derived from the generated motion signal (e.g., respiratory-related motion (or its absence), cardiac-related motion, arousal-related motion, periodic leg movements, etc.). Figure 2AIn the example, optional sleep segmentation processing at point 241 is shown, such as in a sleep segmentation processing module. However, any one or more such processing modules / blocks (e.g., sleep scoring or segmentation, fatigue indication processing, object recognition processing, chronic disease monitoring and / or prediction processing, sleep disorder respiratory event detection processing, or other output processing, etc.) may be optionally added. In some cases, the signal post-processing function at point 230 or the second post-processing stage may be performed using any component, device, and / or method of the apparatus, system, and method described in any of the following patents or patent applications, each of which is incorporated herein by reference in its entirety: International Patent Application No. PCT / US2007 / 070196, filed June 1, 2007, entitled “Apparatus, System, and Method for Monitoring Physiological Markers”; International Patent Application No. PCT / US2007 / 070196, filed October 31, 2007, entitled “System and Method for Monitoring Cardiopulmonary Respiratory Parameters”. 07 / 083155; International Patent Application No. PCT / US2009 / 058020, filed September 23, 2009, entitled "Contactless and Minimal Contact Monitoring of Quality of Life Parameters for Assessment and Intervention"; International Patent Application No. PCT / US2010 / 023177, filed February 4, 2010, entitled "Apparatus, System, and Method for Monitoring Chronic Diseases"; International Patent Application No. PCT / AU2013 / 000564, filed March 30, 2013, entitled "Method and Apparatus for Monitoring Cardiopulmonary Health"; International Patent Application No. PCT / AU2013 / 000564, filed May 25, 2015, entitled... International patent applications filed on October 6, 2014, entitled "Methods and apparatus for monitoring chronic diseases" (PCT / AU2015 / 050273); on October 6, 2014, entitled "Fatigue monitoring and management system" (PCT / AU2014 / 059311); on September 19, 2013, entitled "System and method for determining sleep stages" (PCT / AU2013 / 060652); and on April 20, 2016, entitled "Detection and identification of humans by characteristic signals" (PCT / EP2016 / 058789); 20 PCT / EP2016 / 069496, filed August 17, 2016, entitled "Screening device for sleep-disordered breathing"; PCT / EP2016 / 069413, filed August 16, 2016, entitled "Digital range-gated radio frequency sensor"; PCT / EP2016 / 070169, filed August 26, 2016, entitled "System and method for monitoring and managing chronic diseases"; and U.S. Patent Application No. 15 / 079,339, filed March 24, 2016, entitled "Detection of periodic breathing".Therefore, in some instances, the processing of detected motion (including, for example, respiratory motion) can be used as a basis for determining any one or more of the following: (a) a sleep state indicating sleep; (b) a sleep state indicating wakefulness; (c) a sleep stage indicating deep sleep; (d) a sleep stage indicating light sleep; and (e) a sleep stage indicating REM sleep. In this regard, while the sound-related sensing technology of the present invention provides different mechanisms / processes for motion sensing, such as the use of speakers and microphones and the processing of sound signals, the principle of processing respiratory or other motion signals to extract sleep state / stage information can be achieved through the determination methods of these incorporated references, once a respiratory signal, such as the respiratory rate obtained using the sound sensing / processing methods described in this specification, is obtained.

[0071] 5.1.1 Mobile devices 100

[0072] Mobile device 100 can be adapted to provide an efficient and effective method for monitoring the breathing and / or other motion-related characteristics of a subject. When used during sleep, mobile device 100 and its associated methods can be used to detect the user's breathing and identify sleep stages, sleep states, transitions between states, sleep breathing disturbances, and / or other respiratory conditions. When used during wakefulness, mobile device 100 and its associated methods can be used to detect motion such as the subject's breathing (inspiratory, expiratory, apnea, and exhaustive rates) and / or cardiac impaction waveforms and subsequent exhaustive heart rate. These parameters can be used to control play (thereby guiding the user to reduce their breathing rate for relaxation purposes) or to assess the respiratory status of subjects with chronic diseases such as COPD, asthma, congestive heart failure (CHF), etc. The subject's baseline respiratory parameters change over time prior to a worsening / decompensated event. Respiratory waveforms can also be processed to detect temporary cessation of breathing (such as central apnea, or small chest movements of the obstructive airway seen during obstructive apnea) or decreased breathing (shallow breathing and / or reduced respiratory rate, such as those associated with hypoventilation).

[0073] Mobile device 100 may include integrated chips, memory, and / or other control instructions, data, or information storage media. For example, programming instructions containing the evaluation / signal processing methods described herein may be encoded on an integrated chip in the memory of the device or apparatus to form an application-specific integrated chip (ASIC). Such instructions may also, or alternatively, be loaded as software or firmware using appropriate data storage media. Alternatively, such processing instructions may be downloaded to the mobile device from a server, such as via a network (e.g., the Internet), so that when the instructions are executed, the processing device functions as a screening or monitoring device.

[0074] Therefore, the mobile device 100 may include, for example Figure 3 The multiple components shown. Among others, the mobile device 100 may include components such as a microphone or sound sensor 302, a processor 304, a display interface 306, a user control / input interface 308, a speaker 310, and a memory / data storage 312, as well as processing instructions using the processing methods / modules described herein.

[0075] One or more components of mobile device 100 may be integrated with or operatively coupled to mobile device 100. For example, microphone or sound sensor 302 may be integrated with or coupled to mobile device 100, such as via wired or wireless links (e.g., Bluetooth, Wi-Fi, etc.).

[0076] The memory / data memory 312 may contain a plurality of processor control instructions for controlling the processor 304. For example, the memory / data memory 312 may contain processor control instructions for causing the application program 200 to be executed by the processing instructions of the processing method / module described herein.

[0077] 5.1.2 Exercise and Respiratory Detection Process

[0078] Instances of this technology can be configured to use one or more algorithms or processes, implementable by application 200, to detect movement, breathing, and optional sleep characteristics using mobile device 100 while the user is asleep. For example, application 200 can be characterized by several subprocesses or modules. Figure 2 As shown, the application 200 may include an audio signal generation and transmission subprocess 202, a motion and biological body characteristic detection subprocess 204, a sleep quality characterization subprocess 206, and a result output subprocess 208.

[0079] 5.1.2.1 Generating and transmitting audio signals

[0080] According to some aspects of this technology, audio signals can be generated and sent to a user, such as using one or more tones described herein. A tone provides a pressure change in a medium (e.g., air) at a specific frequency. For the purposes of this specification, the generated tones (or audio signals or sound signals) may be referred to as “sound,” “acoustic,” or “audio” because they can be generated in a manner similar to audible pressure waves (e.g., through a loudspeaker). However, such pressure changes and tones should be understood herein as audible or inaudible, although they are characterized by any of the terms “sound,” “acoustic,” or “audio.” Therefore, the generated audio signal can be audible or inaudible, where the frequency threshold for audibility across populations varies with age. Typical “audio frequency” standards range from approximately 20 Hz to 20,000 Hz (20 kHz). The threshold for high-frequency hearing tends to decrease with age; middle-aged people typically cannot hear sounds above 15–17 kHz, while teenagers may hear sounds at 18 kHz. The most important frequencies for speech are approximately in the 250–6,000 Hz range. Typical consumer smartphones have speaker and microphone signal responses designed to be above 19-20kHz in many cases, with some extending to 23kHz and above (especially if the device supports sampling rates greater than 48kHz, such as 96kHz). Therefore, for most people, a signal in the 17 / 18 to 24kHz range will be usable and inaudible. Younger people who can hear 18kHz instead of 19kHz will be able to use the 19kHz to 21kHz band. It's worth noting that some household pets may be able to hear even higher frequencies (e.g., dogs can hear up to 60kHz and cats up to 79kHz).

[0081] Audio signals can include, for example, sinusoidal waveforms, sawtooth chirps, triangular chirps, etc. For background, the term "chirp" as used herein refers to a short-term non-stationary signal that can, for example, have a sawtooth or triangular shape, with linear or non-linear profiles. Several types of signal processing methods can be used to generate and sense audio signals, including, for example, continuous wave (CW) null, pulse CW null, frequency modulation CW (FMCW), frequency hopping range gating (FHRG), adaptive FHRG (AFHRG), ultra-wideband (UWB), orthogonal frequency division multiplexing (OFDM), adaptive CW, frequency shift keying (FSK), phase shift keying (PSK), binary phase shift keying (BSPK), quadrature phase shift keying (QPSK), and a generalized QPSK called quadrature amplitude modulation (QAM), etc.

[0082] According to some aspects of this technology, calibration functions or calibration modules can be provided for evaluating the characteristics of mobile devices. If the calibration function indicates that hardware, environment, or user settings require audible frequencies, encoding can be overlaid on a spread spectrum signal acceptable to the user, while still allowing active detection of any bio-motion signals within the sensing area. For example, audible or inaudible sensing signals may be “hidden” within audible sounds such as music, television, streaming sources, or other signals (e.g., pleasant repetitive sounds that can be optionally synchronized with detected breathing signals, such as waves crashing on the shore, which may be used to aid sleep). This is similar to audio steganography, where techniques such as phase encoding (where elements of the phase are adjusted to represent encoded data) are used to hide messages that must be perceptibly indistinguishable within audio signals.

[0083] 5.1.2.2 Motion and Biophysical Signal Processing

[0084] 5.1.2.2.1 Technical Challenges

[0085] To achieve a good motion signal to ambient noise ratio, several room acoustics issues should be considered. These issues may include, for example, reverberation, bedding, and / or other room acoustics. Additionally, specific monitoring equipment characteristics, such as the directionality of mobile devices or other monitoring equipment, should also be considered.

[0086] 5.1.2.2.1.1 Reverb

[0087] Reverberation can occur when sound energy is confined to a space with reflective walls, such as a room. Typically, when a sound is first generated, the listener first hears the direct sound from the sound source itself. Afterward, the user may hear discrete echoes caused by sound bouncing off the room's walls, ceiling, and floor. Over time, individual reflections may become indistinguishable, and the listener will hear a continuous reverberation that decays over time.

[0088] Wall reflections typically absorb very little energy at the target frequency. Therefore, it is the air, not the walls, that attenuates the sound. Typical attenuation may be <1%. It takes approximately 400ms in a typical room at an audio frequency to achieve 60dB of sound attenuation. Attenuation increases with frequency. For example, at 18kHz, the typical room reverberation time decreases to 250ms.

[0089] Several side effects are associated with reverberation. For example, reverberation causes room modes. Due to the reverberant energy storage mechanism, inputting acoustic energy into a room results in standing waves at resonant or preferred modal frequencies (nλ = L). For an ideal three-dimensional space with dimensions Lx, Ly, and Lz, the dominant mode is given by the following equation.

[0090]

[0091] These standing waves cause the loudness of a specific resonant frequency to vary at different locations within the room. This can lead to variations in signal level and associated attenuation at sensing components that receive sound signals, such as sound sensors. This is similar to multipath interference / attenuation from electromagnetic waves such as Wi-Fi (e.g., RF); however, room reverberation and associated artifacts (such as attenuation) are more severe for audio.

[0092] Room patterns can introduce design problems for sonar systems. These problems can include, for example, attenuation, 1 / f noise, and large-amplitude motion signals. Large-amplitude motion typically refers to large physiological movements, such as rolling over in bed (or a user getting into or out of bed, which are associated with absence before or after the event, respectively). In one implementation, motion can be considered a binary vector (movement or no movement), with the associated activity index referring to the duration and intensity of the movement (e.g., PLM, RLS, or bruxism are examples of non-large-amplitude motion activities because they involve grinding of limbs or jaws, rather than whole-body movement when changing position in bed). Signals received by sound sensor 302 (e.g., a microphone) may experience attenuation in signal strength and / or attenuation from reflected signals from the user. This attenuation may be due to standing waves caused by reverberation and can lead to signal amplitude variation problems. 1 / f noise (a signal whose power spectral density is inversely proportional to the signal frequency) may occur at the breathing frequency due to indoor airflow disturbing the room pattern and causing an increase in the noise floor, thus resulting in a reduced signal-to-noise ratio. Large-amplitude motion signals from any movement within a room are due to disturbances in the room pattern energy and can produce synthetic variations in the intensity and phase of the sensed / received signal. However, while this is a problem, it can be seen that by deliberately setting the room pattern, useful detection of large-amplitude (large) movements and more subtle activities can be performed, and practical respiratory analysis can be conducted on a primary source of movement within the room, such as a person in a bedroom.

[0093] The term sonar, as used herein, encompasses sound signals, acoustics, sound waves, ultrasound, and low-frequency ultrasound used for ranging and motion detection. Such signals can range from DC to 50 kHz or higher. Some processing techniques (such as FMCW physiological signal extraction) are also applicable to practical short-range (e.g., up to approximately 3 meters—such as for monitoring living spaces or bedrooms) RF (electromagnetic) radar sensors, such as those operating at 5.8 GHz, 10.5 GHz, 24 GHz, etc.

[0094] 5.1.2.2.1.2 Bedding

[0095] The use of bedding (such as comforters, quilts, blankets, etc.) can significantly attenuate sound in a room. In some ways, a comforter can attenuate a sound signal twice as it travels and returns from a sleeping person. The surface of a comforter also reflects sound signals. If breathing is observed on the surface of a comforter, the reflected signals can be used to monitor breathing.

[0096] 5.1.2.2.1.3 Mobile Phone Characteristics

[0097] For acoustic sensing applications similar to sonar, the placement of speakers and microphones on smartphones is not always optimal. Typical smartphone speakers have poor directionality towards the microphone. Often, smartphone audio is designed for human speech, not specifically for sonar-like acoustic sensing applications. Furthermore, speaker and microphone placement varies from smartphone model to model. For example, some smartphones have speakers on the back and microphones on the sides. As a result, the microphone cannot easily "see" the speaker signal unless it is first redirected by reflections. Additionally, the directionality of speakers and microphones increases with frequency.

[0098] Room modes can enhance the omnidirectionality of smartphone microphones and speakers. Reverberation and associated room modes create standing wave nodes at specific frequencies throughout the room. When motion disturbs all room mode nodes, any motion in the room can be seen at these nodes. When the smartphone microphone is located at a node or antinode, omnidirectional characteristics are achieved even if the microphone and speaker are directional, because the sound path is omnidirectional.

[0099] 5.1.2.2.2 Overcoming technical challenges

[0100] According to aspects of the invention, specific modulation and demodulation techniques can be applied to reduce or eliminate the impact of identified and other technical challenges.

[0101] 5.1.2.2.2.1 Frequency Hopping Range Gating (FHRG) and Adaptive Frequency Hopping Range Gating (AFHRG)

[0102] Frequency hopping range gating is a modulation and demodulation technique that utilizes a series of discrete tones / frequencies occupying a specified frequency range. Each tone, tone pair, or multi-tone pulse in the tone sequence is sent for a specific duration, defined by the range requirement. Changing the specific duration of such tone pulses or tone pairs results in a change in the detection range. A series of tone pairs or tones produces tone frames or time slot frames. Each time slot can be considered a tone time slot. Therefore, a frame can include multiple time slots, each of which can contain a tone or tone pair. Typically, to facilitate improved inaudibility, the duration of each tone pair can be equal to the duration of the time slots in the frame. However, this is not required. In some cases, the time slots in the frame can have equal or unequal widths (durations). Therefore, the tone pairs in the frame can have equal or unequal durations within the frame. Sound modulation can include guard bands (silence periods) between tones and then between frames to allow for better separation of sensed / received tones, thus enabling range gating. FHRG can be used with available audio. FHRG can apply frequency hopping with frequency jitter and / or timing jitter. Timing jitter means that the timing of frames and / or internal time slots can vary (e.g., the time interval of a time slot varies, or the start time of a frame or the start time of a time slot within a time slot can vary), such as to reduce the risk of room pattern establishment and the risk of interference between other sonar sensors in the environment. For example, in the case of a four-slot frame, while the frame width can remain constant, timing jitter can allow at least one smaller time slot width (e.g., a shorter duration tone pair) and at least one larger time slot width (e.g., a longer duration tone pair) in the frame, while the other time slot widths can remain constant. This can provide slight variation within a specific range detected in each time slot of the frame. Frequency jitter can allow multiple “sensors” (e.g., two or more sensing systems in a room) to coexist in common proximity. For example, by employing frequency jitter, the frequency of the time slots (tones or tone pairs) is shifted, making interference between “sensors” statistically insignificant during normal operation. Both time and frequency jitter can reduce the risk of interference from non-sonar sources in the room / environment. Synchronous demodulation requires a precise understanding of the alignment between sequences, for example, by using a special training sequence / pulse and then using pattern detection / matching filters to align frames and recover the actual offset.

[0103] Figure 4An exemplary FHRG sonar frame is shown. As illustrated, eight individual pulse sonar signals (similar to those in a radar system) can be transmitted in each frame. Each pulse can contain two orthogonal pulse frequencies, and each pulse can be 16 ms long. As a result, 16 individual frequencies can be transmitted in each 128 ms transceiver frame. The block (frame) is then repeated. It can utilize the orthogonal frequencies of OFDM to optimize frequency usage in limited bandwidth and enhance the signal-to-noise ratio. Orthogonal Dirac comb frequencies also allow for frame timing jitter and tone frequency jitter to help reduce noise. Tone pairing shapes the resulting pulses, and this shape contributes to inaudibility.

[0104] Adaptive Frequency Hopping Range Gating (AFHRG) systems maintain the frequency of the changing tone frequency sequence over time using a defined length (e.g., 16 or 8 ms) of tone before switching to another tone (four tones in total). This generates tone blocks, which are then repeated at 31.25 Hz. For AFHRG systems, the tone pattern varies on each frame. These patterns can change continuously. Therefore, the frequency pattern of a frame can differ from the frequency patterns of other frames in a series of frames. Thus, patterns at different frequencies can change. Frames can also have different frequencies within a frame, such as within a time slot, or relative to different time slots. Furthermore, by adjusting the frequency of each frame within a fixed frequency band, each frame can be adapted to mitigate attenuation. Figure 5 An example of a single 32ms transceiver frame used in an AFHRG system is shown. Figure 5 As shown, four separate zero-difference pulse sonar signals can be transmitted in each frame. Each signal has a flight time of 8 ms. Within this 8 ms flight time, the range, including both directions, is 2.7 meters, the effective actual frame is 1.35 meters, and the pulse repetition frequency is 31.25 Hz. Furthermore, each pulse can include two pulse frequencies (e.g., substantially simultaneous tones), such as... Figure 5 As shown, eight separate pulse zero-difference signals are generated in a single 32ms transceiver frame.

[0105] AFHRG can be configured to optimize available bandwidth using Dirac comb frequencies and quadrature pulse pairs. Each "pair" can reside in a time slot within a frame. Therefore, multiple frames can contain many "pairs" in repeating or non-repeating patterns. Individual frequency pairs can be used for each time slot, and linear or Costas code frequency hopping can be used. Time slots can be determined based on desired range detection. For example, as... Figure 5 As shown, for a detection range of 1.3m, t ts = 8ms. The frequency pair can be generated as: {A Sin(ω1t)–A Sin(ω2t)}, where By adjusting the frequency of each frame within a fixed frequency band, each frame can be adapted to mitigate attenuation once the requirement of frequency equivalence is maintained. Therefore, adaptation is used to maximize the available SNR (signal-to-noise ratio) from any object present in the detection range.

[0106] Each frequency pair can be selected to optimize the desired bandwidth (e.g., 1kHz bandwidth (18kHz to 19kHz)) to provide maximum isolation between frequencies, maximum isolation between pulses, and / or minimize inter-pulse transients. These optimizations can be implemented for n frequencies, each with a frequency spacing df and a time slot width tts, such that:

[0107] df = f n -f n-1 =BW / n

[0108] and

[0109]

[0110] For an 8ms time slot duration and a 1kHz bandwidth (e.g., fn = 18,125Hz and fn-1 = 18,000Hz), the frequency separation df becomes:

[0111] df = 125Hz = 1kHz / 8

[0112] and

[0113]

[0114] Zero crossover can be achieved by using trigonometric identities: Make each time slot contain a frequency of The amplitude of the sinusoidal pulse. Therefore, the duration of a frame's time slot, or each time slot, can be equal to one divided by the difference between the frequencies of the tone pairs. Thus, within a frame, the duration of a time slot can be inversely proportional to the frequency difference between the frequencies of the tones of the tone pairs in the time slot.

[0115] Figure 6A , 6B Figures 6C and 6C show an example of a 5×32ms pulse frame used in an example AFHRG system. Figure 6A Graph 602 in the figure shows 5 × 32 ms pulse frames in x-axis time and y-axis amplitude. Graph 602 is a time-domain representation of five 32 ms frames, where each frame contains four (4) time slots—a total of 5 × 4 = 20 tone pairs. The tone pairs are located in each of time slots 606-S1, 606-S2, 606-S3, and 606-S4. Its envelope varies, illustrating a “real-world” speaker and microphone combination that may be slightly less sensitive at higher frequencies. Figure 6B The formation is considered in more detail in Figure 604. Figure 6A The time slots of the frames shown are illustrated in Graph 604. Figure 6A The frequency domain representation of a single 32ms frame is shown. It illustrates four time slots 606-S1, 606-S2, 606-S3, and 606-S4, where each time slot contains two tones 608T1 and 608T2 (tone pairs). Graph 604 has frequency on the x-axis and arbitrary y-axis. In one instance, the tones of the tone pairs in the frame are each distinct sound tone (e.g., a different frequency) in the range of 18000Hz to 18875Hz. Other frequency ranges can be implemented, such as for inaudible sounds, as discussed in more detail herein. Each tone pair of multiple tone pairs in the frame is generated sequentially (serially) within the frame. The tones of the tone pairs are generated substantially simultaneously in a common time slot of the frame. Figure 6A and Figure 6B The pitch of the time slot can be regarded as a reference. Figure 6C The graph 610 is a temporal view of a single tone pair (e.g., tone 608T1, 608T2) within a frame, showing their formation and occurrence within the time slots of the frame. Graph 610 has an x-axis for time and a y-axis for amplitude. In this example, the amplitude of the tone rises and falls within the time period of the time slot. In the example of graph 610, the tone amplitude ramp begins at zero amplitude at the start of the time slot and slopes down to zero amplitude at the end of the time slot. Such tones reaching zero amplitude simultaneously at the beginning and end of a time slot can improve inaudibility between tone pairs in adjacent time slots with the same end and start time slot amplitude characteristics.

[0116] Now refer to Figure 7 The methods of the processing modules shown describe aspects of an FHRG or AFHRG (here referred to as "(A)FHRG") architecture used for processing audio signals. Figure 7 As shown, the reflected signal can be sensed, filtered by a module of high-pass filter 702, and input to Rx frame buffer 704. The buffered Rx frames can be processed by IQ (in-phase and quadrature) demodulator 706. According to some aspects of the invention, each of the frequencies of n (where n can be, for example, 4, 8, or 16) can be demodulated to baseband as I and Q components, where the baseband represents motion information (breathing / body movement, etc.) corresponding to changes in the sensed distance detected by audio reflections (e.g., transmitted / generated and received / sensed audio signals or tone pairs). The intermediate frequency (IF) stage 708 is a processing module that outputs single I and Q components from multiple signals, the single I and Q components undergoing optimization in the module "IQ Optimization" at 710 to produce a single combined output.

[0117] The multi-IQ input information at 710 can be optimized and condensed into a single IQ baseband signal output at the algorithm input stage. A single IQ output (actually a combined signal from the I and Q components) can be derived based on the selection of candidate IQ pairs, chosen based on the one with the highest signal quality, such as the clearest respiratory rate. For example, the respiratory rate can be detected from the baseband signal (candidate signal or combined signal), as described in International Application WO2015006364, which is incorporated herein by reference in its entirety. The single IQ output can also be the average of the input signals, or essentially the average or median of the derived respiratory rates. Therefore, such a module at 710 can include summation and / or averaging processes.

[0118] The AFHRG architecture allows for the optional addition of an IF stage. Two possible IF stage methods exist: (i) comparing an earlier (e.g., 4ms (0.5m range)) signal with a later (e.g., 4ms (0.5m range)) signal, and (ii) comparing the phase of the first tone with the phase of the second tone in the tone pair. Because the stage compares level or phase changes in the sensed sound signal over a time-of-flight (ToF) period, it can be used to eliminate common-mode signal artifacts (such as motion and 1 / f noise problems). The phase folding recovery module 712 receives the I and Q component outputs of the optimization stage at 710 and, after processing, outputs the IQ baseband output. Minimizing folding in the demodulated signal may be desirable, as this significantly complicates breathing “in-band” detection. Folding can be minimized using I / Q combination techniques (such as arctangent demodulation), real-time center tracking estimation, or more standard dimensionality reduction methods (such as principal component analysis). Folding occurs when chest movement spans approximately 9 mm half a wavelength (e.g., at 20°C, the CW frequency is 18 kHz, the speed of sound is approximately 343 m / s, and the wavelength is approximately 19 mm). A dynamic center frequency strategy can be used to automatically correct for folding—for example, to reduce the severity or probability of this behavior. In this case, if an octave or abnormally jagged breathing pattern (morphology) is detected, the system can shift the center frequency to a new frequency to push the I / Q channels back to equilibrium without folding. Whenever there is movement, reprocessing is generally required. Changes in audio frequency are designed to be inaudible (unless a masking tone is used). If the movement is λ / 4, we can see it on one channel. If the movement is greater than λ / 2, then we are likely to see it on two channels.

[0119] In the FHRG implementation, the frame / tone pair modulator at 714 is processor-controlled (using implicit multi-frequency operation of frames, etc.). In adaptive implementations such as AFHRG or AToF, the system also includes modules for attenuation detection and explicit frequency shifting operations not present in the FHRG implementation. The module of attenuation detector 718 provides a feedback mechanism to adjust system parameters (including frequency shift, modulation type, frame parameters (such as the number of tones), time and frequency intervals, etc.) to optimally detect motion, including breathing, under varying channel conditions. Attenuation can be detected directly from changes in amplitude modulation (extracted via envelope detection) in the sensed sound signal and / or from changes in specific tone pairs. The attenuation detector can also receive auxiliary information from subsequent baseband signal processing, although this may add some processing latency; this can be used to correlate the quality / morphology of the actually extracted respiratory signal with the current channel conditions to provide better adaptability to maximize the useful signal. The attenuation detector can also process I / Q pairs (pre-baseband respiratory / heart rate analysis). By configuring the unattenuated Tx (transmit) waveform, the limited transmit signal power of the loudspeaker can be optimized / optimally utilized to maximize the useful information received by the sound sensor / receiver Rx (receive), as well as demodulation / further processing to baseband. In some cases, if the system is found to be more stable over a longer period (e.g., to avoid multipath variations in time-scale attenuation similar to respiratory rate, which may introduce "noise" / artifacts in the desired demodulated respiratory rate band), the system can select a slightly suboptimal (in terms of short-term SNR) frequency group for Tx.

[0120] Of course, if one or more individual tones are used instead of tone pairs, a similar architecture can be used to implement an adaptive CW (ACW) system.

[0121] 5.1.2.2.2.2 Detailed Information on FHRG Modulation and Demodulation Modules

[0122] 5.1.2.2.2.2.1 FHRG Sonar Equation

[0123] The sound pressure signal generated by the smartphone speaker is reflected by the target and returned to be sensed at the smartphone microphone. Figure 8 Examples of isotropic omnidirectional antenna 802 and directional antenna 804 are shown. For an omnidirectional source, the sound pressure level (P(x)) will decrease with distance x, as follows:

[0124]

[0125] The frequency of a smartphone speaker is 18kHz, therefore:

[0126]

[0127] Where γ is the loudspeaker gain, its value is between 0 and 2, and is usually >1.

[0128] The target (e.g., a user in bed) has a specific cross-section and reflectivity. Reflections from the target are also directional.

[0129]

[0130] Where α is the reflector attenuation, its value is between 0 and 1, and is typically <0.1.

[0131] β is the reflector gain, and its value is between 0 and 2, typically >1.

[0132] σ is the sonar cross section.

[0133] As a result, the transmitted signal of sound pressure P0 will return to the smartphone with the following attenuation level when reflected at a distance d:

[0134]

[0135] The smartphone microphone will then see a portion of the reflected sound pressure signal. The percentage will depend on the effective microphone area Ae:

[0136]

[0137] In this way, a small portion of the transmitted sound pressure signal is reflected by the target and returned to be sensed at the smartphone microphone.

[0138] 5.1.2.2.2.2.2 FHRG motion signal

[0139] A person within the range of the FHRG system at a distance d from the transceiver will reflect the active sonar transmitted signal and generate a received signal.

[0140] Due to the modem signal, for any frequency f nm The sound (pressure wave) produced by a smartphone is:

[0141]

[0142] For any single frequency, a signal with a distance d reaching the target is as follows:

[0143]

[0144] The reflected signal reaching the smartphone microphone is:

[0145]

[0146] If the target distance moves with a sinusoidal change of approximately d, then

[0147] d(x, t, w) b A b )=d0+A b Sin(2πf b t+θ)

[0148] in:

[0149] Respiratory rate: w b =2πf b

[0150] Respiratory amplitude: A b

[0151] Respiratory phase: θ

[0152] Nominal target distance: d0

[0153] Because the maximum respiratory displacement A b Since the distance d from the target is relatively small, its impact on the received signal amplitude can be ignored. As a result, the smartphone microphone's breathing signal becomes:

[0154]

[0155] Because an idealized respiratory motion signal (which may resemble a cardiac motion signal, albeit at a different frequency and displacement) can be considered a sine function of a sine function, but for the same displacement peak, it will have regions of maximum and minimum sensitivity to the peak amplitude.

[0156] To properly recover the signal, it is beneficial to use a quadrature phase receiver or similar signal to mitigate the sensitivity zero point.

[0157] Then, the I and Q baseband signals can be used as phasors I+jQ, to:

[0158] 1. Restore the RMS and phase of the respiratory signal

[0159] 2. Restore the direction of motion (phasor direction)

[0160] 3. Recovering folding information by detecting when phasors change direction.

[0161] 5.1.2.2.2.2.3 FHRG demodulator / mixer ( Figure 7 (Modules at positions 705 and 706 in the middle)

[0162] FHRG softmodem (a software-based modulator / demodulator used in acoustic sensing (sonar) detection) sends modulated sound signals through a speaker and senses echo reflections from a target through a microphone. Figure 7An example of a soft modem architecture is shown. The soft modem receiver, which processes sensed reflections into received signals, is designed to perform a variety of tasks, including:

[0163] • Signal received via high-pass filtering at audio frequencies

[0164] • Frame timing synchronization with modulated audio

[0165] • Demodulate the signal to baseband using a quadrature phase synchronous demodulator.

[0166] • If present, demodulate the IF (intermediate frequency) component.

[0167] • In-phase and quadrature baseband signals generated by low-pass filtering (e.g., Sinc filtering)

[0168] • Regenerate I and Q baseband signals from multiple I and Q signals at different frequencies.

[0169] soft modems (e.g.) Figure 7 The modules at 705 and 706 utilize a local oscillator (shown as "Asin(ω1t)" in the module at 706) whose frequency is synchronized with the frequency of the received signal to recover phase information. The module at 705 (frame demodulator) selects and separates each tone pair for subsequent IQ demodulation of each pair—that is, defines each (ωt) for IQ demodulation of the tone pair. In-phase (I) signal recovery will be discussed. The quadrature components (Q) are identical but have a 90-degree phase shift using a local oscillator (shown as "ACos(ω1t)" in the module at 706).

[0170] Given that the pressure signal sensed at the microphone is accurately regenerated into an equivalent digital signal and input to a high-pass filter (e.g., filter 702), the audio signal received by the smartphone's soft modem is:

[0171]

[0172] in:

[0173] Amplitude A is determined by sonar parameters:

[0174] The distance d is modulated by the target motion: d = d0 + A b Sin(2πf b t+θ)

[0175] The static clutter reflection / interference signal is represented as follows:

[0176]

[0177] When correctly synchronized, for any received signal, the output signal of the in-phase demodulator is:

[0178]

[0179] This demodulation operation follows the trigonometric identity:

[0180]

[0181] When the local oscillator and the received signal have the same angular frequency w1, this will be reduced to:

[0182]

[0183] After low-pass filtering (such as...) Figure 7 (As shown in the LPF) Removing 2w1t produces:

[0184]

[0185] In this way, when correctly synchronized, the low-pass filter synchronization phase demodulator output signal for any received signal is:

[0186]

[0187] 5.1.2.2.2.4 FHRG Demodulator Sinc Filter

[0188] The smartphone soft modem utilizes an oversampling and averaging Sinc filter at 706 as the "LPF" of the demodulator module to remove unwanted components and decimate the demodulated signal to the baseband sampling rate.

[0189] The received frame contains two tone pairs transmitted together in the same tone burst (e.g., where two or more tone pairs are played simultaneously), and also includes tone jumps transmitted in subsequent tone burst times. When the demodulated return signal, such as f(ij), is received, unwanted demodulated signals must be removed. These include:

[0190] • Baseband signals spaced by frequency: |f ij -f nm |

[0191] • Due to the sampling process in |2f nm -f s | and |f nm -f s |The aliasing components at the two locations

[0192] • Demodulated audio components: 2f nm

[0193] ·f nmThe audio carrier at that point may pass through the demodulator due to nonlinearity.

[0194] For this purpose, a low-pass filter (LPF) such as a Sinc filter can be used. The Sinc filter is almost ideal for this purpose because it has a filter response that is zero at all unwanted demodulation components in the design. The transfer function of such a moving average filter of length L is:

[0195]

[0196] Figure 9 An example of the Sinc filter response is shown, while Figure 10 The worst-case attenuation characteristics of the 125Hz Sinc filter due to averaging 384 samples over an 8ms period are shown.

[0197] 5.1.2.2.2.2.5 Demodulator baseband signal

[0198] The smartphone's soft modem sends modulated tone frames. Each frame contains multiple tone pairs (e.g., two tone pairs with four tones), which are transmitted together in the same tone burst, and also include tone jumps transmitted in subsequent tone bursts.

[0199] First, we will discuss the demodulation of tone pairs. Because tone pairs are transmitted, reflected, and received simultaneously, their only demodulation phase difference is caused by the frequency difference between the tone pairs.

[0200] Before “flight time”:

[0201]

[0202] Since there is no reflected signal, the demodulator outputs for both tone pairs are at DC level due to near-end crosstalk and static reflections.

[0203] After the "time of flight" period, this DC level receives contributions from the moving target component. The demodulator output now receives the following:

[0204] a) For the in-phase (I) demodulator signal output of each audio pair frequency, we obtain:

[0205]

[0206]

[0207] b) For the quadrature (Q) demodulator signal output of each audio pair frequency, we obtain:

[0208]

[0209]

[0210] The sum of each of these samples in the demodulated tone burst received signal, ∑, results in a frequency "Sinc" filter. This filter is designed to average the desired received tone burst to enhance the signal-to-noise ratio, and at each f... nm -f 11 The unwanted demodulated tone burst frequencies at intervals produce zeros in the transfer function so that all of these frequencies are rejected.

[0211] This sum occurs during the full-tone burst cycle, but the component caused by motion only occurs after the flight period until the end of the tone burst cycle, i.e.:

[0212] from until

[0213] As a result of the time-of-flight factor, another signal attenuation factor is introduced in signal recovery. The attenuation coefficient is:

[0214]

[0215] This is the amplitude decreasing linearly to the maximum range D.

[0216] When a frame contains n consecutive tone pairs, the demodulated audio signal is then averaged at the frame rate to produce a baseband signal with the following sampling rate:

[0217]

[0218] It samples once per second, providing 2n different I and Q baseband demodulation signals.

[0219] 5.1.2.2.2.2.6 Demodulator I mn and Q mn

[0220] The result of this demodulation and averaging operation is that the tone is demodulated to the audio frequency into four independent baseband signal samples, namely I11, I12, Q11, and Q22, at a sampling rate of samples per second.

[0221]

[0222] During subsequent periods of the frame, this operation is repeated for different frequencies that are frequency-separated from the previous pair of frequencies. To maintain optimal suppression through the Sinc filter function.

[0223] This sampling process is repeated for each frame to generate multiple baseband signals. In the given example, it generates 2×4×I and 2×4×Q baseband signals, each with its own phase characteristics due to the trigonometric identities:

[0224]

[0225] exist Figure 11 The text describes an example of the phase change of the first tone pair with a moving target at a distance of 1m.

[0226] The AFHRG architecture provides tight timing synchronization facilitated by the time, frequency, and envelope amplitude characteristics of tone pairs, such as by utilizing frame synchronization module 703. Audio Tx and Rx frames can be asynchronous. During initialization, it may be expected that the mobile device maintains perfect Tx-Rx synchronization. According to some aspects of the invention, and as... Figure 12 As shown, initial synchronization can be achieved using training tone frames, which can contain nominal frame tones, as shown at 1202. The Rx signal can be IQ demodulated by IQ demodulator 706, as described above regarding... Figure 7 As shown. The envelope can then be calculated (e.g., by taking the absolute value and then low-pass filtering, or using a Hilbert transform), as shown at 1204, and timing can be extracted using level threshold detection, as shown at 1206. The shape of the A Sin(ω1t)-A Sin(ω2t) pulse contributes to autocorrelation and can be configured to improve precision and accuracy. The Rx cyclic buffer index can then be set for proper synchronization timing, as shown at 1208. In some aspects of the invention, training tone frames may not be used; instead, Rx can be activated, and then TX can be activated again after a short time, and a threshold can be used to detect the start of a TX frame after a short recording silence period. The threshold can be selected to be robust to background noise, and prior knowledge of the rise time of the Tx signal can be used to create an offset from the true start of Tx, thereby correcting for threshold detection occurring during said rise time. When considering the estimated envelope of a signal, a detection algorithm that can be adjusted based on the relative signal level can be used to perform peak and "valley" benchmark detection (where valleys are near zero, due to the use of absolute values). Alternatively, only peaks can be processed, as valleys are more likely to contain noise due to reflections.

[0227] In some cases, devices (such as mobile devices) may lose synchronization due to jitter or loss of audio samples caused by other activities or processing. Therefore, a resynchronization strategy, such as using frame synchronization module 703, may be required. According to some aspects of the invention, a periodic training sequence can be introduced to periodically confirm that synchronization is good / true. An alternative method (which may or may optionally use such a periodic training sequence, but not this method) is to iteratively cross-correlate a known Tx frame sequence (typically a single frame) along segments of the Rx signal until a maximum correlation above a threshold is detected. The threshold can be chosen to be robust to expected noise and multipath interference in the Rx signal. This index of maximum correlation can then be used as an index for estimating the synchronization correction factor. Note that desynchronization in the demodulated signal can manifest as a small or significant step in the baseline (usually the latter). Unlike the stepping response seen in some real large-motion signals, desynchronized segments do not retain useful physiological information and are removed from the output baseband signal (which may result in the loss of several seconds of signal).

[0228] Therefore, synchronization typically utilizes prior knowledge of the transmitted signal, performing initial alignment using techniques such as envelope detection, and then selecting the correct burst using the cross-correlation between the received signal Rx and the known transmitted signal Tx. Ideally, synchronization checks for loss should be performed periodically, and optionally, data integrity tests should be included.

[0229] In an instance of this integrity test, for the case of a new synchronization being performed, a check can be performed to compare the new synchronization with the timing of one or more previously synchronized times. It can be expected that the time difference between candidate synchronization times in the sample will be equal to an integer frame within an acceptable timing tolerance. If this is not the case, synchronization loss may have occurred. In this case, the system can initiate reinitialization (e.g., by utilizing a new training sequence (unique within the local time interval), introducing a defined period of "silence" in Tx, or some other marker). In such cases, data sensed since the potential desynchronization event may be flagged as suspicious and may be discarded.

[0230] This desynchronization check can be performed continuously on the device, establishing useful trend information. For example, if regular desynchronization is detected and corrected, the system can adjust the audio buffer length to minimize or stop this undesirable behavior, or modify the processing or memory load, such as deferring some processing until the end of the sleep session and buffering the data used for such processing instead of executing complex algorithms in real-time or near real-time. This resynchronization method is effective and useful for various Tx types (FHRG, AFHRG, FMCW, etc.)—especially when using complex processors such as those in smart devices. As described, a correlation can be performed between the envelope of the reference frame and the estimated envelope of the Rx sequence; however, a correlation can also be performed directly between the reference frame and the Rx sequence. This correlation can be performed in the time domain or the frequency domain (e.g., as a cross-spectral density or cross-coherence measurement). There are also cases where the Tx signal stops generating / playing for a period of time (when expected to play or due to user interaction, such as selecting another application or receiving a call on a smart device), and resynchronization is paused until the Rx sees a signal above a minimum level threshold. Desynchronization may occur over time if the device's main processor and audio codec are not fully synchronized. For suspected long-term desynchronization, a mechanism for generating and playing training sequences can be used.

[0231] Note that desynchronization can produce a signal that looks like a DC offset and / or a step response in the signal, and the average (average) level or trend can change after a desynchronization event. Furthermore, desynchronization can also occur due to external factors such as loud ambient noise (e.g., brief pulses or durations); for example, a loud noise, banging on a device containing a microphone or a nearby table, very loud snoring, coughing, sneezing, shouting, etc., can cause the Tx signal to be submerged (the noise source has a similar frequency content to Tx, but with a higher amplitude), and / or Rx to be submerged (entering saturation, hard or soft clipping, or possibly activating automatic gain control (AGC)).

[0232] Figure 13 An example of a 32ms frame with an intermediate frequency is shown. According to some aspects of the invention, the IF stage can be configured to compare an early 4ms (0.5m range) signal with a later 4ms signal. According to other aspects of the invention, the IF stage can be configured to compare the phase of a first tone with the phase of a second tone. Because the IF stage compares level or phase changes in the received signal during flight time, common-mode signal artifacts such as motion and 1 / f noise can be reduced or eliminated.

[0233] Room reverberation can generate room modes at resonant frequencies. The AFHRG architecture allows for fundamental frequency shifting once frame frequency separation is maintained. Pulse pair frequencies can be shifted to frequencies that do not generate room modes. In other embodiments, frame pulse pair frequencies can be skipped to reduce mode energy accumulation. In other embodiments, a long, non-repeating pseudo-random sequence can be used throughout the frequency-specific reverberation time to make frame pulse pair frequencies skip, preventing homodyne receivers from seeing reflections. Frame frequencies can be jittered to mitigate interference from modes, or sinusoidal frame shifting can be used.

[0234] Using an AFHRG system offers numerous advantages. For example, such systems allow for adaptive transmission frequencies to mitigate room reverberation. The system improves SNR by using active transmissions with a repetitive pulse mechanism and varying frequencies, unlike the typical "quiet period" required by pulse continuous wave radar systems. In this respect, the frequency achieves a "quiet period" by being different from subsequent tone pairs in the frame compared to previous tone pairs. The subsequent frequency allows the propagation of reflected sound from earlier, different frequency tone pairs to be determined as the subsequent tone pairs run (propagate). Time slots provide range gating for the first sequence. Furthermore, the architecture uses dual-frequency pulses to further improve SNR. Including the intermediate frequency phase provides defined and / or programmable range gating. The architecture is designed to allow the use of Costas or pseudo-random frequency hopping codes to slow the sampling period, thus mitigating room reverberation. Additionally, the quadrature Dirac comb frequencies allow for frame timing jitter and tone frequency jitter to further help reduce noise.

[0235] It should be noted that wider or narrower bandwidths can also be selected in AFHRG. For example, if a particular mobile phone can transmit and receive with good signal strength up to 21 kHz, and the user of the system can hear frequencies up to 18 kHz, the system can choose to use a 2 kHz bandwidth of 19-21 kHz. It can also be seen that these tone pairs can be hidden within other transmitted (or detected ambient) audio content and adapt to changes in other audio content to maintain masking of the user. Masking can be achieved by transmitting audible (e.g., from, for example, about 250 Hz upwards) or inaudible tone pairs. When the system operates, for example, above about 18 kHz, any generated music source can be low-pass filtered below about 18 kHz; conversely, where the audio source cannot be directly processed, the system tracks the audio content and injects a predicted number of tone pairs, adapted to maximize SNR and minimize audibility. In effect, elements of the existing audio content provide masking for tone pairs—where tone pairs are adapted to maximize the masking effect; auditory masking is where the perception of tone pairs is reduced or eliminated.

[0236] Another approach when generating or playing acoustic signals (e.g., music) for use with (A)FHRG, FMCW, UWB, etc. (with or without masking) is direct carrier modulation. In this case, the sensed signal is directly encoded into the played signal by adjusting the signal content (specifically, by changing the amplitude of the signal). Subsequent processing involves demodulating the input (received) signal and the output (generated / transmitted) signal by mixing and then low-pass filtering to obtain the baseband signal. The phase change in the carrier signal is proportional to motion near the sensor (such as breathing). It is important to note that if the music naturally changes amplitude or frequency at breathing frequencies due to potential beat differences, this will increase noise on the demodulated signal. Furthermore, during a significant period of playback, additional signals (such as amplitude-modulated noise) may need to be injected to continue detecting breathing during this time, or the system may simply ignore the baseband signal during such silent playback periods (i.e., to avoid the need to generate a "filler" modulated signal). It can also be seen that, instead of (or in addition to) amplitude-modulated audio signals (such as music signals), the signal can be phase-coded with defined phase segments, then the received signal is demodulated, and the phase changes of the received signal during the coding interval are tracked to recover the baseband signal.

[0237] 5.1.2.2.2.3 Adaptive Flight Time

[0238] According to some aspects of the present invention, an Adaptive Time-of-Flight (AToF) architecture can be used to mitigate indoor acoustic problems. The AToF architecture is similar to AFHRG. Like AFHRG, AToF uses a normal silence period to enhance SNR through repetitive pulses. For example, four separate zero-difference pulse sonar signals can be transmitted in each frame. Each signal can have a flight time of 8 ms. In an 8 ms flight time, the range including both directions is 2.7 meters, the effective actual frame is 1.35 meters, and the pulse repetition frequency is 31.25 Hz. Additionally, each pulse can include two pulse frequencies, such as... Figure 14 As shown, eight separate pulse zero-difference signals are generated within a single 32ms transceiver frame. Unlike the AFHRG architecture, AToF uses pulses lasting 1ms instead of 8ms.

[0239] Similar to AFHRG, AToF can use Dirac comb features and orthogonal pulse pairs to help shape pulses. AToF can use separate frequency pairs for each time slot and can use linear or Costas code frequency hopping. Time slots can be determined by the desired range detection. Frequency pairs can be generated as {A Sin(ω1t) - A Sin(ω2t)}, where By adjusting its frequency within a fixed frequency band, each frame can be adjusted to reduce attenuation once the frequency equivalence requirement is met.

[0240] In summary, some advantages of frequency pairs (compared to multiple individual tones at different frequencies) include:

[0241] • Frequency flexibility: It allows the use of any frequency in any time slot (once...) This allows frequency shift adaptability to mitigate room patterns and attenuation.

[0242] • S / N (Impact on SNR): Increased signal-to-noise ratio due to improved bandwidth utilization (allowing scaling of pulse edges)

[0243] • Transient mitigation: Provides near-ideal zero-crossover transition, which mitigates frequency hopping transients in both the transmitter and receiver, and allows for true pseudo-random frequency hopping. This results in an inaudible system.

[0244] • Facilitates synchronous recovery: The unique frequency and amplitude profile of the shaped pulse improves synchronous detection and accuracy (i.e., clear shape).

[0245] • Pulse and frame timing flexibility: Allows for any time slot duration to facilitate the “variable detection range” feature.

[0246] • Enhance jitter: Enable both frequency jitter and time slot jitter to reduce noise.

[0247] • Compact bandwidth: Naturally produces a very compact bandwidth, so it is inaudible even without filtering (filters introduce phase distortion).

[0248] • MultiTone: Allows the use of multiple tones in the same time slot to enhance S / N and reduce attenuation.

[0249] 5.1.2.2.2.4 Frequency Modulated Continuous Wave (FMCW)

[0250] In another aspect of the invention, a frequency modulated continuous wave (FMCW) architecture can be used to alleviate some of the aforementioned technical challenges. FMCW signals are commonly used to provide location (i.e., distance and velocity) because FMCW enables range estimation and thus provides distance gating. An example of a sawtooth chirp audible in everyday life is like the chirping of birds, while a triangular chirp sounds, for example, like a police siren. FMCW can be used in RF sensors and also in acoustic sensors, such as those implemented in typical smart devices (such as smartphones or tablets), using built-in or external speakers and microphones.

[0251] It should be noted that audible guide tones (e.g., played at the start of recording or when the smart device is moved) can be used to convey the relative “loudness” of inaudible sounds. On smart devices with one or more microphones, select an outline to minimize (and ideally make unnecessary) any additional processing in the software or hardware CODEC, such as disabling echo cancellation, noise reduction, automatic gain control, etc. For some phones, the camera microphone or main microphone is configured to be in “speech recognition” mode, which can provide good results, or select an “unprocessed” microphone feed, such as a microphone feed that can be used for virtual reality applications or music mixing. Depending on the phone, “speech recognition” modes, such as those available in some Android OS revisions, can disable effects and preprocessing on the microphone (which is desirable).

[0252] FMCW allows for spatial tracking to determine where a person is breathing (if they have moved) and to separate the breathing of two or more people within the sensor's range (i.e., to recover the breathing waveforms of each object from different ranges).

[0253] FMCWs can have chirps, such as ramp sawtooth, triangular, or sinusoidal shapes. Therefore, unlike the pulses of (A)FHRG type systems, FMCW systems can generate repetitive sound waveforms with varying frequencies (e.g., inaudible). It is important to match the frequency variations at the zero-crossing points of the signal where possible—that is, to avoid jump discontinuities in the generated signal, which can cause unwanted harmonics that may be audible and / or unnecessarily stress the speaker. Ramp sawtooths can be like... Figure 15B The ramp shown is in the form of an upward ramp (from lower frequency to higher frequency). When repeated, by repeating from low to high, the ramp upward forms an upward sawtooth pattern (ramp sawtooth) of waveform from low to high. However, this ramp can alternatively be downward (from higher frequency to lower frequency). When repeated, by repeating from high to low, the ramp downward forms a downward sawtooth pattern (inverted ramp sawtooth) of waveform from high to low. Furthermore, although this ramp can be approximated linearly using a linear function as shown, in some versions, the rise in frequency can form a curve (increasing or decreasing), for example, a polynomial function between low and high (or high and low) waveform frequency changes.

[0254] In some versions, the FMCW system can be configured to change one or more parameters of the form of the repeating waveform. For example, the system can change (a) the position of any one or more frequency peaks in the repeating portion of the waveform (e.g., earlier or later peaks in the repeating portion). The parameter for this change can be a change in the slope of the frequency change of the waveform's ramp (e.g., upward and / or downward slope). The parameter for this change can be a change in the frequency range of the repeating portion of the waveform.

[0255] This section outlines a specific method for implementing this inaudible signal on smart devices using FMCW triangular waveforms.

[0256] For an FMCW system, the theoretical range resolution is defined as: V / (2*BW), where V is velocity (such as sound) and BW is bandwidth. Therefore, for a V = 340 m / s (at room temperature) and an 18-20 kHz chirped FMCW system, a target separation resolution of 85 mm can be achieved. Each of one or more moving targets (such as the breathing motion of an object) can then be detected with a finer resolution (similar to a CW system), assuming (in this example) each object is separated by at least 85 mm in the sensor's field of view.

[0257] Depending on the relative frequency sensitivity of the speaker and / or microphone, the system may optionally use emphasis on the transmitted signal TxFMCW waveform so that each frequency is corrected to have the same amplitude (or other modifications to the transmitted signal chirp to correct nonlinearities in the system). For example, if the speaker response decreases with increasing frequency, the 19kHz component of the chirp is produced with a higher amplitude than the 18kHz component, making the actual Tx signal have the same amplitude across the entire frequency range. Avoiding distortion is, of course, important, and the system can examine the received signal to obtain the adjusted transmitted signal to determine if distortion (such as clipping, unwanted harmonics, sawtooth waveforms, etc.) occurs—that is, adjusting the Tx signal to be as close to linear as possible at the largest possible amplitude to maximize the SNR. The deployed system may reduce the waveform and / or decrease the volume to meet the target SNR to extract breathing parameters (a tradeoff between maximizing SNR and hard-driving the speaker)—while maintaining a signal that is as linear as possible.

[0258] Depending on channel conditions, chirp can also be adjusted by the system—for example, by adjusting the bandwidth used, chirp is robust to strong interference sources.

[0259] 5.1.2.2.2.5 FMCW Chirp Types and Inaudibility

[0260] FMCW has a period of 1 / 10 ms = 100 Hz (and associated harmonics) at audible frequencies, and its hum is clearly audible unless further digital signal processing is performed. The human ear has an incredible ability to distinguish very low-amplitude signals (especially in quiet environments such as a bedroom), even if a high-pass filter shifts the component down to -40 dB (unwanted components may need to be filtered down to below -90 dB). For example, the start and end of the chirp can be de-emphasized (e.g., using Hamming, Hanning, Blackman, etc. windows to aim at the chirp) and then re-emphasized in subsequent processing for correction.

[0261] Chirps can also be isolated—for example, by repeating them only once every 100-200 ms to introduce a much longer silence period between chirps than the duration itself; a side benefit is that it minimizes the detection of reverberation-related standing waves. The disadvantage of this approach is that it uses less available transmit energy than a continuously repeating chirp signal. The trade-off is that reducing the audible clicks and reducing the continuous output signal does have some benefits in terms of driving devices (e.g., smart devices such as telephone speakers) speakers and amplifiers less effortfully; for example, coil speakers can be affected by transients if driven at maximum amplitude for a long period.

[0262] Figure 15 shows an example of an audible version of the FMCW ramp sequence. Figure 15(a) shows the standard chirp in a graph illustrating the amplitude versus time relationship. Figure 15(b) shows the spectrogram of the standard audible chirp, while Figure 15(c) shows the spectrum of the standard audio chirp.

[0263] As mentioned above, FMCW tones can also use a sine profile instead of a ramp. Figure 16 shows an example of an FMCW sine profile with a high-pass filter applied. Figure 16A The relationship between signal amplitude and time is shown. Figure 16B The spectrum of the signal is shown. Figure 16C The spectrum of the signal is shown. The following example uses a start frequency of 18 kHz and an end frequency of 20 kHz. A “chirp” of 512 samples with a sinusoidal profile is generated. The intermediate frequency is 19,687.5 Hz, with a deviation of + / - 1,000 Hz. The sequence of these chirps has continuous phase across the chirp boundaries, which helps to make the sequence inaudible. A high-pass filter can optionally be applied to the resulting sequence; however, it should be noted that this operation may cause phase discontinuities, which in turn make the signal easier to hear than less audible as expected. The FFT of the shifted received signal is multiplied by the conjugate of the FFT of the transmitted sequence. The phase angle is extracted and expanded. The optimal straight line is found.

[0264] As mentioned above, it is not ideal to filter a sinusoidal signal to make it inaudible, as this would distort the phase information. This is likely due to phase discontinuities between scans. Therefore, an inaudible FMCW sequence using triangular waveforms can be used.

[0265] A triangular waveform signal can be generated as a phase-continuous waveform, preventing the speaker from producing a clicking sound. Although continuous within a scan, a phase difference may exist at the start of each scan. The equation can be modified to ensure that the phase at the start of the next scan begins as a multiple of 2π, allowing individual portions of the waveform to be cycled in a phase-continuous manner.

[0266] A similar approach can be applied to ramp chirped signals, although the frequency jump discontinuities from 20kHz to 18kHz (assuming the chirp is from 18kHz to 20kHz) can strain amplifiers and speakers in commercial smart devices. Due to the added downscan, the triangular waveform provides more information than the ramp, and the more "gentle" frequency changes of the triangular waveform are less harmful to telephone hardware than the discontinuous jumps of the ramp.

[0267] The following is an example equation that can be implemented in the signal generation module to generate this phase-continuous triangular waveform. The phase of the triangular chirp for the up and down scans and time index n can be calculated by the following closed expression:

[0268]

[0269]

[0270] in

[0271] f s Sampling rate

[0272] f1 and f2 represent low frequency and high frequency, respectively.

[0273] N is the total number of samples in the upper and lower scans, i.e., the number of samples in each scan. N samples (assuming equal distribution N).

[0274] The final stage at the end of the next scan is by Give

[0275] To bring the phase sine back to zero at the start of the next scan, we provide:

[0276]

[0277] For example, suppose N and f2 are fixed. Therefore, we choose m to make f1 as close as possible to 18kHz:

[0278]

[0279] For N = 1024 and f2 = 20kHz, m = 406, then f1 = 18,062.5kHz

[0280] Demodulation of this triangular waveform can be considered as demodulating the upper and lower scans of the triangle separately, or by processing the upper and lower scans simultaneously, or in fact by processing frames of multiple triangular scans. Note that processing only the upper scan is equivalent to processing a single frame of the ramp (sawtooth).

[0281] By using up and / or down scanning, such as in respiratory sensing, inspiration can be separated from the expiratory portion of breathing (i.e., if inspiration or expiration occurs, it is known at a certain point in time). Figure 17 An example of a triangular signal is shown, and... Figure 18A and Figure 18B The demodulation of the detected respiratory waveform is shown in the figure. For example... Figure 17 As shown, the smart device's microphone has recorded an inaudible triangle emitted by the same smart device's speaker. Figure 18A The graph shows the demodulation results of the reflection of the triangular signal relative to the upper scan of the signal, and shows the respiratory waveform extracted at 50.7 cm. Figure 18B The graphs show the demodulation results of the triangular signal reflection relative to the lower scan of the signal, illustrating the respiratory waveform extracted at 50.7 cm. As shown in these figures, Figure 18B The recovered signal is inverted relative to the demodulated unfolded signal of the corresponding upscan due to the phase difference.

[0282] Alternatively, an asymmetrical “triangular” ramp can be considered instead of a symmetrical triangular ramp, where the duration of the upscan is longer than that of the downscan (and vice versa). In this case, the shorter duration is (a) to maintain inaudibility, which may be subject to transient trade-offs due to the upscan ramp alone, and (b) to provide a reference point in the signal. Demodulation is performed during the longer upscan (as it will be used for a sawtooth waveform with a quiescent period (but the “quiescent period” is the shorter duration of the downscan); this may allow for reduced processing load (if necessary) while maximizing the Tx signal used and maintaining an audible, phase-continuous Tx signal.

[0283] 5.1.2.2.2.6 FMCW Demodulation

[0284] Figure 19The diagram illustrates an example flow of a sonar FMCW processing method. It shows a block or module providing front-end signal generation at 1902, signal reception at 1904, synchronization at 1906, down-conversion at 1908, and “2D” analysis at 1910, including generating signals or data representing estimates of any one or more of activity, motion, presence / absence, and respiration. Based on these data / signals, wakefulness and / or sleep segmentation can be provided at 1912.

[0285] The transmitted chirp can be correlated with the received signal, particularly for checking synchronization in the module at 1906. As an example, the resulting narrow pulse can be used to determine fine details caused by spikes following correlation operations. There is often a correlation between the width of the signal spectrum and the width of the correlation function; for example, a relatively wide FMCW signal associated with an FMCW template produces a narrow correlation peak. Additionally, correlating the standard chirp with the response at the receiver provides an estimate of the echo impulse response.

[0286] In FMCW, the system effectively considers the set of all responses across the entire frequency range as the “strongest” large-amplitude response (e.g., while some frequencies in the chirp may experience severe attenuation, other frequencies may provide a good response, so the system considers the set).

[0287] An exemplary method, such as that of one or more processors in a mobile device, for recovering a baseband signal (e.g., large movement or breathing) as part of an FMCW system is described below. According to some aspects of the invention, a chirp sequence can be transmitted, such as at 1902 using an FMCW transmit generation module to operate one or more speakers. The chirp sequence can be, for example, an inaudible triangular acoustic signal. An input received signal (such as a received signal operated by a receive module with one or more microphones at 1904) can be synchronized with the transmitted signal using, for example, the peak correlation described above. A continuous resynchronization check of the chirps or blocks (e.g., several chirps) can be performed, for example, at 1906 using a synchronization module. A mixing operation, such as by multiplication or summation, can then be performed for demodulation, such as at 1908 using a down-conversion module, wherein one or more scans of the received signal can be multiplied by one or more scans of the transmitted signal. This produces the frequency of the sum and difference between the transmitted and received frequencies.

[0288] An example of a signal transmitted from a telephone using a transmitting module at 1902 is a phase-continuous triangular chirp (to ensure inaudibility) and is designed to loop sound data according to a speaker controlled by a processor based on a module running on a computing device (e.g., a mobile device). An exemplary triangular chirp uses 1500 samples at 48kHz for both up and down scanning. This gives an up scan time of 31.25ms (which is likely the same for the down scan), providing a baseband sampling rate of 32Hz if up and down scans are used continuously (in contrast to the audio sampling rate), or a baseband sampling rate of 16Hz if they are averaged (or only alternating scans are used).

[0289] When using more than one scan (e.g., four scans), the calculation can be repeated on a block-by-block basis by moving the scan one step at a time (i.e., introducing overlap). In this case, the relevant peaks can be estimated by optionally extracting the relevant envelope using methods such as filtering, maximum preservation, or other methods (e.g., Hilbert transform); outliers can also be removed—for example, by removing those that fall outside the average of the relative peak positions by one standard deviation. When extracting each scan of the received waveform, the starting index can be determined using the pattern of the relative peak positions (outlier removal).

[0290] The correlation metric may sometimes be lower than others, possibly due to intermittent unwanted signal processing affecting the transmitted waveform (causing a decrease or "dipping" in the chirp), or because a loud noise in the room environment causes the signal to be "drowned out" for a short period of time.

[0291] At 1906, finer phase level synchronization (albeit with greater computational complexity) can be performed by examining the correlation with a template that shifts in the phase in degree increments up to a maximum of 360 degrees. This estimates the phase level offset.

[0292] Unless the timing of the system is controlled very precisely (i.e., the level of delay in the system is known), a synchronization step is required to ensure that the demodulation works correctly and does not produce noise or erroneous output.

[0293] As mentioned earlier, the down-conversion processing in module 1908 generates a baseband signal for analysis. Down-conversion can now be performed by synchronizing data transmission and reception. This module utilizes various submodules or processes to process the synchronized transmission and reception signals (generated and reflected sounds) to extract the "beat" signal. An example of this audio processing is... Figure 20As shown in the diagram; here, I and Q transmit signals can be generated, where the phase offset determined using a fine-grained synchronization method is used for mixing. A nested loop can be repeated at 2002 for mixing in the module or process, for example, at 2004: the outer loop iterates over each scan of the current received waveform held in memory and is applied by the access process at 2006 to mix at 2004 (each chirp and overbite is considered a separate scan); then the inner loop iterates first over the I channel, then the Q channel.

[0294] Another example of the system ( Figure 20 (Not shown in the image) Optionally, multiple received scans are buffered, a median is calculated, and a moving median number containing the multiple scans is provided. This median is then subtracted from the current scan to provide another method for signal cancellation.

[0295] For each scan (up or down), the relevant portion of the received waveform is extracted (using the synchronization sample index as a reference) and mixed with the phase-shifted transmit chirp at 2004. Using the transmitted and generated received portions, the waveforms are mixed (i.e., multiplied) together at 2004. For example, it can be seen that for a down scan, the received down scan can be mixed with the TX down scan, the flipped received down scan can be mixed with the TX up scan, or the flipped received down scan can be mixed with the flipped TX down scan.

[0296] The output of the mixing operation (e.g., the received waveform multiplied by a reference-aligned waveform (e.g., a reference chirp)) can be low-pass filtered to remove higher sum frequencies, as in filtering processes or modules such as those in 2008. As an example, an implementation could use a triangular chirp of 18–20–18 kHz with a sampling rate of 48 kHz; for such a frequency band, the higher sum frequencies in the mixing operation are effectively undersampled and produce approximately 11–12 kHz of aliasing components (for Fs = 48 kHz, the sum is approximately 36 kHz). The low-pass filter can be configured to ensure that if this aliasing component occurs in a particular system implementation, it is removed. The components of the remaining signal depend on whether the target is static or moving. Assuming an oscillating target at a given distance (e.g., a breathing signal), the signal will contain a “beat” signal and a Doppler component. The beat signal is the difference in frequency position between the transmitted and received scans due to time-delayed reflections from the target. The Doppler is the frequency shift caused by a moving target. Ideally, this beat should be calculated by selecting only the portion of the scan from the point of arrival of the received signal to the end of the transmitted scan. However, this is difficult to achieve in practice for complex targets such as humans. Therefore, a full scan can be used. According to some aspects of the invention, more advanced systems can utilize adaptive algorithms that take advantage of the fact that once the object's location is clearly identified, an appropriate portion of the scan is taken, further improving the system's accuracy.

[0297] In some cases, the mixed signal can optionally have its mean (average value) removed, such as in the detrending process at 2010 (average detrending), and / or be high-pass filtered (such as in the HPF process at 2012) to remove unwanted components that would ultimately appear in the baseband. Other detrending operations can be applied to the mixed signal, such as linear detrending, median detrending, etc. The level of filtering applied at 2008 can depend on the quality of the transmitted signal, the strength of the echo, the number of interference sources in the environment, static / multipath reflections, etc. For higher quality conditions, less filtering can be applied because the low-frequency envelope of the received signal may contain breathing and other motion information (as well as the ultimately demodulated baseband signal). For lower quality, more challenging conditions, noise in the mixed signal may be significant, and more filtering is desired. Indications of the actual signal quality (e.g., breathing signal quality) can be fed back as feedback signals / data to these filtering processes / modules to select the appropriate filtering level at the filtering module stage at 2008.

[0298] Following the downconversion at 1908, a two-dimensional (2D) analysis is performed at 1910 using a complex FFT matrix to extract respiration, presence / absence, large-amplitude motion, and activity. Thus, the downconverter can produce a frequency domain transform matrix. For this type of process, blocks of mixed (and filtered) signals are windowed (e.g., using a Hanning window module at 2014). A Fourier transform (such as a Fast Fourier Transform or FFT) is then performed at 2016 to estimate and produce the “2D” matrix 2018. Each row is a scanned FFT, and each column is an FFT window. It is these FFT windows that are transformed into ranges, hence the term “range windows.” The matrices are then processed by a 2D analysis module, resulting in a set of orthogonal matrices for each matrix based on I-channel or Q-channel information.

[0299] refer to Figure 21 The module shown illustrates an example of this processing. As part of the 2D analysis at 1910, body (including limb) motion and activity detection can be performed on this data. The motion and activity estimation of the processing module extracts information about the motion and activity of the object from the analysis of the composite (I+jQ) mixed signal in the frequency domain at 2102. It can be configured to operate based on body (e.g., rolling or limb movement) motion being uncorrelated with and much larger than chest displacement during breathing. Independent of (or dependent on) this motion processing, multiple signal quality values ​​are calculated for each range window. These signal quality values ​​are calculated for the amplitude and unfolded phase of both the I and Q channels.

[0300] The processing at 2102 generates an "activity estimate" signal and a motion signal (e.g., body motion markers). The "body motion marker" (BMF) generated by the activity / motion processing at 2102 is a binary marker indicating whether motion has occurred (output at 1 Hz), while the "activity count," generated at 2104 in conjunction with the counter processing module, is a measure of activity for each epoch between 0 and 30 (output at 1 / 30th of a second). In other words, the "activity count" variable captures the amount (severity) and duration of body motion within the defined period of a 30-second block (updated every 30 seconds), while the "motion" marker is simply a yes / no update per second of motion. The original "activity estimate" used to generate the activity count is a measure estimated based on the correlation between the mixed chirped signals during non-motion periods and their uncorrelatedness during motion periods. Since the mixed scans are represented in the complex frequency domain, the modulus or absolute value of the signal can be obtained before analysis. Pearson correlation is calculated on a series of mixed scans spaced 4 chimes apart within the target range. For a waveform transmitted with 1500 samples at 1500 kHz, this interval corresponds to 125 ms. The chirp interval determines the velocity distribution being examined and can be adjusted as needed. The correlation output is inverted, for example, by subtracting from 1, so the decrease in correlation is related to the increase in metric. The maximum value per second is calculated and then passed through a short 3-tap FIR (Finite Impulse Response) boxcar mean filter. For some devices, the response to motion calculated using this method can be inherently nonlinear. In this case, the metric can be remapped using a natural log function. The signal is then detrended by subtracting the minimum value observed in the first N seconds (corresponding to the maximum observed correlation) and passed to a logistic regression model with a single weight and bias term. This produces a 1 Hz raw activity estimate signal. Binary motion markers are generated by applying a threshold to the raw activity estimate signal. Activity counts are generated by comparing the 1 Hz raw activity estimates above the threshold to a nonlinear mapping table, where values ​​per second are selected from 0 to 2.5. These are then summed over a 30-second time period and constrained to a value of 30 to generate the activity count for each epoch.

[0301] Therefore, a one-second (1Hz) motion marker has been created, along with an activity intensity estimate and an associated activity count up to a maximum of 30 (the sum of activities within a 30-second epoch), which is associated with the activity / motion module at 2012 and the activity counter at 2104.

[0302] Another aspect of the 2D processing at 1910 is the calculation of the distance between the object and the sensor (or, in the case of two or more objects within the sensor's range, several distances). This can be achieved using a range window selection algorithm, which processes the resulting two-dimensional (2D) matrix 2018 to produce a 1D matrix (or matrix) at the target window. Although not in Figure 21 As shown, however, the output of the module used for this selection process can inform the processing module of 2D analysis, including, for example, any extraction processing at 2016, computational processing at 2108, and breathing decision processing at 2110. In this regard, unfolding the phase at the resulting beat frequency will return the desired oscillatory motion of the target (assuming no phase unfolding problem occurs). Optional phase unfolding error detection (and potential correction) can be performed to mitigate possible jump discontinuities in the unfolded signal due to total motion. However, these “errors” can actually provide useful information, i.e., they can be used to detect motion in the signal (typically large motion, such as large amplitude movements). If the input signal for phase unfolding is only very low amplitude noise, a flag can optionally be set for the duration of this “below-threshold” signal, since the unfolding may be invalid.

[0303] The start and end times, as well as the location (i.e., range) of motion, can be detected in a 2D unfolded phase matrix using various methods; for example, this can be achieved using 2D envelope extraction and normalization, followed by thresholding. The normalization timescale is configured to be large enough to exclude phase changes caused by breathing-like motions (e.g., >>k*12 seconds for a minimum respiratory rate defined as 5 bpm). Detection of such motions can serve as an input trigger for window selection processes; for example, a window selection algorithm can "lock" to a previous range window during the motion period and hold that window until a valid breathing signal is subsequently detected in or near that window. This reduces computational overhead in cases where an object breathes quietly, moves in bed, and then lies quietly; they are still seen within the same range window. Therefore, the start and / or end positions of the detected motion can be used to limit the search range for window selection. This can also be applied to multiple object monitoring use cases (e.g., simultaneously monitoring two objects in a bed where the motion of one object might obscure the breathing signal of the other, but such a locked window selection can help recover the valid range windows (and breathing signal extraction) for both objects—even during a large motion of one object).

[0304] This method generates a signal produced by oscillations formed by phase unrolling. The composite output produces a demodulated baseband IQ signal containing body motion (including respiration) data from a living person or animal for subsequent processing. For an example of a triangular Tx waveform, the system can process the linear up and down scans separately and treat them as separate IQ pairs (e.g., to increase SNR and / or separate inspiratory and expiratory breathing). In some implementations, optionally, the beat frequency from each (up and down scan) can be processed to average any possible coupling effects in the Doppler signal, or to produce a better range estimate.

[0305] The baseband SNR metric can be configured to compare the respiratory noise band (0.125-0.5 Hz, equivalent to 7.5 to 30 breaths per minute – although this may expand to approximately 5-40 breaths per minute depending on usage – i.e., the primary respiratory signal content) to the motion noise band (i.e., the band primarily containing motion other than breathing) from 0.125-0.5 Hz, where the baseband signal is sampled at 16 Hz or higher. Baseband content below 0.083 Hz (equivalent to 5 breaths per minute) can be removed using a high-pass filter. Heart rate information can be contained in a band of approximately 0.17-3.3 Hz (equivalent to 25 to 200 heartbeats per minute).

[0306] During sonar FMCW processing, the difference between beat estimates can optionally be used to exclude (remove) static signal components, i.e., by subtracting adjacent estimates.

[0307] The range window selection algorithm in this method can be configured to track range windows corresponding to the positions of one or more objects within the sensor's range. This may require partially or completely searching possible range windows based on user-provided location data. For example, if the user notices their distance relative to the device while sleeping (or sitting), the search range can be narrowed. Typically, the blocks or epochs of data considered might be 30 seconds long and non-overlapping (i.e., one range window per epoch). Other implementations may use longer epoch lengths and employ overlap (i.e., multiple range interval estimates per epoch).

[0308] When using the detection of an effective respiratory rate (or the probability that the respiratory rate exceeds a predetermined threshold), this limits the minimum possible window length, as at least one respiratory cycle should be able to contain the respiratory rate within an epoch (e.g., a very slow respiratory rate of 5 breaths per minute means one breath every 12 seconds). At the upper limit of the respiratory rate, 45-50 breaths / minute is a typical limitation. Therefore, when using epoch-level spectral estimation (e.g., removing the mean, optionally windowing, and then performing an FFT), relevant peaks in the desired respiratory band (e.g., 5-45 breaths / minute) are extracted, and the power in the band is compared to the power of the full signal. The relative power of the maximum peak (or the maximum peak-to-mean ratio) across the band is compared to a threshold to determine if candidate respiratory frequencies exist. In some cases, several candidate windows with similar respiratory frequencies can be found. These may occur due to reflections within the room, producing distinct breathing in multiple ranges (this may also be related to FFT sidelobes).

[0309] Processing the triangular waveform can help mitigate this uncertainty in the range (e.g., due to Doppler coupling), but the system can choose to use it if the reflection contains a better signal than the direct path over a period of time. In cases where the user has a comforter / duvet and the room contains soft furniture (including, for example, a bed and curtains), indirect reflections are unlikely to produce a higher signal-to-noise ratio than the direct component. Since the dominant signal is likely to be a reflection from the comforter surface, it may appear slightly closer to a person's chest when considering the actual estimated range.

[0310] It can be seen that knowledge of the estimated range window from previous epochs can be used to inform subsequent range window searches, in an effort to reduce processing time. For systems that do not require near real-time (epoch-by-epoch) analysis, longer timescales can be considered.

[0311] 1.1 Signal Quality

[0312] Various signal quality metrics can be calculated as part of the FFT2D metric, such as Figure 21 As shown in the calculator processing of module 2108, the signal quality of respiration (i.e., the detection and correlation quality of respiration) is determined, which can be achieved through methods as previously described. Figure 20 The filtering at 2008 can be considered. These metrics can also be considered in the respiratory decision processing at 2110. Optionally, Figure 21 The 2D analysis processing can include processes or modules for determining absence / presence or sleep stages, based on some intermediate values ​​in the output and description.

[0313] Regarding the calculator processing at point 2108, various metrics, such as those used for respiration estimation, can be determined. In this example, Figure 21 The text indicates four metrics:

[0314] 1. "I2F" - In-band (I) squared divided by the full frequency band

[0315] 2. "Ibm - In-band (I) unique metric"

[0316] 3. "Kurt" - a measure of kurtosis based on covariance.

[0317] 4. "Fda" - Frequency Domain Analysis

[0318] 1. "I2F"

[0319] The I2F metric can be calculated as follows (a similar method applies to Q (orthogonal) channels that can be invoked in BandPwrQ):

[0320]

[0321] Where “inBandPwrI” is the total power, for example, from about 0.1 Hz to 0.6 Hz (e.g., within a selected target breathing band), and “fullBandPwrI” is the total power outside this range. This metric is based on the broadband assumption that the full-band and in-band power subsets are similar, and the full-band power is used to estimate the in-band noise.

[0322] 2. "IBM"

[0323] The next metric does not share the same assumptions as I2F and provides an improved estimate of the in-band signal and noise. It does this by finding the peak power within the band and then calculating the power around it (e.g., by taking three FFT windows on each side of the peak). It then multiplies that signal power by the peak itself and divides by the next peak (for wide sinusoidal signals). In cases where the breath contains strong harmonics, the choice of window can be re-evaluated so that the harmonics are not confused with the noise components. The noise estimate then includes everything outside the signal subband but still within the breath band:

[0324]

[0325] 3. Kurt

[0326] Kurtosis provides a measure of the "tails" of a distribution. This can provide a means of separating respiratory signals from other non-respiratory signals. For example, signal quality output can be the reciprocal of the kurtosis of the covariance of a DC (using IIR HPF) signal removed at a specified distance from the peak covariance. In cases of poor signal quality, this metric is set to "invalid".

[0327] 4. "Fda"

[0328] A “Fda” (Frequency Domain Analysis) can be performed; such statistics can be calculated using a 64-second overlapping data window with a step size of 1 second. Causal relationships are calculated using retrospective data. This process can detect respiratory rates within a specific respiratory rate window. For example, respiratory rates can be detected as described in International Application WO2015006364, which is incorporated herein by reference in its entirety. For example, respiratory rates can be detected within a rate window corresponding to 6 to 40 breaths per minute (bpm), corresponding to 0.1–0.67 Hz. This frequency band corresponds to the actual human respiratory rate. Therefore, “in-band” refers to the frequency range of 0.1–0.67 Hz. Each 64-second window can contain 1024 data points (64 seconds at 16 Hz). Therefore, the algorithm calculates a 512-point (N / 2) FFT for each (I and Q) data window. The results of these FFTs are used to calculate the in-band spectral peak (which can then be used to determine the respiratory rate), as described below. The in-band frequency range is used to calculate the respiratory rate for each 64-second window, as described below.

[0329] Other types of "Fda" analysis are shown below.

[0330] For typical heart rates, alternative frequency bands can also be considered (e.g., where HR of 45 to 180 beats per minute corresponds to 0.75-3 Hz).

[0331] It can also determine the spectral peak ratio. It identifies the maximum in-band and out-of-band peaks and uses this information to calculate the spectral peak ratio. This can be understood as the ratio of the maximum in-band peak to the maximum out-of-band peak.

[0332] It is also possible to determine the in-band variance. The in-band variance quantifies the power in the frequency band. In some cases, this can also be used for subsequent presence / absence detection.

[0333] Spectral peaks in the target frequency band are identified by implementing a quality factor that combines the spectral power level at each window with the distance to adjacent peaks and the frequency of the window. The window with the highest value of the aforementioned quality factor is selected.

[0334] As part of the 2D analysis at 1910, four metrics, already outlined in 2108, are calculated to identify one or more living individuals within the sensing vicinity of the mobile device and to track whether they move to different ranges (distance from the sensor), actually leave the sensing space (go to the restroom), or return to the sensing space. This provides input for absence / presence detection, as well as the presence of interference within the sensing range. These values ​​can then be further evaluated, such as during the breathing decision processing module at 2110, to generate a final breathing estimate of one or more objects monitored by the system.

[0335] 1.2 Absence / Presence Assessment

[0336] like Figure 19 As shown, the 2D analysis at 1910 can also provide processing for absence / presence estimation and an output indication of such estimation. Detecting a person's body movements and their respiratory parameters can be used to evaluate a range of signals to determine whether an object is absent or present within sensor range. "Absence" (and clearly distinguishable from apnea) can be triggered within 30 seconds or less, depending on signal quality detection. In some versions, absence / presence detection can be seen to include a feature extraction method followed by absence detection. Figure 22 An example flowchart of such a process is shown, which can use some of the computational features / metrics / signals of the previously described process at 1910, but other processes can be computed as described herein. Various types of features can be considered when determining human presence (Prs) and human absence (Abs), including one or more of the following, for example:

[0337] 1. Non-zero activity count

[0338] 2. Pearson correlation coefficient between the current 2D signal (respiration and range) and the previous window – limited to respiration and sensing range only.

[0339] 3. The maximum value of the I / Q Ibm quality metric in the current respiratory window.

[0340] 4. Maximum variance of the I / Q respiratory rate channel within a 60-second window.

[0341] 5. Maximum variance of the I / Q respiratory rate channel within a 500-second window.

[0342] Features are extracted from the current buffer and preprocessed by taking percentiles over the buffer length (e.g., after 60 seconds). These features are then combined into a logistic regression model, which outputs the probability of absence occurring in the current epoch. To limit false positives, averaging over a short window (several epochs) can be used to produce the final probability estimate. Periods of motion considered relevant to body movement (human or animal movement) are prioritized for identification. Figure 22 In the example shown, activity counts, IBM, Pearson correlation coefficients, and absence probabilities are compared to appropriate thresholds to assess a person's presence. In this example, any one positive assessment might be sufficient to determine presence, while each test could be assessed as negative to determine absence.

[0343] Using a holistic system view, movements such as limb movements and final rolls tend to disrupt respiratory signals, resulting in higher-frequency components (but can be separable when considering “2D” processing). Figure 23A , 23BThe 23C provides an example of a 2D data segment (top panel), which has two sub-sections ( Figure 23B and Figure 23C The diagram illustrates human activity sensed within an acoustic range by a computing device equipped with a microphone and speaker. Each trace or signal in the graph represents acoustic sensing at a different distance (range) from a mobile device with the applications and modules described herein. In the accompanying figures, breathing signals can be visualized at several distances and moving over time within several ranges (distances). For this example, a person is breathing at approximately 0.3 m from the mobile device, with their torso facing the device. In areas BB and... Figure 23B In the middle, a person gets out of bed (a significant physical movement) and leaves the vicinity of the sensing area (e.g., leaving the room to go to the bathroom). Shortly afterward, regarding the area CC and Figure 23C The person returned to the vicinity of the sensing area or room, such as back to bed (with significant body movement) but slightly away from the mobile device. Now they are facing away from the mobile device (about 0.5m away), and their breathing pattern is visible. At the simplest level of interpreting the attached figure, there is no trace of breathing or significant movement indicating a period of absence from the room.

[0344] This motion detection can be used as input for sleep / wake processing, such as in... Figure 19 In the processing module at position 1912, sleep state can be estimated and / or wake-up or sleep indicators can be generated. Large movements are also more likely to be precursors to changes in range window selection (e.g., the user has just changed position). On the other hand, in the case of an SDB event detected from identifiable respiratory parameters such as apnea or hypopnea, breathing may decrease or cease for a period of time.

[0345] If a velocity signal is required (which may depend on the end use case), an alternative method for handling FMCW chirped signals can be applied. Using this alternative method, the received signal can arrive as a delayed chirp, and an FFT operation can be performed. The signal can then be multiplied by its conjugate to eliminate the signal and preserve its phase shift. A best-fit line to the slope can then be determined using multiple linear regression.

[0346] Graphically, this method generates the FMCW slope when the phase angle is plotted in radians relative to frequency. The slope detection operation needs to be robust to outliers. For a 10ms FMCW chirp in a 18kHz to 20kHz sequence, the effective range is 1.8m with a distance resolution of 7mm. Overlapping FFT operations are used to estimate points on the recovered baseband signal.

[0347] In another approach to processing the FMCW signal directly as a one-dimensional signal, a comb filter can be applied to the synchronization signal (e.g., a block of four chirps). This is to eliminate the direct path from Tx to Rx (directly from the speaker to the microphone component) and static reflections (clutter). An FFT can then be performed on the filtered signal, followed by windowing. A second FFT can then be performed, followed by maximum ratio combining.

[0348] The purpose of this processing is to detect the oscillations of the side lobes and output velocity rather than displacement. One advantage is that it directly estimates the 1D signal (without complex window selection steps) and can optionally be used to estimate the possible range window in the case of a single motion source (e.g., a person) in the sensor field.

[0349] It can also be seen that this type of FMCW algorithm processing technology can also be applied to 2D composite matrices, such as those output by radar sensors that utilize various types of FMCW chirps (e.g., sawtooth waveforms, triangles, etc.).

[0350] 5.1.3 Additional System Considerations – Speakers and / or Microphones

[0351] According to some aspects of the invention, the speaker and microphone can be located on the same device (e.g., on a smartphone, tablet, laptop, etc.) or on different devices having a common or otherwise synchronized clock signal. Furthermore, if two or more components can transmit synchronization information via an audio channel or other means such as the Internet, a synchronized clock signal may not be necessary. In some solutions, the transmitter and receiver can use the same clock, thus eliminating the need for special synchronization techniques. Alternative methods for achieving synchronization include utilizing a Costas ring or PLL (phase-locked loop)—a method that can be “locked” to the transmitted signal—such as a carrier signal.

[0352] Considering any buffering in the audio path, understanding the cumulative impact on the synchronization of transmitted and received samples in order to minimize potentially unwanted low-frequency artifacts can be important. In some cases, it may be necessary to place the microphone and / or speaker near or inside bedding—for example, by using a telephone headset / microphone plug in the device (such as for hands-free calling). One example is that the device is often bundled with Apple and Android phones. During calibration / setup, the system should be able to select appropriate microphone (if the phone has multiple microphones), speaker, amplitude, and frequency settings to suit the system / environment settings.

[0353] 5.1.3.1.1 System calibration / adaptation to component changes and environment

[0354] The goal is to develop a technology that optimally adapts system parameters to the specific telephone (or other mobile device) in use, the environment (e.g., a bedroom), and the user of the system. This means the system learns over time, as the device can be portable (e.g., moved to another living space, bedroom, hotel, hospital, nursing home, etc.), adapts to one or more living objects in the sensing area, and is compatible with various devices. This also implies equalization of the audio channels.

[0355] The system can automatically calibrate channel conditions by learning (or essentially pre-programming by default) device or model-specific characteristics and channel characteristics. Device and model-specific characteristics include the baseline noise characteristics of the speaker and microphone, the ability of mechanical components to vibrate stably at the expected frequency or range, and the amplitude response (i.e., the actual transmitted volume of the target waveform and the signal response of the receiving microphone). For example, in the case of FMCW chirping, the amplitude of the received direct path chirping can be used to estimate the system's sensitivity, and if automatic adjustment of signal amplitude and / or telephone volume is needed to achieve the desired system sensitivity (which may be related to the performance of the speaker and microphone combination, including the sound pressure level (SPL) emitted by the speaker / sensor at a given telephone volume). This allows the system to support a wide variety of smart devices, such as the vast and diverse ecosystem of Android phones. This can also be used to check if the user has adjusted the master volume of the phone and whether automatic readjustment is needed, or if the system needs to be adjusted to new operating conditions. Different operating system (OS) revisions may also produce different characteristics, including audio buffer length, audio path latency, and the likelihood of accidental loss / jitter in the Tx and / or Rx streams.

[0356] Key aspects of the capture system include the quantization levels (available bits) of the ADC and DAC, the signal-to-noise ratio, simultaneous or synchronous clocks for TX and RX, and the room temperature and humidity (if available). For example, it may be detected that the device has an optimal sampling rate of 48kHz, and if a 44.1kHz sample is provided, a suboptimal resampling is performed; in this case, the preferred sampling rate would be set to 48kHz. An estimate is formed from the system's dynamic range, and an appropriate signal composition is selected.

[0357] Other characteristics include the spacing (angle and distance) between the transmitting and receiving components (e.g., the distance between the transmitter and receiver), and the configuration / parameters of any active automatic gain control (AGC) and / or active echo cancellation. Devices (especially telephones using multiple active microphones) may implement signal processing measures that can become cluttered due to the continuous transmission of signals, resulting in unwanted oscillations in the received signal that need correction (or effectively adjusting the device configuration to disable unwanted AGC or echo cancellation as much as possible).

[0358] This system can be pre-programmed with the reflection coefficients of various materials and frequencies (e.g., a reflection coefficient of 18 kHz). If we consider echo cancellation in more detail, for CW (single-tone) signals, the signal is continuous; therefore, unless there is perfect acoustic isolation (unlikely on smartphones), the TX signal is much stronger than the RX signal, and the system can be negatively affected by the echo canceller built into the phone. The CW system may experience strong amplitude modulation due to the activity of the AGC system or the basic echo canceller. As an example of the CW system on a specific phone (the Samsung S4 running the "Lollipop" operating system), the original signal may contain amplitude modulation (AM) of the returning signal. One strategy to address this is to perform very high-frequency AGC on the original sample to smooth out the AM components unrelated to breathing motion.

[0359] Different types of signals, such as the FMCW “chirp,” can appropriately defeat echo cancellers implemented in the device; in fact, (A)FHRG or UWB methods can also be robust to acoustic echo cancellers directed at speech. In the case of FMCW, a chirp is a short-lived, non-stationary signal that can access instantaneous information about the room and motion within it at a single point in time and continue forward; the echo canceller tracks it in a hysteretic manner but can still see the returning signal. However, this behavior depends on the exact implementation of the third-party echo canceller; generally, it is expected (where possible) that any software or hardware (e.g., in a CODEC) echo cancellation is disabled for the duration of physiological sensing Tx / Rx usage.

[0360] Another approach is to use continuous wideband (UWB) signals. This UWB method is highly flexible in situations where there is no good response at a specific frequency. Wideband signals can be confined to inaudible frequency bands or propagate as hiss within the audible band; such signals can be at low amplitudes that will not disturb people or animals, and can optionally be further “shaped” through a window to make them sound pleasant.

[0361] One approach to creating an inaudible UWB sequence is to take an audible probe sequence, such as a maximum-length sequence (MLS—a pseudo-random binary sequence), and modulate it into an inaudible frequency band. It can also be kept at audible frequencies, for example, in situations where it can be masked by existing sounds such as music. Due to the need to repeat the maximum-length sequence, the resulting sound is not pure spectral-flat white noise; in fact, it can produce a sound similar to that produced by commercial audio equipment—while simultaneously slowing down the detection of the breathing rate of one or more people. The impulse can be narrow in time or in the autocorrelation function. “Magic” sequences like MLS are periodic in both time and frequency; MLS, in particular, has excellent autocorrelation properties. The distance to a door can be determined by estimating the impulse response of a room. It is necessary to select subsamples of motion to reconstruct the breathing signal of objects in the room. This can be done using group delay extraction methods; one example is using the centroid (at the first moment) as a surrogate for filtering the impulse response (i.e., the group delay of the segment corresponds to the centroid of the impulse response).

[0362] Channel models can be created automatically (e.g., first used on a smart device or hardware via an "app") or manually prompted by the user (launched), and channel conditions can be monitored periodically or continuously and adjusted as appropriate.

[0363] The room can be questioned using correlated or adaptive filters (such as echo cancellers). Helpfully, bits that cannot be canceled are primarily due to motion within the room. After estimating the room parameters, the original signal can be transformed using a characteristic filter (derived through optimizing an objective function), and then further transformed using the room's impulse response. In the case of pseudo-white noise, the target masking signal can be shaped into the frequency characteristics of ambient noise and improve the sleep experience while also recording bio-motion data. For example, automated frequency assessment of the environment can reveal unwanted standing waves at specific frequencies and adjust the transmitted frequency to avoid such standing waves (i.e., resonant frequencies—the points and inverse nodes of maximum and minimum motion in the air medium / pressure nodes).

[0364] In contrast, flutter echoes can affect sounds above 500Hz and have noticeable reflections from parallel walls (with hard surfaces), dry walls, and glass. Therefore, active noise cancellation can be applied to reduce or eliminate unwanted reflections / effects seen in the environment. Regarding device orientation, system settings will remind the user to face the phone in the optimal position (i.e., to maximize SNR). This may require pointing the phone's speaker towards the chest within a specific distance range. Calibration can also detect and correct for the presence of manufacturer- or third-party phone covers that may alter the system's acoustic characteristics. If the microphone on the phone appears damaged, the user may be prompted to clean the microphone opening (e.g., the phone may have picked up lint, dust, or other material in the microphone opening that can be cleaned).

[0365] According to some aspects of the invention, the continuous wave (CW) method can be applied. Unlike distance gating systems that use time-of-flight, such as FMCW, UWB, or A (FHRG), CW uses a single continuous sinusoidal tone. Unmodulated CW can utilize the Doppler effect (i.e., the return frequency deviates from the transmission frequency) when the object is moving, but may not be able to assess distance. For example, CW can be used in the case of a single person in bed with no other motion sources nearby, and can provide a high signal-to-noise ratio (SNR) in this case. The demodulation schemes outlined in FHRG can be used for the special case of single tone (without the frame itself) to recover the baseband signal.

[0366] Another approach is adaptive CW. This is not range-gated (although it can practically have a limited range, such as detecting the nearest person in the bed due to limitations imposed by Tx power and room reverberation), and it can utilize room patterns. Adaptive CW maintains a continuously transmitted audio tone within the inaudible range, but within the capabilities of the transmitting / receiving devices. By progressively scanning inaudible frequencies, the algorithm iteratively searches for the optimal available breathing signal—both in frequency content and temporal morphology (breathing shape). Intervals of only 10 Hz in the Tx signal can result in entirely different shapes of demodulated breathing waveforms; the cleared morphology is best suited for apnea (central and obstructive) and hypopnea analysis.

[0367] Holography involves wavefront reconstruction. Holography requires a highly stable audio oscillator, where the coherent source is a loudspeaker, and takes advantage of the fact that the room stores energy due to reverberation (i.e., selecting the CW signal to specifically have a strong standing wave, rather than other methods previously discussed that attempt to shift out of the pattern, or actually not use frequency hopping and / or adaptive frequency selection to create a standing wave at all).

[0368] In systems with two or more speakers, the “beam” can be adjusted or manipulated in a specific direction (e.g., to optimally detect a single person in the bed).

[0369] When we consider further Figure 7 In the case of data being stored after a high-pass filter 702 (e.g., an HPF with a 3dB point at 17kHz) for later offline analysis (as opposed to an online processing system that reads and discards the raw audio data), the high-pass filter can act as a privacy filter because the data blocked (removed / filtered out) in the stopband contains the main speech information.

[0370] 5.1.3.1.2 Adaptable to various cooperative or non-cooperative devices and interference.

[0371] The system may include multiple microphones or require collaboration with neighboring systems (e.g., two phones running an application placed on either side of a double bed to independently monitor two people). In other words, multiple transceivers need to be managed in an environment—using channels or other means, such as wireless signals or data transmitted over the Internet—to allow coexistence. Specifically, this means that waveforms can be adapted to selected frequency bands to minimize interference, whether by adjusting the coding sequence or, in the simplest case, detecting a sine wave at approximately 18 kHz and selecting to transmit at approximately 19 kHz. Therefore, the device may include a setup mode that can be activated upon startup to examine the device's environment or nearby sound signals using signal analysis of nearby sounds (e.g., frequency analysis of sounds received by the microphones) and, in response to the analysis, select different frequency ranges for operation, such as non-overlapping frequency groups from the received sound frequencies. In this way, multiple devices can operate in a common vicinity using the sound generation and modulation techniques described herein, where each device generates sound signals within a different frequency range. In some cases, the different frequency ranges may still be within the low ultrasonic frequency range described herein.

[0372] Since an FMCW on a single device can detect multiple people, it is unlikely that an FMCW will be running on more than one nearby device at a time (as opposed to CW, for example); however, if more than one FMCW transmitter is running nearby in order to maximize the available SNR, the frequency bands can be automatically (or by user intervention) adapted to be non-overlapping in frequency (e.g., 18-19kHz and 19.1-20.1kHz, etc.) or non-overlapping in time (the chirps occupy the same frequency bands as each other, but have non-overlapping quiet periods with guard bands to allow reflections from other devices to dissipate).

[0373] When using (A)FHRG with tone pairs, it can be seen that these can be frequency and / or time jitter. Frequency jitter means frequency changes between frames, while time jitter means changes in time of flight (based on changes in pulse duration). Such methods of one or both jitter methods can include reducing the probability of generating room patterns and / or allowing two sonar systems to coexist within each other's "hearing" distance.

[0374] It can also be seen that the FHRG system can be released by defining a pseudo-random sequence of tones / frames—assuming that the conversion is such that sound harmonics are not introduced into the resulting transmitted signal (or by applying appropriate comb filtering to remove / attenuate unwanted subharmonics).

[0375] When TX and RX signals are generated on different hardware, a common clock may not be usable; therefore, cooperation is required. When two or more devices are within "hearing" distance of each other, cooperative signal selection is optimal to avoid interference. This allows the transmitted signal to be optimally returned from the source.

[0376] 5.1.3.1.3 Adapting to User Preferences

[0377] A simple audio scan test allows the user to select the lowest frequency they can no longer hear (e.g., 17.56kHz, 19.3kHz, 21.2kHz, etc.); this (with a small guard band offset) can be used as the starting frequency for generating the signal. A pet setup process can also be included to check if a dog, cat, pet mouse, or similar animal responds to a sample sound; if they do, it may be best to use a low-amplitude (to humans) audible sound signal with modulation information. A "pet setup mode" can be implemented to check, for example, if a dog responds to a specific sound, and the user can record this fact, allowing the system to examine different sounds to find a signal that will not cause discomfort. Similarly, the system can be configured to use an alternative signal type / band if user discomfort is noticed. Including the white noise characteristic of an active signal TX is very useful when pets / children cannot tolerate a particular waveform and / or require a calm masking noise ("white noise" / hiss). Therefore, the mode can cycle through one or more test sound signals and prompt the user with input regarding whether the test sound signal (which may be inaudible to humans) is problematic, and select the frequency to use based on the input.

[0378] 5.1.3.1.4 Automatically pause Tx playback when the user interacts with the device.

[0379] Because the system may be playing high-amplitude, inaudible signals, it's expected that mobile devices, such as smart devices, will be muted (Tx paused) if a user is interacting with the device. Specifically, if the device is near the user's ear (e.g., making or receiving a call), it must be muted. Of course, once the signal is unmuted, the system needs to resynchronize. The system may also drain the battery, so the following methods can be used. If the system is running, it should only be prioritized when the device is charging; therefore, if the smart device is not connected to a wired or wireless charger, the system can be designed to pause (or not start). Regarding pausing Tx (and processing) while the device is in use, input can be taken from one or more interactions between the user and the phone—pressing a button, touching the screen, moving the phone (detected by a built-in accelerometer (if present) and / or gyroscope, and / or infrared proximity sensor), or location changes detected by GPS or assisted GPS, or an incoming call. In the case of notifications (e.g., text messages or other messages), if the device is not in "mute" mode, the system can temporarily pause Tx to anticipate the user picking up the device for a period of time.

[0380] Upon startup, the system can wait a period before activating Tx to check if the phone is already placed on the surface and if the user is no longer interacting with it. If a distance estimation method such as FMCW is used, the system can reduce the Tx power level or pause (mute) Tx if breathing motion is detected very close to the device, based on the minimum possible acoustic output power being sufficient to meet the required SNR level. Alternatively, the Tx volume can be actively and smoothly reduced using gestures recovered from the demodulated baseband signal when the user reaches for the device, before muting it when the user actually interacts with the device, or before the user smoothly increases the volume to the operational level by withdrawing their hand from the device.

[0381] 5.1.3.1.5 Data Fusion

[0382] For systems that collect (receive) audio signals with active transmit components, it is also desirable to process the full-band signal to reject or utilize other modes. For example, this can be used to detect speech, background noise, or other modes that may overwhelm the TX signal. For respiratory analysis, data fusion with extracted respiratory audio waveform characteristics is highly desirable; direct detection of respiratory sounds can be combined with the demodulated signal. Furthermore, aspects of dangerous sleep or respiratory conditions with audio components such as coughing, wheezing, and snoring can be extracted—particularly for bedroom applications (typically quiet environments). Respiratory sounds can originate from the mouth or nose and include snoring, wheezing, panting, whistling, and so on.

[0383] The complete audio signal can also be used to estimate motion (such as large movements like a user rolling around in bed) and distinguish it from other background, non-physiologically generated noise. Motion estimated by sonar (from the processed baseband signal) and full-band passive audio signal analysis can be combined, often prioritizing sonar motion (and the estimated activity derived from that motion) through non-range-specific passive acoustic analysis in FMWC, (A)FHRG ToF, etc., due to the advantages of range gating (range detection). The duration and intensity of detected noise can also provide insight into unwanted desynchronization—such as when a fan is placed close or when the heating or cooling system is excessively noisy. In sonar baseband signal analysis, acoustic fan noise can also be considered as an increased 1 / f signal. In cases where actual user voices can be detected on a noise floor but the sonar Rx signal quality is very poor, the system can revert to a processing mode where physiological sounds are mapped to activity and directly drive a sleep / wake detector with reduced accuracy, purely based on these activity indices. The system can also provide feedback to the user on optimal fan placement and / or adapt its operating frequency to a less congested band. Other noise sources, such as fluorescent tubes / bulbs and LED ballasts, can be detected, and the system is adapted to preferred sonar operating frequencies. Therefore, the device can also process audio signals received via the microphone using conventional sound processing methods (in addition to the demodulation processing techniques described herein) to evaluate any one or more of ambient sounds, speech, and breathing sounds to detect user motion and related characteristics.

[0384] Voice detection can also be used as a privacy feature to proactively discard any temporary data that may contain private information. Processing can be performed to eliminate potential confounding factors—tickling analog clocks, streaming media on televisions and tablets, fans, air conditioning units, forced heating systems, traffic / street noise, and so on. Therefore, bio-motion information can be extracted from the audible spectrum and combined with data extracted from the demodulation scheme to improve overall accuracy.

[0385] According to some aspects of the invention, a light sensor on a device such as a smartphone can provide a separate input to the system indicating whether the person is trying to sleep or is watching TV, reading, using a tablet, etc. Using a temperature sensor on the phone and / or available humidity sensing (or location-based weather data) can be used to enhance channel estimation / propagation of transmitted signals. Interaction with the phone itself can provide additional information about the user's alertness level and / or fatigue state.

[0386] Understanding the motion of a sensing device through internal motion sensors, such as MEMS accelerometers, can be used to disable acoustic sensing when the phone is in motion. The fusion of sensor data and accelerometer data can be used to enhance motion detection.

[0387] Note that while the systems and methods described herein are described as being implemented by mobile devices, in other aspects of the invention, the systems and methods can be implemented by fixed devices. For example, the systems and methods described herein can be implemented by bedside consumer monitoring devices or medical devices, such as flow generators (e.g., CPAP machines) for treating sleep apnea or other respiratory conditions.

[0388] 5.1.3.1.6 Cardiac Information

[0389] As previously mentioned, in addition to respiratory information, various technical versions of sound generation and reflection analysis, as described earlier, can be implemented based on the generated motion-related signals for other periodic information (such as cardiac information detection or heart rate). Figure 21 In this example, cardiac decision processing can be an optional addition module. Cardiac peak detection based on cardiac impact mapping can be applied during the FFT-based 2D processing phase (as is performed for respiratory detection), but higher frequency bands can be viewed.

[0390] Alternatively, wavelets (as a replacement or supplement to FFT) – such as discrete wavelet transforms that process each of the I and then Q signals – or discrete complex wavelet transforms that can be used in the 2D processing stage to separate large-amplitude motion, respiratory, and cardiac signals.

[0391] Time-frequency processing, such as wavelet-based methods (e.g., Discrete Continuous Wavelet Transform—DCWT—using, for example, Daubechies wavelets), can perform detrending and direct extraction of body motion, respiratory, and cardiac signals. Cardiac activity is reflected in higher-frequency signals and can be accessed by filtering using a bandpass filter with a passband range of 0.7 to 4 Hz (48 to 240 heartbeats per minute). Activity caused by large-amplitude movements is typically in the range of 4 to 10 Hz. It should be noted that overlap may exist within these ranges. Strong (clean) respiratory traces can produce strong harmonics, and these need to be tracked to avoid confusion. For example, signal analysis in some versions of this technique may include any of the methods described in International Patent Application Publication No. WO2014 / 047310, which is incorporated herein by reference in its entirety, including, for example, wavelet denoising methods for respiratory signals.

[0392] 5.1.3.1.7 System Instance

[0393] Typically, the technical version of this application can be implemented by one or more processors configured to monitor related methods, such as the algorithms or methods of the modules described in more detail herein. Therefore, the technology can be implemented using an integrated chip, one or more memories and / or other control instructions, data, or information storage media. For example, programming instructions containing any of the methods described herein can be encoded on an integrated chip in the memory of a suitable device. These instructions can also or alternatively be loaded as software or firmware using a suitable data storage medium. Therefore, the technology can include a processor-readable medium or a computer-readable data storage medium storing processor-executable instructions that, when executed by one or more processors, cause the processors to perform any of the methods or aspects of the methods described herein. In some cases, a server or other networked computing device may include or be otherwise configured to access such a processor-readable data storage medium. The server can be configured to receive requests for downloading processor-executable instructions from a processor-readable data storage medium to a processing device over a network. Therefore, the technology can relate to a method of a server accessing a processor-readable data storage medium. The server receives requests to download processor-executable instructions from a processor-readable data storage medium, such as for downloading instructions over a network to a processing device such as a computing device or a portable computing device (e.g., a smartphone). The server can then, in response to the request, send processor-executable instructions from a computer-readable data storage medium to the device. The device can then execute the processor-executable instructions, for example, if the processor-executable instructions are stored on another processor-readable data storage medium of the device.

[0394] 5.1.3.1.8 Other portable or electronic processing devices

[0395] As previously mentioned, the sound sensing method described herein can be implemented by one or more processors or computing devices in an electronic processing device (such as smartphones, laptops, portable / mobile devices, sports phones, tablets, etc.). These devices can generally be understood as portable or mobile. However, other similar electronic processing devices can also be implemented using the techniques described herein.

[0396] For example, many homes and vehicles contain electronic processing devices capable of emitting and recording sound, such as in the low-frequency ultrasonic range just above the human hearing threshold—e.g., smart speakers, active soundbars, smart devices, and other devices that support voice and other virtual assistants. Smart speakers or similar devices typically include communication components via wired or wireless devices such as Bluetooth, Wi-Fi, ZigBee, mesh, peer-to-peer networks, etc., for communication with other home devices, such as for home automation and / or networks such as the Internet. Unlike standard speakers designed to simply emit acoustic signals, smart speakers typically include one or more processors, one or more speakers, and one or more microphones. Microphones can be used to interact with intelligent assistants (artificial intelligence (AI) systems) to provide personalized voice control. Some examples are Google Home, Apple HomePod, Amazon Echo, and devices that use the phrases "OK Google," "Hey Siri," and "Alexa" for voice activation. These devices can be portable and are often used in specific locations. The sensors they connect to can be considered part of the Internet of Things (IoT). Other devices can also be used, such as active soundbars (i.e., those that include microphones), smart TVs (which can typically be fixed devices), and mobile smart devices.

[0397] Such devices and systems can be adapted to perform physiological sensing using the low-frequency ultrasound technology described herein.

[0398] For devices with multiple transducers, beamforming can be implemented—that is, signal processing is used to provide directional or spatial selectivity for signals sent to or received from a sensor array (e.g., a loudspeaker). This is typically a “far-field” problem, where the wavefront is relatively flat for low-frequency ultrasound (as opposed to the “near-field” problem in medical imaging). For a pure CW system, the audio waves emanate from the loudspeaker, resulting in maximum and minimum regions. However, if multiple sensors are available, this radiation pattern in our favor can be controlled—a method called beamforming. Multiple microphones can also be used on the receiving side. This allows acoustic sensing to be preferably directed in one direction (e.g., guiding the emitted sound and / or received sound waves) and sweep across an area. In the case of a user in bed, sensing can be directed toward an object—or toward multiple objects (in this case, for example, two people in bed).

[0399] As another example, the techniques described herein can be implemented in wearable devices, such as non-invasive devices (e.g., smartwatches) or even invasive devices (e.g., implanted chips or implantable devices). These portable devices can also be configured with this technology.

[0400] 5.2 Other Notes

[0401] This patent document contains a portion of copyrighted material. Because it appears in the patent office's patent documents or records, the copyright holder does not object to any person making a copy of this patent document or the patent disclosure, but otherwise retains all copyright rights.

[0402] Unless explicitly stated in the context and a numerical range is provided, it should be understood that every intermediate value between the upper and lower limits of the range, up to one-tenth of the lower limit unit, and any other such value or intermediate value within the range are broadly included within this technique. The upper and lower limits of these intermediate ranges may be independently included within the intermediate range and within the scope of this technique, but are subject to any explicitly excluded boundaries within the range. When the range includes one or both of these boundaries, the range excluding one or both of those included boundaries is also included within this technique.

[0403] Furthermore, where one or more values ​​of this technology are implemented as part of this technology, it should be understood that such values ​​may be approximate unless otherwise stated, and such values ​​may be used to the extent permitted or required by the practical implementation of the technology for any appropriate valid digits.

[0404] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this art pertains. Although any methods and materials similar to or equivalent to those described herein may be used in the practice or testing of this art, a limited number of exemplary methods and materials are described herein.

[0405] When a particular material is identified for use in a component, a readily available alternative material with similar properties is used as its substitute. Furthermore, unless otherwise stated, any and all components described herein are to be understood as being capable of being manufactured and therefore can be manufactured together or separately.

[0406] It must be noted that, unless the context clearly specifies otherwise, the singular forms “a,” “an,” and “the” as used herein and in the appended claims include their plural equivalents.

[0407] All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials that are the subject of those publications. The publications discussed herein provide only disclosures prior to the filing date of this application. None of this document should be construed as an admission by prior invention that the present technology is not claimed prior to such publications. Furthermore, the publication dates provided may differ from the actual publication dates, and independent verification may be required.

[0408] The terms “comprising” and “including” should be interpreted as referring to an element, component or step in a non-exclusive manner, indicating that the referenced element, component or step may be present or used or combined with other elements, components or steps not expressly referenced.

[0409] The main headings used in the detailed description are included for the reader's convenience only and should not be used to limit the subject matter of the invention as found throughout the disclosure or claims. These headings should not be used to interpret the scope or limitation of the claims.

[0410] Although the techniques described herein have been illustrated with reference to specific examples, it should be understood that these examples merely illustrate the principles and applications of the techniques. In some cases, terms and symbols may imply specific details not required for practicing the techniques. For example, although the terms "first" and "second" may be used, unless otherwise specified, they are not intended to indicate any order, but can be used to distinguish different elements. Furthermore, although the process steps in a method may be described or illustrated sequentially, they do not need to be ordered. Those skilled in the art will recognize that such order can be modified and / or aspects can be performed simultaneously or even concurrently.

[0411] Therefore, it should be understood that many modifications can be made to the illustrative examples, and other arrangements can be designed without departing from the spirit and scope of the technology.

Claims

1. A processor-readable medium having processor-executable instructions stored thereon, which, when executed by a processor, cause the processor to detect physiological movements of a user, the processor-executable instructions comprising: Instructions for controlling the generation of sound signals, including those near the user, by a speaker coupled to an electronic processing device, wherein, The processor-executable instructions also include instructions for initiating the generation of the sound signal based on the detection that no user interaction with the electronic processing device is detected; Instructions for controlling the sensing of sound signals reflected from the user via a microphone coupled to the electronic processing device; Instructions used to process the sensed sound signals; and Instructions used to detect respiratory signals from processed sound signals. The sound signal is a continuous wave (CW) signal using only a single continuous sinusoidal tone during the detection of the respiratory signal.

2. The processor-readable medium according to claim 1, wherein, The sound signal is within the range of inaudible sounds.

3. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor-readable medium contains instructions for controlling an adaptive continuous wave scheme.

4. The processor-readable medium according to claim 3, wherein, The instructions for controlling the adaptive continuous wave scheme include instructions for controlling the scanning of inaudible frequencies to iteratively search for the frequency of the respiratory signal.

5. The processor-readable medium according to claim 4, wherein, The scan takes into account both frequency content and temporal morphology to search for the best available respiratory signal.

6. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor executable instructions include signal processing instructions for generating and sensing audio using continuous wave zero-difference technology.

7. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor-executable instructions include instructions for performing high-frequency automatic gain control on the sensed signal samples to smooth amplitude modulation independent of respiratory movements.

8. The processor-readable medium according to any one of claims 1 to 2, wherein, Instructions for processing the sensed sound signal reflected from the user include a demodulator to generate one or more baseband motion signals that include the breathing signal.

9. The processor-readable medium according to claim 8, wherein, The demodulator generates multiple baseband motion signals, which include orthogonal baseband motion signals.

10. The processor-readable medium of claim 9, further comprising instructions for processing the plurality of baseband motion signals, the instructions for processing the plurality of baseband motion signals comprising an intermediate frequency processing module and an optimization processing module to generate a combined baseband motion signal from the plurality of baseband motion signals, the combined baseband motion signal comprising a breathing signal.

11. The processor-readable medium of claim 10, wherein, The instructions for detecting the respiratory signal include determining the respiratory rate based on the combined baseband motion signal.

12. The processor-readable medium according to any one of claims 1 to 2, wherein, The instructions for controlling the sensing of sound signals reflected from the user include storing sound data sampled from the microphone.

13. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor-executable instructions further include: Instructions for calibrating sound-based body motion detection by evaluating one or more features of the electronic processing device; and Instructions for generating the sound signal based on the evaluation.

14. The processor-readable medium according to claim 13, wherein, Instructions for calibrating sound-based detection determine at least one hardware, environment, or user-specific characteristic.

15. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor-executable instructions further include: Instructions for operating the pet's settings mode, wherein the frequency at which the sound signal is generated is selected based on user input, and one or more test sound signals are generated.

16. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor executable instructions also include the following instructions: The device is configured to stop generating the sound signal based on detected user interaction with the electronic processing device, the detected user interaction including any one or more of the following: detecting movement of the electronic processing device and the accelerometer, detecting pressing a button, detecting touching the screen, and detecting an incoming call.

17. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor executable instructions also include the following instructions: This is used to detect large-amplitude body movements based on the processing of sensed sound signals reflected from the user.

18. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor executable instructions also include the following instructions: Used to process audio signals sensed through the microphone to evaluate any one or more of ambient sounds, speech, and breathing sounds to detect user movement.

19. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor executable instructions also include the following instructions: The respiratory signals are used to determine any one or more of the following: (a) a sleep state indicating sleep; (b) a sleep state indicating wakefulness; (c) a sleep stage indicating deep sleep; (d) a sleep stage indicating light sleep; (e) a sleep stage indicating REM sleep.

20. The processor-readable medium according to any one of claims 1 to 2, wherein, The processor executable instructions also include the following instructions: The operation setting mode is used to detect sound frequencies near the electronic processing device and select the frequency range of a sound signal that is different from the detected sound frequency.

21. The processor-readable medium according to claim 20, wherein, The instructions for operating the setting mode select a frequency range that does not overlap with the detected sound frequency.

22. A server capable of accessing a processor-readable medium according to any one of claims 1 to 2, wherein, The server is configured to receive requests for downloading processor-executable instructions from the processor-readable medium to an electronic processing device via a network.

23. A mobile electronic device comprising: one or more processors; a speaker coupled to the one or more processors; a microphone coupled to the one or more processors; and a processor-readable medium according to any one of claims 1 to 2.

24. A method of a server capable of accessing a processor-readable medium according to any one of claims 1 to 2, the method comprising: receiving at the server a request to download processor-executable instructions of the processor-readable medium to an electronic processing device via a network; and sending the processor-executable instructions to the electronic processing device in response to the request.

25. A method for a processor to detect body motion using a mobile electronic device, comprising: Using a processor to access the processor-readable medium of any one of claims 1 to 2 The processor-executable instructions of the processor-readable medium are executed in the processor.

26. A method for a processor to detect physiological body movements using a mobile electronic device, comprising: Control generates sound signals, including those from the user's vicinity, via a speaker coupled to the mobile electronic device. The generation of the sound signal is initiated based on the detection that there is no interaction between the user and the mobile electronic device; Control is achieved by sensing sound signals reflected from the user through a microphone coupled to the mobile electronic device; Process the sensed reflected sound signals; and Detecting physiological body movement signals from processed reflected sound signals. The sound signal is a continuous wave (CW) signal using only a single continuous sinusoidal tone during the detection of the physiological body movement signal.

27. A method for detecting motion and breathing using a mobile electronic device, comprising: Sound signals are transmitted to the user via a speaker on the mobile electronic device, wherein... The transmission of the sound signal begins based on the detection that there is no interaction between the user and the mobile electronic device; The sound signal reflected from the user is sensed by the microphone on the mobile electronic device; and Respiratory and movement signals are detected from the reflected sound signals. The sound signal is a continuous wave (CW) signal using only a single continuous sinusoidal tone during the detection of the breathing and motion signals.

28. The method according to claim 27, wherein, The sound signal is an inaudible sound signal.

29. The method of claim 27, further comprising: When sensing the reflected sound signal, the reflected sound signal is demodulated, wherein demodulation includes: Perform a filtering operation on the reflected sound signal; and The filtered reflected sound signal is synchronized with the transmitted sound signal in terms of timing.

30. The method according to any one of claims 27 to 29, wherein, Generating the sound signal includes: Perform calibration functions to evaluate one or more characteristics of the mobile electronic device; and A sound signal is generated based on the calibration function.

31. The method according to claim 30, wherein, The calibration function is configured to determine at least one hardware, environment, or user-specific characteristic.

32. A processor-readable medium having processor-executable instructions stored thereon, which, when executed by a processor, cause the processor to detect physiological movements of a user, the processor-executable instructions comprising: Instructions for controlling the generation of sound signals, including those near the user, by a speaker coupled to an electronic processing device; Instructions for controlling the sensing of sound signals reflected from the user via a microphone coupled to the electronic processing device; Instructions used to process the sensed sound signals; and Instructions used to detect respiratory signals from processed sound signals. The sound signal comprises pitch pairs forming pulses and a sequence of frames, each frame containing a series of pitch pairs, each pitch pair being associated with a corresponding time slot within the frame.

33. The processor-readable medium according to claim 32, wherein, The sound signal is within the range of inaudible sounds.

34. The processor-readable medium according to claim 32, wherein, The pitch pair includes a first frequency and a second frequency, wherein the first frequency is different from the second frequency.

35. The processor-readable medium according to claim 34, wherein, The first frequency and the second frequency are orthogonal to each other.

36. The processor-readable medium according to claim 32, wherein, The series of tone pairs in the frame includes a first tone pair and a second tone pair, wherein the frequency of the first tone pair is different from the frequency of the second tone pair.

37. The processor-readable medium according to claim 32, wherein, The pitch of the time slot of the frame has zero amplitude at the beginning and end of the time slot, and a ramp amplitude of round-trip peak amplitude between the beginning and the end.

38. The processor-readable medium according to claim 32, wherein, The time width of the frame changes.

39. The processor-readable medium according to claim 38, wherein, The time width is the width of the time slot of the frame.

40. The processor-readable medium according to claim 38, wherein, The time width is the width of the frame.

41. The processor-readable medium according to any one of claims 32 to 33, wherein, The frame sequence forms different frequency patterns relative to different time slots of the frame.

42. The processor-readable medium according to claim 41, wherein, The different frequency patterns are repeated in multiple frames.

43. The processor-readable medium according to claim 41, wherein, The different frequency patterns change for different time slot frames in multiple time slot frames.

44. The processor-readable medium according to any one of claims 32 to 33, wherein, The instructions for controlling the generation of the sound signal include a pitch-to-frame modulator.

45. The processor-readable medium according to any one of claims 32 to 33, wherein, The instructions for controlling the sensing of sound signals reflected from the user include a frame buffer.

46. ​​The processor-readable medium according to claim 32, wherein, The duration of the corresponding time slot of the frame is equal to 1 divided by the difference between the frequencies of the tone pairs.

Citation Information

Patent Citations

  • Detection of periodic breathing

    US20160287122A1

  • System and method for determining sleep stage

    WO2014047310A1

  • Method and system for sleep management

    WO2015006364A2

  • Control for pressure of a patient interface

    WO2015061848A1

  • Respiratory therapy apparatus and method

    WO2016145483A1