Apparatus, system and method for motion sensing
Processor-activated audio equipment with integrated sensors in devices like smartphones effectively monitors biological movements, addressing the challenge of specialized hardware requirements by deriving physiological parameters through acoustic sensing, enhancing sleep monitoring and environmental control.
Patent Information
- Application Number
- JP2024053605
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-12-22
- Filing Date
- 2024-03-28
- Publication Date
- 2026-02-26
- Estimated Expiration
- 2038-12-21
AI Technical Summary
Existing technologies face challenges in efficiently and effectively monitoring biological movements, such as breathing and cardiac activity, using audio equipment, particularly in portable devices like smartphones, due to the need for specialized hardware and antennas, which are not widely available in developing countries.
The use of processor-activated audio equipment, including integrated or externally connectable speakers and microphones, to detect physiological movements like breathing, cardiac, and limb movements, utilizing acoustic sensing and processor-executable instructions to derive physiological parameters, including respiratory and cardiac characteristics, through methods like dual-tone frequency modulation and ultra-wideband audio signals.
Enables efficient and effective monitoring of biological movements, including sleep-related characteristics, using common audio devices, providing insights into sleep stages and physiological parameters, and enabling interactive responses and environmental control.
Smart Images

Figure 0007821217000006 
Figure 0007821217000007 
Figure 0007821217000008
Abstract
Description
[Technical Field]
[0001] 1 Cross-reference to related applications This application is a joint venture of U.S. Provisional Patent Application No. 62 / 610,013 (filed December 22, 2017). The benefit of the present application is claimed in its entirety. It will be called the department.
[0002] 2. Technical Background 2.1 Technology field This technology relates to the detection of biological motion (bio-motion) of living organisms using audio equipment. More specifically, the present technology utilizes acoustic sensing with audio equipment to detect physiological Movement (e.g., breathing, cardiac, and / or other low-periodic biological It involves detecting physiological characteristics such as the body's movements. [Background technology]
[0003] 2.2 Description of Related Art For example, monitoring or observing a person's breathing and body movements (including limbs) while sleeping. Such a monitor can be useful in many situations. For example, Useful in monitoring and / or diagnosing respiratory disorder conditions (e.g., sleep apnea) Traditionally, barriers to adoption of active radiopositioning or related applications have included special However, some hardware circuits and antennas may be required.
[0004] Smartphones and other portable, discreet processing or electronic communication devices Even in developing countries where landline communication lines are not available, communication services have become a ubiquitous part of everyday life. For example, many homes have audio equipment that can emit and record sound. devices are used (e.g., smart speakers, active sound bars, smartphones, etc.) devices, smart TVs and other devices). Such devices may The assistant can be used to support voice commands. It processes incoming verbal commands and responds with audio output. Summary of the Invention [Problem to be solved by the invention]
[0005] Monitoring biological (i.e., physiological) movement What is desired is a method for doing this in an efficient and effective manner. Achieving this will pose considerable technical challenges. [Means for solving the problem]
[0006] 3. Brief description of the technology The present technology provides systems, methods, and apparatus for detecting movement of a subject, for example, while the subject is sleeping. Based on such motion detection (e.g., breathing motion, subject motion) , sleep-related characteristics, respiratory characteristics, cardiac characteristics, sleep states, etc. may be detected. Processor-activated audio equipment (e.g., interactive audio devices) or relates to other processing devices (e.g., smartphones, tablets, smart speakers) In applications where motion detection is required, the processing device may include sensors (e.g., integrated and / or or externally connectable speaker(s) and microphone(s) The term "device" as used herein is broadly defined as a device that provides audio functionality to a device. It is used meaningfully and may include centralized or decentralized systems. The system includes one or more speakers, one or more microphones, and one or more processors. However, in some versions this may include, for example, components and Any one or more of the functions of the sensing method / detection method described in the specification may be implemented as a unit. For example, a processing device and its associated components may be integrated into a single device to provide the desired quality of the device. The method and sensing component may be implemented in, for example, a handheld system (e.g., integrated into a portable system such as a sensing-configurable smartphone or other such system It can be integrated into the housing.
[0007] Some versions of the technology may include a processor-readable medium. The medium stores instructions that can be executed by a processor. When executed by the method, the processor detects physiological movements of the user, such as by detecting physiological movements of the user. Physiological parameters are detected. Physiological movements include breathing, cardiac, limb movements, and gestures. The physiological parameters may include any one or more of the following: It may also include one or more characteristics derived from the measured physiological movement (e.g., breathing amplitude, phase, etc.). relative respiratory amplitude, respiration rate, respiratory rate variability, cardiac amplitude, relative cardiac amplitude of heart rate, heart rate variability), as well as other physiological parameters. Other characteristics (e.g., (a) presence state (presence / absence), (b) sleep state (e.g., wakefulness or (c) sleep stage (e.g., N-REM1 (light sleep substage of non-REM)) N-REM2 (non-REM light sleep substage 2), N-REM3 (non-REM light sleep substage 3) Deep sleep (also known as slow wave sleep (SWS)), REM sleep, etc.), or (d) fatigue and / or (e) other sleep-related parameters, such as sleepiness). Executable by a processor. The available instructions may include the user via a speaker connected to the interactive audio device. The processor-executable instructions may include instructions for controlling generation of a nearby audio signal. ,reflected from the user via a microphone connected to the interactive audio processing device. The processor-executable instructions may include instructions for controlling sensing of a received audio signal. The processor-executable instructions may include instructions for processing the sensed audio signal. Audible speech (ve) sensed via a microphone connected to a talking audio device The instructions executable by the processor may include instructions to perform an evaluation of the communication. The instructions may include instructions to derive a physiological movement signal using the sound signal and the reflected sound signal. The processor-executable instructions are derived in response to the sensed audible verbal communication. The method may include instructions for generating an output based on an evaluation of the physiological movement signal.
[0008] Some versions of the technology include a processor-readable medium. The audio data is stored on the memory card and can be executed by a processor. When executed by a processor in the device, the processor may The processor-executable instructions include: and instructions for controlling generation of an audio signal in the vicinity of the interactive audio device via a speaker. The processor-executable instructions may include: may include instructions for controlling sensing of reflected sound signals from nearby areas via a microphone. The processor-executable instructions include: and instructions for deriving a physiological movement signal using a signal indicative of at least a portion of the audio signal. The processor-executable instructions may include: Both may include instructions to generate an output based on some of the evaluation.
[0009] In some versions, at least a portion of the generated audio signal (e.g., The generated audio signal (part of which is used for sensing purposes) may be substantially in the inaudible range. At least a portion of the audio signal may be a low frequency ultrasonic acoustic signal. May include internally generated oscillator signals or direct path measurement signals. The instructions for deriving the signal may include: (a) converting the physiological movement signal into at least one of the generated audio signal; (b) together with at least a portion of the sensed reflected sound signal; or (c) together with at least a portion of the sensed reflected sound signal. a portion of the sound signal and an associated signal that can be associated with at least a portion of the generated audio signal; Optionally, the related signal may be derived using an internally generated oscillator. The physiological motion signal may be a data signal or a direct path measurement signal. The oscillator signal may be configured to multiply the oscillator signal by a portion of the sensed reflected sound signal. The physiological motion signals emitted may be one or more of the following: respiratory motion, whole body motion, or cardiac motion. or detection of any one or more of such movements. audible verbal communication sensed via a microphone connected to the audio device; The output may further include processor-executable instructions to evaluate the The generating instructions are configured to cause an output to be generated in response to the sensed audible verbal communication. The processor-executable instructions for deriving a physiological movement signal include: The method includes demodulating a portion of the detected reflected sound signal using at least a portion of the audio signal. Demodulation may involve multiplying a portion of the audio signal with a portion of the sensed reflected sound signal.
[0010] The processor executable instructions for controlling generation of the audio signal include dual tone frequencies A dual-tone frequency modulated continuous wave signal can be generated. a first sawtooth frequency change overlapped with a second sawtooth frequency change; This medium generates ultra-wideband (UWB) audio signals as audible white noise. The processor-readable medium may include instructions executable by a processor. B. The medium may include instructions for detecting user movement using audio signals. Probing from the speaker during the setup process to calibrate the distance measurement The medium may further include instructions executable by the processor to generate the audio sequence. , estimation of the distance between the microphone and another microphone of the interactive audio device The setup process involves time synchronization from one or more speakers, including the and a processor-executable instruction for generating a calibration acoustic signal. The medium may activate a beamforming process to further localize the detected area. The method may include instructions executable by a processor to evaluate the derived physiological movement signals. The output generated based on the value may include monitored user sleep information.
[0011] In some versions, the evaluation of the derived physiological movement signal portion is The method may include detecting one or more physiological parameters. is the breathing rate, the relative amplitude of breathing, the cardiac amplitude, the relative cardiac amplitude, The monitored user may include any one or more of heart rate and heart rate variability. Sleep information includes sleep score, sleep stages, and time in sleep stages. The generated output may include an interactive query and response presentation. The presentation of the generated interactive queries and responses may be performed via a speaker. The generated interactive queries and responses are presented to improve the user's sleep monitoring. Based on an evaluation of portions of the derived physiological movement signal, The output generated by the It may further be based on performing a search of resources. The search may be based on historical user data or recorded derived physiological movement signals. The output generated based at least in part on the evaluation of The medium may include a control signal for controlling an automated device or system. The network may further include processor control instructions for transmitting the network control instructions to the network system. may be the Internet.
[0012] Optionally, the media is adapted to modify settings of the interactive audio device based on physiologically derived generating a control signal for performing the action based at least in part on an evaluation of the signal of the desired movement; The control signals may further include control instructions for changing settings of the interactive audio device. The signal may be related to the user's distance from the interactive audio device, the user's state, or the user's location. The processor-executable instructions may include changing the volume based on the detection of the plurality of users. configured to evaluate movement characteristics at different acoustic sensing ranges for monitoring sleep characteristics of a person. The instructions controlling the generation of the sound signals may be adapted to detect different users at different frequencies. The generation of the sound signals may be controlled to generate simultaneous sensing signals having different sensing frequencies. The instructions include interleaved acoustic sensing for sensing different users at different times. The medium may control the generation of the physiological motion signal. and further comprising instructions executable by a processor to detect the presence or absence of a user based on the It can be seen.
[0013] In some versions, the medium comprises at least one of a derived physiological movement signal. and a processor-executable method for performing biometric recognition of a user based in part on the received biometric information. The medium may further include instructions for generating a communication over a network. a biometric assessment determined from an analysis of at least a portion of the physical movement signals; and / or (b) determined from an analysis of at least a portion of the derived physiological movement signals. The medium may further include processor-executable instructions for generating based on the presence detection. The method may further include instructions executable by the processor. These instructions may include instructions for detecting a user's movement. acoustically detecting the presence of a user, audibly calling out to the user, audibly recognizing the user Authenticating and authorizing users or alarm messages about users The interactive audio device is configured to audibly communicate to the user The processor-executable instructions for effectively authenticating a user sensed by a microphone include: The system may be configured to compare the sound waves of the user's words with sound waves of pre-recorded words. Optionally, the processor executable instructions may include: The system may be configured to use the authentication method to authenticate a user.
[0014] In some versions, the medium comprises at least one of a derived physiological movement signal. and causing an interactive audio device to detect movement gestures based in part on the audio signal. The medium may include processor-executable instructions for detecting a detected movement gesture. to generate control signals, send notifications, or (a) use automated devices and / or (b) a processor that controls changes to the operation of an interactive audio device. The audio device may include instructions executable by the audio processor to control changes to the operation of the audio device. The control signal for the interactive voice assistant operation of the interactive audio device is Activating a microphone sense may be performed to start the Interactive voice assistant processes that are not connected to the device may be linked. Initiating the process may further be based on the detection of linguistic keywords by microphone sensing. This medium is at least partly a physiological movement signal derived from an interactive audio device. and detecting a user's breathing movement based in part on the A processor-executable instruction to generate an output queue that triggers It may include.
[0015] In some versions, the media is a separate object controlled by a separate processing device. Communication from the interactive audio device is recorded by the microphone of the interactive audio device. and a process for causing an interactive audio device to receive the inaudible sound waves sensed by the The interactive audio device may include instructions executable by the device to detect a sound signal in the vicinity of the interactive audio device. The instructions for controlling the generation include instructions for generating at least a portion of the sound signal based on the communication. The adjusted parameters allow the interactive audio device to be differentiated from the Interference between the device and the user is reduced. The evaluation of at least some of the signals may further include detecting sleep onset or wake onset. The output based on the service control signal may include a service control signal. The controls may include one or more of: a music player control, a volume control, a thermostat control, and a window covering control. The system includes a prompt for the user to collect feedback and physiological responses derived in response. and providing advice based on at least one of the following: a portion of the biological motion signals and feedback. The medium may further include instructions executable by a processor to generate the environmental data. The advice may include instructions executable by a processor to determine the determined environment. The medium may further be based on the data. instructions and at least a portion of the determined environmental data and derived physiological movement signals. a processor executable to generate a control signal for an environmental control system based on the The medium may include instructions executable by a processor to provide a sleep enhancement service. Sleep enhancement services may include any of the following: (a) Advice in response to detected sleep state and / or collected user feedback and (b) setting sleep environment conditions and / or sleep-related activities. Generates control signals for controlling environmental equipment to provide user advice messages To do this.
[0016] In some versions, the medium comprises at least one of a derived physiological movement signal. and detecting a gesture based in part on an interactive audio device. By the processor to start microphone sensing for initiation of voice assistant operation. The medium may include instructions executable by the user to transmit at least one of the derived physiological movement signals. and processor-executable instructions for initiating and monitoring a nap session based on the portion. The medium may include a packet of sound signals generated to track the user's movements through the vicinity. a processor that varies the detection range by varying at least some of the parameters; The medium may include instructions executable by the device. The medium may include instructions for transmitting at least a derived physiological movement signal. and processor-executable instructions for detecting unauthorized activity using at least a portion of the and the process of generating an alarm or communication to notify the user or a third party. This communication may include instructions executable by the smartphone or An alarm may be provided on the watch.
[0017] Some versions of the technology may include a server. The server may have access to any processor-readable medium such as a The processor-readable medium may include a processor-executable instruction readable from the processor via a network. and receiving a request to download the audio data to the interactive audio device. do.
[0018] Some versions of the technology may include interactive audio devices. An interactive audio device may include one or more processors. The interactive audio device may include a speaker connected to one or more processors. The interactive audio device may include a microphone connected to one or more processors. may include any processor-readable medium described herein; and / or , one or more processors may be configured in any of the server(s) described herein. Optionally, the processor may be configured to access instructions executable by the processor. The interactive audio device is a portable and / or handheld device. Interactive audio devices include mobile phones, smart watches, and tablet computers. It may include a computer or a smart speaker.
[0019] Some versions of the technology may be implemented using processor-readable media as described herein. This may include a method of a server having access to any of the entities. The method includes transmitting processor-executable instructions on a processor-readable medium to a network. The server sends a request to the interactive audio device via the network. The method may include receiving a request executable by a processor in response to the request. This may include sending instructions to an interactive audio device.
[0020] Some versions of the technology include a method for a processor of an interactive audio device. The method may be implemented by transmitting the process to any processor-readable medium described herein. The method may include accessing the data using a processor-readable medium. This may include executing processor-executable instructions in a processor.
[0021] Some versions of this technology are designed for interactive audiovisual detection of the user's physiological movements. The method may include a method for an interactive audio device, the method comprising: The method may include generating an audio signal in the vicinity of the interactive audio device via a speaker. The method involves detecting reflected sound from nearby sources via a microphone connected to an interactive audio device. The method may include instructions for controlling sensing of the reflected sound signal. and processing the physiological movement signal with a signal indicative of at least a portion of the audio signal. The method may include deriving at least one of the derived physiological movement signals within the body. The method may also include generating an output from the processor based in part on the evaluation.
[0022] In some versions, at least a portion of the generated audio signal (e.g., Some of the noise generated by the oscillating sound source (used for sensing applications) may be substantially in the inaudible range. The portion of the voice signal may be a low frequency ultrasonic acoustic signal. The generated oscillator signal or the direct path measurement signal may include: or (b) using at least a portion of the detected sound signal and a portion of the sensed reflected sound signal. ) associated with a portion of the sensed reflected sound signal and at least a portion of the generated sound signal. deriving a physiological movement signal using a related signal that may be obtained. , the relevant signal can be an internally generated oscillator signal or a direct path measurement signal. The method generates a physiological response by multiplying an oscillator signal by a portion of the sensed reflected sound signal. The method may include deriving a movement signal. This is done in the processor via a microphone connected to a talk-type audio device. Generating the output may be in response to a sensed audible verbal communication. The method includes demodulating a portion of the sensed reflected sound signal with at least a portion of the sound signal. The demodulation may include deriving a physiological movement signal by The generation of the sound signal may include multiplying the signal by a portion of the sensed reflected sound signal. A dual-tone frequency modulated continuous wave signal may be generated. The signal is a first sawtooth frequency change superimposed with a second sawtooth frequency change in a repeating waveform. The method may include generating an ultra-wideband (UWB) audio signal as audible white noise. The method may include: detecting a user's movements using the UWB audio signal. During the setup process, a probing sound sequence is generated from the speaker and The method may include calibrating the distance measurement of the high frequency ultrasonic echoes.
[0023] In some versions, the method includes time-synchronizing the setup process. A calibration signal is generated from one or more speakers, including a speaker, and the microphone and may include estimating a distance between another microphone of the interactive audio device. The method may further localize the detected region by activating a beamforming process. Optionally, evaluating at least a portion of the derived physiological movement signals. The output generated based on the derived output may include monitored user sleep information. The evaluation of at least a portion of the physical movement signals may include detecting one or more physiological parameters. The one or more physiological parameters may include respiratory rate, relative amplitude of respiration, heart rate, The parameters may include any one or more of: heart rate, cardiac amplitude, relative cardiac amplitude, and heart rate variability. The monitored user sleep information includes sleep score, sleep stages and sleep stage time. The generated output may include an interactive query answer presentation. The generated interactive query answer presentation may be performed via a speaker. Interactive query and response presentation provides recommendations to improve monitored user sleep information. The device may include a chair.
[0024] In some versions, at least a portion of the derived physiological movement signal The output generated based on the evaluation may be accessed and / or used by a server on a network. The search may further be based on performing a search of network resources. The physiological motion signal may be generated based at least in part on an evaluation of the derived physiological motion signal. The produced output may include a control signal for control of an automated device or system. The method includes transmitting a control signal over a network to an automated device or system. The network may be the Internet. and evaluating at least a portion of the derived physiological movement signals by changing the settings of the audio device. generating a control signal in the processor for performing the operation based on the value. The control signals for changing the settings of the interactive audio device are This includes changing the volume based on detecting the user's distance from the chair, the user's state, or the user's location. The method includes evaluating the motion characteristics of different acoustic sensing ranges in a processor. It may further include monitoring sleep characteristics of multiple users.
[0025] The method comprises: detecting different users at different frequencies; The method may further include controlling generation of the acoustic sensing signal at different times. and controlling generation of interleaved acoustic sensing signals for sensing at different times. The method further comprises determining a user's activity based at least in part on the derived physiological movement signal. The method may include detecting the presence or absence of a derived physiological movement signal. and performing, by the processor, biometric recognition of the user based at least in part on the The method may include (a) analyzing at least a portion of the derived physiological movement signal. (b) biometric assessments determined from and / or derived physiological movements and transmitting communications over the network based on presence detection determined from analysis of at least a portion of the signal. The method may include generating, by a processor, an audio signal indicating the presence of user movement. audibly detecting the user; audibly addressing the user; audibly authenticating the user; and authorizing users or communicating alarm messages about users. Optionally, the method may include using an interactive audio device to prompt the user to Audible authentication combines the sound waves of the user's words picked up by a microphone with the sound of the event. The method may include comparing the speech sound waves with previously recorded speech sound waves. This may include authenticating the user through metric sensing.
[0026] In some versions, the method comprises: and detecting, by an interactive audio device, movement gestures based in part on the audio signal. The method may include sending a notification or (a) using an automated device and / or (b) using an automated device and / or (c) using an automated device and / or (d) using an automated device and / or (e) using an automated device and / or (f ... b) detecting a control signal for controlling a change to the operation of the interactive audio device; and generating the motion gestures in the processor of the interactive audio device based on the motion gestures. A control for controlling changes to the operation of the interactive audio device may include The signal is used to initiate the interactive voice assistant operation of the interactive audio device. activating microphone sensing, thereby enabling the interactive voice assistant to The method can be implemented by using microphone-sensing language keywords. The method may include initiating an interactive voice assistant action based on detecting the code. detecting a user's respiratory movement based at least in part on the derived physiological movement signal; and generating output cues to trigger the user to regulate their breathing. by the interactive audio device. Interactive audio device communication from another interactive audio device controlled by the to an interactive audio device via inaudible sound waves sensed by a microphone in The method may include receiving at least a portion of the sound signal from an interactive audio device. and adjusting parameters for generating the signal based on the communication in the vicinity of the device. The parameters for adjustment allow for the interaction of interactive audio devices with other interactive audio devices. Interference between the derived physiological movement signal and the audio device can be reduced. Some evaluations may further include detecting sleep onset or wake onset. The output may include a service control signal. The service control signal may be a signal for lighting control, appliance control, May include one or more of a volume control, a thermostat control, and a window covering control.
[0027] The method involves prompting a user to collect feedback and generating a derived result in response. Providing advice based on at least one of physical movement signals and feedback The method may include generating the environmental data. The method may be further based on the determined environmental data. and, based on at least a portion of the determined environmental data and the derived physiological movement signal. and generating a control signal for an environmental control system based on the measured temperature. The sleep enhancement service may further include controlling the provision of the sleep enhancement service. (a) the detected sleep state and / or the collected sleep data; (b) generating advice in response to user feedback; and (c) adjusting the sleep environment. To provide users with state setting and / or sleep-related advisory messages The method further comprises generating a control signal for controlling an environmental device based on the derived physiological information. Gesture detection based on biological movement signals and interactive audio device - Patents.com Initiating microphone sensing for initiation of the interactive voice assistant operation. The method may include determining whether to predict a nap based at least in part on the derived physiological movement signal. The method may include initiating and monitoring a session. and varying at least some parameters of the generated sound signal to track the sound through the The method may include varying the detection range by detecting unauthorized movement. detecting the physiological movement using at least a portion of the derived physiological movement signal; and generating an alarm or communication for notification to a user or third party. The communication provides an alarm on the smartphone or smartwatch.
[0028] The methods, systems, devices and apparatus described herein allow a processor to functionality (e.g., general-purpose or special-purpose computers, portable computer processing units ( For example, mobile phones, smart watches, tablet computers, smart speakers, smartphones, etc. Smart TV, etc.), respiratory monitor and / or microphone and speaker. This may allow for an improvement in the functionality of the processor of other processing devices. , systems, devices and apparatuses for automated smart audio device technology This may allow for improvements in the field of technology.
[0029] Of course, some of the above aspects may form sub-aspects of the present technology. and / or various combinations of the various aspects may be used to provide further aspects of the present technology. Or it may constitute a sub-embodiment.
[0030] Other features of the present technology are included in the following detailed description, abstract, drawings, and claims. This becomes clear in light of the information available.
[0031] 4 Brief description of the drawings The present technology is illustrated by way of example and not limitation in the accompanying drawings, in which like reference characters refer to: It contains the following similar elements: [Brief explanation of the drawings]
[0032] [Figure 1] 1 illustrates an exemplary voice-enabled audio device (e.g., low-frequency ultrasound bio-motion sensing using the signal generation and processing techniques described herein). [Figure 2] 1 illustrates an exemplary processing device that receives audio information from a vicinity of the device and a schematic diagram of an exemplary process of the device. [Figure 3] FIG. 1 is a schematic diagram of a processing device (e.g., a smart speaker device) as may be configured in accordance with some aspects of the present technology. [Figure 4A] For example, the frequency characteristics of a single-tone chirp for frequency modulated continuous wave sensing (FMCW) are shown. [Figure 4B] For example, the frequency characteristics of dual-tone chirp for frequency modulated continuous wave sensing (FMCW) are shown. [Figure 5] 1 illustrates an exemplary demodulation for dual-tone FMCW, which may be implemented for a sensing system of the present technology. [Figure 6] 1 illustrates an exemplary operation of a voice-enabled audio device (e.g., low-frequency ultrasound bio-motion sensing using the signal generation and processing techniques described herein). [Figure 7] For example, exemplary audio processing modules or blocks for the processing described herein are shown. [Figure 8] 10 illustrates exemplary output (e.g., sleep stage data) produced by processing audio-derived movement features. [Figure 8A] For example, various signal characteristics of a triangular single tone for FMCW systems are shown. [Figure 8B] For example, various signal characteristics of a triangular single tone for FMCW systems are shown. [Figure 8C] For example, various signal characteristics of a triangular single tone for FMCW systems are shown. [Figure 9A] For example, various signal characteristics of triangular dual tone for FMCW system are shown. [Figure 9B] For example, various signal characteristics of triangular dual tone for FMCW system are shown. [Figure 9C] For example, various signal characteristics of triangular dual tone for FMCW system are shown. DETAILED DESCRIPTION OF THE INVENTION
[0033] 5 Detailed Description of the Embodiments of the Present Technology Before describing the present technology in more detail, it is important to note that the present technology is not limited to the different methods described herein. It should be understood that the present disclosure is not limited to the specific examples that may be used. The terminology used herein is for the purpose of describing the specific embodiments described herein. It should also be understood that this is not limiting.
[0034] The following description is provided in connection with various aspects of the present technology that may share common characteristics or features. One or more features of any one aspect may be combined with one or more features of another aspect or other aspects. It should be understood that combinations are possible. Any single feature or combination of features in any of these may be used in further exemplary forms. It can constitute a state.
[0035] 5.1 Screening, Monitoring and Detection by Audio Equipment This technology can detect the movements of a subject while they are sleeping (e.g., whole body movements, breathing movements, etc.). Audio systems, methods and methods for detecting heart and / or chest movements related to heart More particularly, the present invention relates to interactive audio devices (e.g., smart speakers) and devices. ) processing application that works with audio devices. The service is available on smartphones, tablets, mobile devices, mobile phones, smart TVs, The device may be a laptop, a mobile phone, or the like, which may include device sensors (e.g., speakers, Such movements are detected using sensors (cameras and microphones).
[0036] The following is a particularly minimal and unobtrusive version of an exemplary system suitable for implementing the present technology: The interactive audio device will be described with reference to FIGS. 1 to 3. embodied as a processing device 100 having a processor (e.g., a microcontroller). These processors execute an application 200 for detecting the movement of the object 110. The processing device 100 may be a smart speaker configured with a It can be placed on a bedside table or anywhere else in the room. Alternatively, the processing device 100 may be, for example, a smartphone or a tablet computer. a computer, laptop, smart television or other electronic device. The processor(s) of the processing device 100 may, among other things, 200 (e.g., generally open or unrestricted) Typically, audio is transmitted through air as a medium (e.g., the medium in a room near the device). The processing device may be, for example, a transducer (e.g., a microphone). The processing device may receive reflections of the transmitted signal by sensing the reflections with a microphone. The sensed signal is processed (e.g., by demodulation with the transmitted signal) to determine the body motion (e.g., For example, the processing device 100 may determine the motion of the whole body, the cardiac motion, and the respiratory motion. Typically, this would include a speaker and a microphone, among other components. The speaker is connected to the generated audio signal and a microphone to receive the reflected signal. The system may be implemented to transmit the generated audio for sensing and processing. The signal is disclosed in International Patent Application PCT / EP2017 / 073613 (filing date: September 1, 2017). 9), which is incorporated herein by reference in its entirety. In the version shown in Figs. 1 shows various processing devices with integrated sensing devices (e.g., sensing devices in a housing). Includes all sensing devices or components (e.g., microphones and speakers) In some versions, the sensing devices are housed in separate or individually may be components connected or unconnected via wired and / or wireless connection(s). works.
[0037] Sensing devices are referred to herein primarily in the context of acoustic sensing (e.g., low frequency ultrasonic sensing). Although described, it should be understood that these methods and devices may be implemented using other sensing technologies. It will be understood that, for example, the processing device may alternatively be configured to function as a sensing device. The RF sensor may be implemented with a radio frequency (RF) transceiver, thereby generating The transmitted and reflected signals are RF signals. Such an RF sensing device may be used in conjunction with a processing device. The technology and sensor controller described below may be integrated with or connected to a processing device. The present invention may be embodied using any of the components of International Patent Application No. PCT / US 2013 / 051250 (Title: "Range Gated Radio Freq "Efficiency Physiology Sensor," application date: July 19, 2013 , International Patent Application No. PCT / EP2017 / 070773 (title: "Digital Radio Frequency Motion Detection Sensor , filing date: August 16, 2017), and International Patent Application No. PCT / EP2016 / 06 9413 (Title: "Digital Range Gated Radio Fre "Quency Sensor," application date: August 16, 2017. Similarly, another In this version, such a sensor for transmitting a signal and sensing its reflection is The sensing device may include an infrared radiation generator and an infrared radiation detector (e.g., an IR emitter and Such signal processing for motion detection and characterization can also be implemented using a can be similarly embodied.
[0038] Combining two or more of these different sensing technologies can combine the advantages of each technology. For example, the acoustic sensing technology described can be used to It is quite tolerable for use in the noisy environments of everyday life. For sensitive users, the noise is significantly lower when using this technology at night, resulting in a lower perceived signal. Similarly, IR sensing can be problematic at night, as it can make signals more audible. It has a good signal-to-noise ratio, but daytime light (and heat) can make it problematic to use. In this case, the use of IR sensing at night can be complemented by the use of acoustic sensing during the day. It is possible.
[0039] Optionally, the audio-based sensing method of the processing device is adapted to detect other types of devices (e.g. , bedside devices (e.g., respiratory therapy devices (e.g., continuous positive airway pressure (e.g., , "CPAP") device or high flow therapy device) (where the therapy device is a treatment device 5, which may function as a separate processing device 100 or may operate in conjunction with a separate processing device 100. embodied in or by the illustrated respiratory treatment device 5000) Examples of such devices include pressure devices or blowers (e.g., motor and impeller in the lute), one or more sensors of the pressure device or blower and the central control device, see International Patent Publication WO / 2015 / 061848 (Application No. PC T / AU2014 / 050315) (Application date: October 28, 2014) and international patent Publication WO / 2016 / 145483 (Application No. PCT / AU2016 / 050117) The device described in the application date: March 14, 2016 can be considered. The entire disclosure of such respiratory therapy devices is incorporated herein by reference. The chair 5000 may include an optional humidifier 4000 and may be connected to the patient interface 3000. The supply of blood may be provided via a patient circuit 4170 (e.g., a conduit). Therefore, the respiratory treatment device 5000 may be configured to perform external audio-related acoustic conditions (such as those described throughout this application). Internal voice-related conditions in and through the patient circuit 4170 (as opposed to conditions It may have individual sensors (eg, microphones) that sense it.
[0040] The processing device 100 may be configured to determine the effectiveness of monitoring the subject's respiration and / or other movement-related characteristics. When used during sleep, the treatment The device 100 and associated methods may, for example, detect the user's breathing and determine sleep stages, sleep patterns, and may be used to identify transitions between sleep states, states, breathing and / or other respiratory characteristics When used during wakefulness, the processing device 100 and its associated methods may be used to monitor the behavior of a person or subject. Movements such as the presence or absence of inspiration (inspiration, expiration, pauses, and derived rate or number) and / or or for detecting ballistocardiogram waveforms and the subsequent derived heart rate. Such movements or movement characteristics may be used to perform a variety of functions as described herein. It can be used for more detailed control.
[0041] The processing device 100 may include integrated chips, memory and / or other control instructions, data or For example, the assessment / signal processing methods described herein may include an information storage medium. The included programmed instructions are then used to create a device that forms an application specific integrated chip (ASIC). Such instructions may be coded on an integrated chip in the memory of a device or apparatus. Additionally or alternatively, software or firmware may be stored on a suitable data carrier. Optionally, such processing instructions may be transmitted over a network, for example. may be downloaded to a processing device from a server (e.g., the Internet) via a network, When these instructions are executed, the processing device may then be a screening device or acts as a monitor device.
[0042] Thus, processing device 100 may include multiple components as shown in FIG. The processing device 100 may include, among other components, a microphone (or microphones). a processor(s) 304, or an audio sensor 302, an optional Display Interface 306 Optional User Control / Input Interface 30 8, speaker(s) 310, and memory / data storage 312 (e.g., (using the processing instructions of the processing methods / modules described in this specification). In some cases, the microphone and / or speaker may be connected to the device's processor (singular It may function as a user interface with the The processing device responds to the audio and / or verbal commands sensed by the speaker. In response via the processing device, the processing device may control the operation of the processing device. 100 can function as a voice assistant using natural language processing, for example.
[0043] One or more of the components of the processing device 100 may be integrated with the processing device 100. For example, a microphone (single or multiple microphones) may be connected to the The sound sensor(s) 302 may be integrated with the processing device 100 or For example, via a wired link or a wireless link (e.g., Bluetooth, Wi-Fi, etc.) The processing device 100 may be coupled to the data processing device 100. A communication interface 314 may be included.
[0044] The memory / data storage device 312 stores multiple processors for control of the processor 304. For example, the memory / data storage device 312 may include control instructions for the processes described herein. A processor for executing an application 200 by a processing method / module processing instruction It may include control instructions.
[0045] Examples of the present technology may be configured to use one or more algorithms or processes. These algorithms or processes may be used while the user is asleep when using the processing device 100. application to detect movement, breathing and optionally sleep characteristics when For example, application 200 may be embodied by , which may be characterized by several sub-processes or modules. The application 200 includes an audio signal generation and transmission subprocess 202; Motion and biophysical property detection subprocess 204 and, for example, object absence / presence detection, motion detection for biometric identification, sleep characterization, respiratory or cardiac related characterization, etc. A sex analysis sub-process 206 and a method for, for example, presenting information or as described in more detail herein. and a result output subprocess 208 for controlling various devices.
[0046] For example, the optional sleep staging in process 206 may be implemented as, for example, a sleep stage processing model. However, among such processing modules / blocks, Any one or more of the following may optionally be added (e.g., sleep scoring or staging, Object recognition processing, motion monitoring and / or prediction processing, device control logic processing, or other In some cases, signal post-processing functionality is described in the following patents or patent applications: Any of the components, devices and / or methods of any of the apparatus, systems and methods described The disclosures of the following documents are incorporated herein by reference in their entirety: and is incorporated herein by reference in its entirety. / 070196 (Application date: June 1, 2007, Title: "Apparatus, S system, and Method for Monitoring Physiol "Ogical Signs"), International Patent Application No. PCT / US2007 / 083155( Application date: October 31, 2007, Title: "System and Method for Monitoring Cardio-Respiratory Parameters ters"), International Patent Application No. PCT / US2009 / 058020 (filing date: 2009 September 23, 2017, Title: "Contactless and Minimal-Con tact Monitoring of Quality of Life Param eters for Assessment and Intervention”), International Application No. PCT / US2010 / 023177 (filing date: February 4, 2010, Title Ru: “Apparatus, System, and Method for Chr. "Sonic Disease Monitoring"), International Patent Application No. PCT / AU2 013 / 000564 (Application date: March 30, 2013, Title: "Method a nd Apparatus for Monitoring Cardio-Pulmo "Nary Health"), International Patent Application No. PCT / AU2015 / 050273 (issued Date of application: May 25, 2015, Title: "Methods and Apparatu" "S for Monitoring Chronic Disease" International patent issued Application No. PCT / AU2014 / 059311 (filing date: October 6, 2014, title: “Fatigue Monitoring and Management System m"), International Patent Application No. PCT / EP2017 / 070773 (filing date: August 2017) 16th, Title: "Digital Radio Frequency Motion "Detection Sensor"), International Patent Application No. PCT / AU2013 / 06 0652 (Application date: September 19, 2013, Title: "System and Met "hod for Determining Sleep Stage"), international patent application No. PCT / EP2016 / 058789 (Application date: April 20, 2016, Title: Detection and Identification of a Human from Characteristic Signals), International Patent Application No. PCT / EP2016 / 080267 (Application date: December 8, 2016, Title: "Peri odic Limb Movement Recognition with Sens ors), International Patent Application No. PCT / EP2016 / 069496 (filing date: 2016) April 17th, Title: "Screener for Sleep Disorders "D Breathing"), International Patent Application No. PCT / EP2016 / 058806 (issued Application date: April 20, 2016, Title: "Gesture Recognition with Sensors"), International Patent Application No. PCT / EP2016 / 069413( Application date: August 16, 2016, Title: "Digital Range Gated "Radio Frequency Sensor"), International Patent Application No. PCT / EP2 016 / 070169 (Application date: August 26, 2016, Title: "Systems and Methods for Monitoring and Management t of Chronic Disease), International Patent Application No. PCT / US2014 / 045814 (Application date: July 8, 2014, Title: "Methods and "Systems for Sleep Management" (U.S. Patent Application No. 15 / 079,339 (Application date: March 24, 2016, Title: "Detection Thus, in some examples, In this case, the processing of detected movements (e.g., breathing movements) may include any one or more of the following: These may serve as criteria for determining: (a) sleep states indicative of sleep; (b) wakefulness; (c) sleep stages indicating deep sleep; (d) sleep stages indicating light sleep; and (e) sleep stages indicative of REM sleep. The technique involves the use of different mechanisms / processes for motion sensing (e.g., speakers and microphones). Although these cited references provide Compared to radar or RF sensing techniques described in some of the literature, respiratory signals (e.g., Respiratory rate obtained using the audio sensing / processing methods described herein) followed by sleep state / stage The principles of processing respiratory or other movement signals for information extraction are discussed in these cited references. For example, the respiratory rate and movement and activity counts can be determined by the method of determining the movement and activity counts. Once determined by RF or SONAR, sleep stages are shared based on As a further example, the sensing wavelength may be implemented by RF pulse CW and sonar. This may differ between FMCW implementations, e.g., range (different sensing distances). ), the velocity can be determined differently. ,Motion detection can be performed in multiple ranges.,Thus, tracking of one or more moving targets ( The presence of a person, whether two people or actually different parts of one person, relative to the sonar sensor. This can be done depending on the angle.
[0047] Typically, the audio signal from the speaker includes one or more tones as described herein. Tones can be generated and transmitted to the user to sense audio signals, etc. This gives the pressure change of a medium (e.g., air) at one or more specific frequencies. For purposes of this description, the generated tone (or audio signal or voice signal) is Audible pressure waves can be generated (e.g., by a speaker), and therefore can be referred to as "voice," "sound," or " However, in this specification, such pressure modifications and tones ( (singular or plural) refers to any of the terms "voice," "sound," or "audio" Regardless of how they are characterized, they should be understood as being audible or inaudible. Thus, the generated audio signal may be audible or inaudible, and may be of audible or inaudible to the human population. The frequency threshold for speech (e.g., above 18 kHz) changes with age. This signal is effectively inaudible, since no audio can be distinguished (to the extent that it can be heard). The standard range for typical "audio frequencies" is approximately 20Hz to 20,000Hz. (20 kHz). High-frequency hearing thresholds tend to decrease with age, and Humans often cannot hear sounds with frequencies above 15-17 kHz, and teenagers In this case, 18 kHz can be heard. The most important frequencies in conversation are around 250 ~6,000Hz. Typical consumer smartphone speakers and microphones The Phonon signal response often rolls off above 19-20kHz. Some are designed to roll off above 23kHz ( In particular, devices that support sampling rates above 48kHz (e.g., 96kHz) Therefore, for most humans, the 17 / 18 to 24 kHz range It is possible to use signals within 18kHz and remain inaudible. For younger people who cannot hear kHz, the band from 19kHz to, for example, 21kHz is used. In the case of some household pets, higher frequencies can also be heard (e.g. Please note that the maximum frequency range for dogs is 60 kHz, and for cats it is 79 kHz. A suitable range of sensed audio signals may be in the low ultrasonic frequency range (e.g., 15 ~24kHz, 18~24kHz, 19~24kHz, 15~20kHz, 18~20k Hz or 19-20kHz).
[0048] For example, a low frequency ultrasonic sensing signal as described in PCT / EP2017 / 073613 may be used. Any of the audio sensing arrangements and methods described herein may be used in conjunction with the processing data. However, in some cases, dual-tone FMCW (Dual-tone FMCW) The ALUMP technology may be implemented as described herein.
[0049] For example, a triangular FMCW waveform with one "tone" (i.e., the frequency is swept The generation of what is picked up and what is down is determined by the processing device being used to generate the ) can be used. The waveform has the frequency vs. time characteristics shown in Figure 4A. Only sweep or downsweep processing or both processes are evaluated for distance detection. The reason why a phase-continuous triangular shape for one tone is highly desirable is that This minimizes any audible artifacts in the reproduced audio that may be caused by phase discontinuities. There are points where the signal is removed or eliminated. Amplitude audio playback from a much lower (or much higher) amplitude at similar amplitudes in sample space. When the speaker(s) are asked to jump to a higher (higher) frequency, Can cause a pleasant audible buzzing noise, i.e. due to mechanical changes in the speaker Clicking sounds can be produced by chirping, and buzzing sounds can be produced by frequent repetition of chirps. The user hears a sound (a number of closely spaced clicks).
[0050] Alternatively, in some versions of the present technology, a ramp waveform (e.g., up-switching) may be used. Special dual "tone" with only upsweep or only downsweep By implementing the acoustic sensing signal as FMCW, one ramp (ramp of frequency) The frequency ramp (the sudden change in frequency from the end of the up and down) to the next ramp (frequency ramp This dual "tone" The frequency characteristics of the frequency modulation waveform compared with time are shown in Figure 4B. This allows for extremely simple data processing in the system, and high amplitude at each point of the triangular waveform This also eliminates the possibility of transitions. If sudden transitions occur repeatedly, the system's low-level DSP / CODEC / Firmware) may trigger strange behavior.
[0051] 4A and 4B show FMCW single-tone (FIG. 4A) and dual-tone (FIG. 4B) implementations. B) Comparison of the frequency domain between runs. To ensure inaudibility, a single tone (Figure 4) was used. A) may preferentially include a downsweep (a decrease in the generated frequency over time). The tones may be omitted, but in that case a constant audibility may be produced. By using a pair (Figure 4B), the time domain representation is shaped to be inaudible. This can help avoid the need for such a downsweep. One tone 4001 and an optional second tone 4002 are shown overlapping. In this figure, the received echoes (ie, reflection signals) are not shown. Thus, the tone creates a first sawtooth frequency change superimposed with a second sawtooth frequency change in the repeating waveform. These tones are repeated in a sequence that is repeated over the sensing period. To be continued.
[0052] Therefore, when implementing a low-frequency ultrasonic sensing system using the FMCW approach, the acoustic sensing signal There are different ways in which the signal can be generated, depending on the waveform shape in the frequency domain (e.g. triangular). (symmetric or asymmetric), ramp (inclined, sinusoidal), period ("chirp" ) duration), and bandwidth (frequencies covered by the "chirp" (e.g., 19 In FMCW configurations, two or more simultaneous tones can be used. It is also possible that
[0053] The number of samples selected determines the possible output demodulation sampling rate (e.g., sampling For 512 samples at a rate of 48 kHz, the frequency is 93.75 Hz (48,000 / 51 2), the sweep time for a 4096 sample duration is 11.72 Hz (48,000 (equivalent to 1 / 4096). Triangle waveform has 1500 sample uptime and 1500 sample When used with pull-down time, the output sampling rate is 16Hz (48,0 In this type of system, the signal is multiplied by a reference template, for example. By calculating the time, synchronization can be achieved.
[0054] Regarding the selection of output sampling rate, experimental testing has shown that: In other words, the range of roughly 8 to 16 Hz is preferable. This reduces the possibility of 1 / f noise (air movement, strong fading and / or room motion). low frequency effects due to the harmonics are largely avoided, and The reverberant region seen is avoided (i.e., the perceived reverberant region is reached before the next similar component in the next "chirp"). Allowing time for energy decay at any one frequency in the waveform "chirp" In other words, if the bin is too wide, changes in airflow and temperature (e.g., door If there are openings in the room and heat coming in and out, any block you are looking at now is a breathing This means that the data may contain unwanted baseline drift that appears to be This means that waves are seen across a band (across range bins) as the air moves. This means that the heat generated by a tabletop or pedestal fan or by an air conditioner or other HVAC system In fact, if the blocks are made too wide, the system On the other hand, if the system has too high a refresh rate, the system will "look like" a CW system. If run in slate (i.e. slope too short) reverb can be produced .
[0055] As shown in FIG. 4A, a triangular FMCW waveform has one "tone" (i.e., frequency If the number goes up and down, the system can, for example, perform only up sweeps or only down sweeps. or in fact both may be processed for distance detection. The reason why a triangular shape with continuous phase is highly desirable is that it can be easily removed due to phase discontinuities. This minimizes or eliminates any audible artifacts in the resulting playback sound. A variant of this, tilt, is the sampling from a particular amplitude speech playback at a certain frequency. Jump to much lower (or much higher) frequencies at similar amplitudes in space. When the speaker(s) are requested, they cause an extremely annoying audible buzzing noise. That is, mechanical changes in the speaker can cause clicks and The frequent repetition of the chirp creates a buzzing sound (a large number of closely spaced clicks) that the user It sounds like...
[0056] Therefore, in some versions of the present technology, a ramp waveform ( For example, a special design with only up sweeps or only down sweeps By implementing the acoustic sensing signal as FMCW using dual "tones", (the sudden change in frequency from the end of a frequency ramp up and down) to the next ramp (frequency ramping up and down) without audible artifacts. Dual "tone" frequency modulation waveforms exhibit frequency characteristics over time and have at least two Two varying frequency ramps overlap in one period, and each of these frequency ramps is At any instant in time during a period such as the ramping duration, This is illustrated in FIG. 4B in relation to the dashed-dotted line versus the solid line. This allows the data processing in the system to be extremely simple, The possibility of high amplitude transitions at each point in the waveform is also eliminated. This triggers strange behavior in the system's low-level DSP / CODEC / firmware This may be the case.
[0057] An important consideration when implementing such dual tone signals is the speaker / system. The signal shape is created so that the system has no sharp transitions and has a zero point. This allows for the filtering to be done to make the signal inaudible. For example, the signal can be made inaudible and sensitive, while the high-pass filter can be used to reduce the need for Or bandpass filtering can be avoided. This simplifies the synchronization of transmission and reception (e.g., for demodulation) of such signals. This simplifies signal processing. Dual tone results in more than one tone being used. This provides an element of fading robustness. The frequency and phase of the signal can be varied or varied with frequency (e.g., dual A 100Hz offset can be used between FMCW tones in a tone system. ).
[0058] The performance of the FMCW single tone in Figure 4A and the FMCW dual tone in Figure 4B is shown in Figures 8 and 9. 8A, 8B and 8C are the FMCW single-mode signals of FIG. 7A. 9A, 9B, and 9C show the signal characteristics of the FMCW duplexer of FIG. 7B. 1 shows example signal characteristics for Altone.
[0059] FIG. 8A shows the transmitted (Tx) signal 8001 and the received (Rx) reflection 8001-R (Echo) is a triangular single tone FMCW operating in acoustic sensing systems. Figure 8B shows the time domain waveform. Figure 8C shows the spectrum of the signal. As can be seen, the peak area is at lower frequencies (relative to the bandwidth of the FMCW signal). Therefore, there is content even in the lower frequencies (outside the audible range). This can fall within a range of wavenumbers, leading to undesirable performance characteristics.
[0060] FIG. 9A shows a dual tone ramp FMCW signal in signal graph 9002. mp FMCW signal). Signal graph 9002 shows both tones, and signal graph 900 2-R shows the received echo of these two tones / multitones. The shape of the dual-tone cosine-like function is shown in Figure 9, along with the zero points (leading to the zero crossings). C shows a much smaller peak at lower frequency and a larger output amplitude. When comparing the rope region SR with the slope region SR in Figure 8C, the / at lower frequencies The figure shows the rapid drop in power (dB) of the dual-tone ramp FMCW for lower frequencies. from high frequencies (which are practically inaudible and are used for sensing) to lower frequencies ( A steeper roll-off to frequencies (audible and typically not used for sensing) This is a desirable acoustic sensing characteristic because it is less intrusive to the user. The output at lower frequencies (outside the peak region related to the signal bandwidth) is shown in Figure 8C. This can be 40 dB lower than the single-tone FMCW triangular mode, as shown in Figure 9C. Compare the upper smooth peak region PR of FIG. 9C with the multi-edge peak region PR of FIG. 8C. The dual-tone ramp FMCW signal may have better acoustic sensing characteristics and may be used with a speaker. There are few demands for
[0061] Such multiple tone FMCW or dual tone FMCW systems (e.g., (running on a single-board computer based on NACS) It is possible to obtain detection that can identify multiple people within the detection range. It can detect heart rate and respiratory rate (singular or plural) at 1.5 meters from the chair. ) can be detected up to about 4 meters or more. An exemplary system uses two tones It can be used at 18,000 Hz and 18,011.72 Hz. The frequency is ramped up to, for example, 19,172 Hz and 19,183.72 Hz, respectively. ; tilt) can be.
[0062] For this 1,172Hz ramp, for example, FF of size 4096 points It is conceivable to use T with a bin width of 48,000 Hz / 4096 = 11.72. Since the speed of sound is 340 m / s, please note the following: 340 m / s / 11.72 / 2 (outgoing and returning) = 14.5m (per 100 bottles) or 14.5cm (per bin). Each "bin" can hold, for example, up to one person per bin. It is possible to detect this (but in reality, people are farther apart than this). As part of the process, for example, a more computationally expensive method (multiplying the signal by a reference template) can be used. To avoid highly correlated behavior of the signal, the signal can be squared. Independently, the maximum size resolution is speed of sound / (bandwidth*2)=340 / (1172*2)= However, the perceived reflected signal and the perceived direct path signal are A synchronization process may optionally be provided that includes cross-correlation of the reference temperature. The method may optionally include multiplying the plate by at least a portion of the sensed reflected sound signal. do.
[0063] Such multiple tone FMCW or dual tone FMCW systems (e.g., (running on a single-board computer based on NACS) It is possible to obtain detection that can identify multiple people within the detection range. It can detect heart rate and respiratory rate (singular or plural) at 1.5 meters from the chair. ) can be detected at distances of up to about 4 meters or more. An exemplary system uses two tones can be used at 18,000 Hz and 18,011.72 Hz, The tones are ramped, for example, to 19,172 Hz and 19,183.72 Hz, respectively. (slanted) possible.
[0064] For this 1,172 Hz ramp, let's take an FFT of size 4096 points and bin it. It can be considered to use with a width of 48,000Hz / 4096=11.72. Note that since the velocity is 40 m / s, it is 340 ms / s / 11. 72 / 2 (outgoing and returning) = 14.5 m (per 100 bins) or 14.5 cm (per bin). Each "bin" can detect, for example, up to one person (per bin). (But in reality, people are farther apart than this.) As a result, the computationally more expensive correlation (multiplying the signal by a reference template) can be achieved. To avoid correlations, the signal can be squared. Independent of the FFT size used, The resolution of the maximum size is speed of sound / (bandwidth*2)=340 / (1172*2)=14.5c m.
[0065] Figure 5 shows the " 1 shows an example of "self-mixing" demodulation. In this case, the received echo signal is compared with the generated transmitted signal (e.g., a signal from an oscillator). signal) to determine the distance or movement within the range of the speaker or processing device 100. This process produces an "intermediate" frequency (IF) signal and a signal that reflects the FMCW provides a "beat frequency" signal, also called a local oscillator. The received Rx signal is either transmitted by the transmitter or by the signal itself as described in more detail herein. When demodulated and low-pass filtered, it is not considered baseband as is. If the IF signal is subjected to, for example, a Fast Fourier Transform (FFT) process, an abnormal "intermediate" signal may be generated. ), this signal can be brought to baseband (BB).
[0066] As shown in Figure 5, demodulation is performed only on the received (reflected sound signal) Rx signal. A possible reason is that the Rx signal is a large percentage of the signal that represents the transmit (Tx) signal. (e.g. the generated sound is partly transmitted from the speaker to the microphone) There are points on the direct path that can travel and be sensed along with the reflected sound. The signal Rx can be multiplied by itself (e.g., by squaring) and the demodulation can be viewed as a multiplication operation. After that, a filtering process (e.g. low pass) is performed. It can be.
[0067] In Figure 5, self-mixing is illustrated, where the motion signal is mixed with the reflection signal and Several different methods are available for deriving the sensed signal (i.e., Tx or audio signal). In one such version, the local A local oscillator (LO) (which may also generate an audio signal) generates a copy of the Tx signal for demodulation. The actual Tx signal generated may be delayed or distorted. Therefore, the internal signal from the local oscillator LO may differ slightly. The signal from (Tx)*Rx can be demodulated, and then filtered (e.g. For example, a low pass filter may also be performed.
[0068] In another version, two local oscillators are implemented to generate two LO signals. For example, sine and cosine copies of the LO signal can be implemented to generate a Typically, a signal (sine or cosine) is generated from an oscillator. ) is transmitted. The exact Tx signal may differ from the local oscillator due to delay or distortion. In this version, (a)RX*LO(Si n) and (b) RX*LO(Cos), and then in each case In order to perform filtering (e.g., low-pass filtering) on the I and Q demodulation components, It is possible to generate a
[0069] [Sensing. Sound sensing and other audio (music, speech) played back by the system. , snoring, etc.)]
[0070] Some versions of the present technology may utilize the speakers and / or microphones of the processing device 100. When the microphone is used for purposes other than ultrasonic sensing as described herein, Further processes may be implemented to allow for such simultaneous functionality. For example, for simultaneous audio content generation and ultrasound sensing, The signal bitstream (acoustic sense signal) and the audio signal being played by the speaker as described above. It can be digitally mixed with any other audio content (audible) that is available. For the execution of audible audio content and ultrasonic processing such as One approach is to use other audio content ( This is true for example for mono channels in a multi-channel surround sound system. , which may be stereo channels or more channels) to perform preprocessing It is necessary to remove all spectral content that overlaps with the waveform. For example, In a sequence, for example, frequencies above 18 kHz overlap with 18-20 kHz sensing signals. In this case, low-pass filtering is applied to the music components around 18 kHz. As a second option, overlap detection (direct path and echo) can be performed. -) Adaptive filtering is performed on the music at that time to extract short-term frequency components. You have the option to remove it and leave the music unfiltered. The roach is designed to preserve the fidelity of the music. There is also the option to make no changes whatsoever.
[0071] The audio source on a specific channel (e.g., Dolby Pro Logic, Digital, A Intentionally slowing down the processor (tmos, DTS, etc. or actually virtualization specializer features) If delay is added, any such in-band signals will be treated accordingly and the sensed waveform will be delayed. Note that the echo delay is either not allowed or allows for a delay in echo processing.
[0072] [Sensing. Coexistence with voice assistants]
[0073] Certain implementations of ultrasonic sensing waveforms (e.g., triangular FMCW) have spectral coverage within the audible range. It has the capability to provide voice recognition services (e.g., Google Home). Please note that this may have unintended or unwanted effects on certain voice assistants. Use dual ramp tone pairs or pre-filter the sensed waveform (triangle High-pass or band-pass filtering during chirp) or speech recognition signal By adapting the processing to be robust to the ultrasonic sensing signal component, This avoids the possibility of such crosstalk.
[0074] Consider an FMCW ramp signal y such that:
[0075]
number
[0076] The ramp goes from frequency f_1 to frequency f_2 over time period T. This has sub-harmonics because it is switched over a period of T.
[0077] Analysis of this suggests that out-of-band noise is audible because it occurs at lower frequencies. It is clear that there are harmonics.
[0078] Consider a particular dual ramp pair y as follows:
[0079]
number
[0080] Thus, the sub-harmonics are cancelled (subtracted as above) and the signal is preserved. is highly specific, i.e., using (1 / T) or even -(1 / T) This cancels out the switching effect during the period T. Therefore, the resulting signal This is mathematically simple, so the device (e.g., smart phone) This is advantageous because there is no computational load on the phone device.
[0081] Dual tones switch at DC level ("0"), so for example, click sounds ( i.e., avoiding loudspeaker movement (on and off) To achieve this, natural points can be turned off in the waveform chirp (at the beginning and end of the signal). "0" allows for dereverberation and / or specific transmitter identification (i.e., between each chirp or indeed several It also allows for periods of quiet between groups of short chirps.
[0082] The absence of sub-harmonics also reduces interference sources when two devices are operating simultaneously in a room. This is advantageous because it eliminates the possibility of two different devices overlapping (in frequency). Non-overlapping tone pairs or tone pairs that actually overlap in frequency (and are not overlapping) The addition of a silent period (due to the absence of time-overlapping tone pairs) can be used. In the latter case, the loudspeaker / microphone combination is not available. Limited hearing bandwidth (i.e., sensitivity rolls off significantly to 19 or 20 kHz) This can be advantageous in terms of
[0083] When comparing the relatively inaudible triangular FMCW signal with a dual-tone ramp, the latter Subharmonic levels are much lower (closer to the noise floor on real-world smart devices, e.g. For example, close to the quantum level).
[0084] Dual tone ramps (rather than triangles) can be ramped up or down and banded Since there is no external component, the problem of blurring between lamps that can occur in the case of triangular tilt is eliminated. It will disappear.
[0085] To make a standard ramp audio signal inaudible, the phase and amplitude of the resulting waveform must be adjusted. Extensive filtering is essential due to the potential for distortion.
[0086] [Sensing. Calibration / Indoor Mapping for Optimal Performance]
[0087] The processing devices may be configured with a setup process. During startup (or periodic setup during operation), the device Acoustic probing and seeking to map the presence and / or number of people within If the device is subsequently moved or a degradation in the perceived signal quality is detected, If the system detects a problem, the process can be repeated. Check the capabilities of the microphone(s) and microphone(s) and adjust the equalization parameters. It may also release acoustic training sequences for data estimation. In the case of converters, the ultrasonic frequency used by the system and the temperature and on-state characteristics There may be some nonlinearity in the signal (e.g., a loudspeaker may take several minutes to stabilize). (There is a possibility).
[0088] [Sensing. Beamforming for localization]
[0089] It is possible to implement dedicated beamforming or to use existing beamforming functions. That is, it allows for directional or spatial selectivity of the signals transmitted to and received from the sensor array. Signal processing is used to ensure that the signal is transmitted in a timely manner. This is typically a "far-field" problem, and Wavefronts are relatively low for low frequency ultrasound (as opposed to medical imaging, which is "near field" In a purely CW system, the sound waves travel from the speaker to the maximum area and However, if multiple transducers are available, this It is possible to advantageously control the radiation pattern (an approach known as beamforming). At the receiving end, multiple microphones can also be used. Prioritize operation in any direction (for example, when there are multiple speakers, the sound emitted (manipulation of voice and / or received sound waves), capable of sweeping an area If the user is in bed, it can be directed at an object, or, for example, two people in a bed. When people are present, the senses can be manipulated to be directed towards multiple targets. System manipulation can be implemented on the transmitting or receiving side. The transducer (microphone or speaker) is highly directional (for example, for a small transducer, the wavelength is (similar to the size of the transducer) may limit the area in which such a transducer can operate. .
[0090] [Sensing, Demodulation and Downconversion]
[0091] Returning to FIG. 5, demodulation of the sensed signal can be achieved, for example, by using a multiplier (mixer) module as shown in FIG. This is done by module 7440 or by the demodulator of FIG. 5 to generate a baseband signal. This baseband signal is further processed to detect "presence" (i.e. , disturbances in the demodulated signal related to changes in the received echo (related to the characteristic movements of the person) are effective. Detects whether the received "direct path" (high-quality signal from the speaker to the microphone) When crosstalk is present (e.g. solid-to-air transmission and / or from speakers) Short-distance (signal to microphone), demodulate the received echo signal plus multiply the resulting sum If not, the received echo can be filtered (not acoustically) ) can be multiplied (mixed) with a portion of the original transmitted signal extracted in electronic form. In a particular example, the system may include a multiplication of the received signal and the transmitted signal during demodulation. Instead, the system uses a (damped) The received signals (including the transmitted signal and the received echo(s)) are multiplied by each other as follows: It can be calculated.
[0092] Send = A TX (Cos(P) - Cos(Q))
[0093] Receive = A (Cos (P) - Cos (Q) ) + B (Cos (R) - Cos (S) )
[0094] Self-mixer = [A (Cos (P) - Cos (Q) ) + B (Cos (R) - Cos (S) )] x [A (Cos (P ) - Cos (Q) ) + B (Cos (R) - Cos (S) )], That is, receive x receive.
[0095] The (demodulated) self-mixer components after low-pass filtering can be expressed as This can be done.
number
[0096] The demodulated output of the self-mixer after the simplified equation can be expressed as follows: .
number
[0097] The demodulated components containing the reflected signal information (which may be static or motion related) are: It can be expressed as:
number
[0098] The advantage of this is that all timing information is contained at the receiver, so transmission and reception are synchronized. , and the computation is fast and simple (squaring the array).
[0099] After I and Q (in-phase and quadrature) demodulation, low frequency components related to air turbulence, multipath Reflections (including multipath-related fading) and other slow-moving (typically Select a method to separate out the physiological (non-physiological) information. In some cases, this process involves: This can be called clutter removal. It involves subtracting the DC level (average) or some other trend. D A high-pass filter may be applied to remove the C component and very low frequency (VLF) components. The "removed" information can be processed to remove such DC and VLF data (e.g., Estimate the strength of the influence of airflow or multipath (whether there is a significant influence). Then, filtering The demodulated signal can be sent to a spectrum analysis stage. The unfiltered signal is directly processed by the spectrum analysis processing block without using a filter. There is an option to send the signal to the network and perform DC and VLF estimation at this stage.
[0100] [Example System Architecture]
[0101] Figure 6 shows the voice-enabled sleep improvement system using low-frequency ultrasonic bio-motion sensing. 1 illustrates an exemplary system architecture. The system utilizes the sensing technology described herein ( For example, this can be implemented using multi-tone FMCW acoustic sensing. You can talk to a pre-activated voice-activated speaker to monitor your sleep. For example, verbal instructions can cause a smart speaker to generate a query. This verbal command is received by a microphone. The query is evaluated by a processor to determine the query content or instructions. , determine sleep score, respiratory (SDB) events or audible reporting of sleep statistics Based on this report, the system processes and provides audible notifications for improved sleep. Advice (e.g., recommendations for therapeutic devices to aid sleep) can also be generated. For example, in response to a query received by the microphone, such as "How was my sleep last night?" In response, sleep-related data (e.g., sleep score, time in sleep stages) shown in FIG. a speaker (singular or plural) of the processing device 100 to provide a verbal summary and audible explanation to the user of the processing device 100; In response, the processing device 100 may provide an output (e.g., processing data (e.g., sound / audio) detected by the device (e.g., sonar sensing) For example, in response to this, the processing The device may generate an audio report with the following explanation: "Your total sleep time The duration of sleep was 6 hours, with only two awakenings during sleep. Of these, deep sleep lasted for 4 hours. There were 2 hours of light sleep. Apnea and hypopnea counts were 2 events. "There were some incidents."
[0102] Speaker-enabled processing devices enabled by the low-frequency ultrasonic sensing of the present technology The system process for detecting motion in the vicinity of the chair 100 is illustrated in the exemplary module shown in FIG. The processing device 7102 may be considered in relation to a speaker (single or multiple speakers). 7302 and one or more microphones 7310 and optionally microphone(s) 7303. and a microcontroller 7401 equipped with the above programmable processor. The modules can be programmed into the memory of the microcontroller. Regarding the point, the audio sample or audio content is It can be upsampled by an optional upsampling processing module, as if Whether any audio content is generated by the speaker (simultaneously with the sensed signal) In this regard, the adder module 7420 is a signal processing system for combining audio content with an FMCW signal (e.g., a signal in the desired low ultrasonic frequency range). FMCW process module 74430 to generate dual-tone FMCW signals within the FMCW range Optionally combine with an FMCW signal within the desired frequency range from The FMCW signal is then fed to a converter module for output from, for example, a speaker 7310. This FMCW signal may be processed by a demodulator such as a multiplier module 7440. In a demodulator such as multiplier module 7440, the FMCW signal is The received echo signal observed at the microphone 7302 and the processed (e.g., mixed / multiplied) Prior to such mixing, the received echo signals are subjected to the filtering process as described herein. Adaptive filtering removes unwanted frequencies outside the frequency spectrum of interest The audio output processing module(s) 7444 may Downsampling the filtered output and / or converting the signal to audio Next, the multiplier module 7440 may generate a multiplication signal. The demodulated signal output from these can be further processed, for example, by a post-processing module 7450. For example, this output can be processed by frequency processing (e.g., FFT) and digital signal processing. By analyzing the movement of the heart, it is possible to detect (a) respiratory movements or movements, (b) cardiac movements or movements, and (c) ) detecting a whole body movement or motion (e.g., whole body movement or whole body motion) to isolate the whole body movement or motion The signal of the detected or otherwise individual movements is divided into frequency ranges. Next, various information outputs such as those mentioned above (e.g., sleep, sleep To characterize various motions in the signal to detect stage, motion, and respiratory events The 7460 characteristic processing allows recording of physiological movement signal(s). or otherwise processed, for example digitally.
[0103] In relation to detecting whole body movements or whole body actions, such movements may include arm movements, head movements, etc. The movements may include any of torso movements, limb movements, and / or whole body movements. Such a method of detection from transmitted and reflected signals for motion detection is called sonar-sound type This can be applied to motion detection, e.g., International Patent Application PCT / EP2016 / 05888 6 and / or PCT / EP2016 / 080267. The entire document is incorporated herein by reference. By their very nature, such RF or sonar technology captures all body movements (including (or at least most of it) can be seen at once, and the "beam" can be directed For example, this technique primarily irradiates the head and chest or the whole body. If leg movements are, for example, periodic, the leg movements may be characterized as certain movements based on the frequency of the movements. and optionally by performing a separate automatic gain control (AGC) operation. Respiration detection is most effective when there is minimal whole-body movement, and Separating the frequency and shape of respiratory waveform signals (typical COPD or (including CHF rate of change and inspiratory / expiratory ratio, SDB events, and longer-term SDB modulation).
[0104] If the movement is associated with a person in bed, the largest amplitude signal is associated with whole body movement (e.g., sleeping). Hand or leg movements may be faster (e.g., I / Q The relative amplitude is lower than that of the velocity from the motion signal analysis. Different components and / or component sequences may be used, e.g., for whole body movements and The acceleration and speed of arm movement can be taken into account in determining whether it started and then stopped. This identification can be targeted to different movement gestures.
[0105] [Sensing. Coexistence with other audio digital signal processing (DSP)]
[0106] The sensing frequency used by the acoustic signal generated by the speaker of the processing device 100 For the number(s) and waveform shape(s), the device or associated Suppress any existing echo cancellation in the attached hardware or software (e.g. (If necessary, it can be disabled.) Reflected " Automatic gain control (AGC) is used to minimize "echo" signal disturbances (unintended signal processing). C) and noise suppression can also be disabled.
[0107] The resulting received signal (e.g., received at a speaker) is digitally systematically bandpass filtering each intended transmit and sense waveform at different frequencies. other signals in the vicinity (e.g., speech, background noise, co-located or is to distinguish between different sensing signals carried out in co-located devices.
[0108] [Sensing. Multi-mode / hybrid sensing]
[0109] Continuous wave (CW) systems rapidly detect "anything" (i.e., general movement in a room) However, this is insufficient for high-precision distance gating and As an improvement measure, multi-tone to prevent fading is being used. CW is available. Range gating can be used to localize the motion. Therefore, FMCW, UWB or some other modulation scheme (e.g., FSK or PS It is desirable to use K. In the case of FMCW, there are no strong nulls like in CW. It supports distance gating and tolerates indoor mode accumulation.
[0110] In other words, a wide range of waveforms, such as dual or multi-tone continuous wave (CW), can be used. It can detect all motion within an area such as a room. Multipath reflection and / or or to minimize nulls due to standing or traveling waves generated by reverberation. The advantage of this approach is that it is possible to detect any motion. This allows for the use of a larger signal that fills the space. It can act as a detector and can function as an intruder detector. If a condition is detected, the system identifies possible candidate physiological signals (e.g., After searching for a typical activity sequence (user walking into a room, breathing, heart rate, and For example, we search for characteristic movements such as gestures. Our system does not directly provide range information. From CW-type systems, frequency modulated continuous wave (FMCW) or ultra-wideband (UWB) signals It switches to a system that can detect and track movements within a specific range, such as obtain.
[0111] UWB systems can be audible or inaudible depending on the frequency response of the speaker and microphone. This means that these components can support higher frequencies. In the case of more typical consumer speakers, wideband signals may be outside the range of human hearing. , UWB speech is more likely to be audible than filtered white noise. sounds (e.g., pink noise or any sound that is not "unpleasant" to the human ear) This can be done by using a device such as a white noise filter during sleep. In other cases, UWB may be able to mimic the noise generator. Another option for providing distance-based sensing when away from home for security applications This becomes:
[0112] In some versions, sensing is performed using multiple sensing devices (e.g., any two One or more types of sensing devices (e.g., acoustic sensing devices, RF sensing devices, and IR sensing devices) For example, the processing device may be configured to perform RF sensing and sound sensing. The processing device may detect motion by IR sensing and Acoustic sensing (e.g., FMCW) may be used to detect motion. The processing device may also use IR sensing and The processing device may detect motion through IR sensing, RF sensing and acoustic sensing. Sensing (eg, FMCW) may be used to detect motion.
[0113] [Sensing. Physiological signals]
[0114] After separation of DC and VLF (e.g., airflow), breathing signals, heart rate signals, and whole-body movements The signals are separated by searching for bins in the FFT window and tracking over the window. traces and / or direct peaks / troughs of the time domain signal at specified distances or Zero-crossing analysis (e.g., a specified distance range extracted using complex FFT analysis of the demodulated signal) The result is that a range of user motion choices can be estimated via the time domain signal. This is also known as "2D" (two-dimensional) processing, such as FFT.
[0115] [Biometric sensing using smart speakers]
[0116] In the home environment, so-called "smart" devices (e.g. "smart speakers") is rapidly expanding and provides opportunities for novel physiological sensing services. .
[0117] A smart speaker or similar device is typically installed in a home Automation (e.g., smart appliances, smart lighting, smart thermostats) or smart power switches for powering appliances) and networks (e.g., internet to and from other connected home devices for the Internet Wireless means (e.g., Bluetooth, Wi-Fi, ZigBee, mesh, It includes communication via peer-to-peer networking. It is designed to simply output an acoustic signal. Unlike standard speakers, smart speakers typically have no other electronics besides the processing. Also includes one or more speakers and one or more microphones. (a) a speaker and (b) a processor for providing personalized voice Intelligent assistants (artificial intelligence (AI) systems) to enable control Some examples include Google H There's the HomePod, the Apple HomePod, the Amazon Echo, and you can just say, "OK, Google it." Using catchphrases like "Hey Siri," "Hey Siri," or "Alexa" These devices and connected sensors are voice activated. It can be considered part of the Internet of Things (IoT).
[0118] The ultrasonic detection techniques described above can be applied to smart devices (using audible or inaudible acoustic signals). When using it in a smart speaker, (estimated and and / or updated based on the actual performance of a particular device, Specific optimization is required. Broadly speaking, the speaker(s) and microphone The maximum frequency supported by both the phone(s) allows for the highest available A high frequency inaudible sense signal is ultimately defined, which is based on the specific device manufacturing tolerances. This may vary slightly depending on the
[0119] For example, the speaker in the first device (e.g., a Google Home device) , in a second device (e.g., a Samsung Galaxy S5 smartphone) The primary device speaker may have different characteristics than the primary speaker. The primary device speaker has a sensitivity up to 24 kHz. It is possible to have a microphone response that is nearly flat for similar frequencies. The speaker of this device may roll off at lower frequencies, resulting in peak sensitivity and The trough can exceed 18kHz. For Amazon Alexa device speakers , for example, may roll off at 20 kHz. For an exemplary Google device, A passive reflex speaker design can be used for returning phase-inverted waves, and these waves The sound characteristics of the speaker change because it can send the sound to the side (for example, a "10kHz" speaker (When the speaker is turned on, it changes to a 25kHz speaker.)
[0120] Some devices use microphone arrays (e.g., multiple microphones on a flat plate). ) and may implement gain and averaging functions, but Such operations are numerically (i.e., This can be done (i.e., in the digital domain using digital signal processing).
[0121] One possible difference that needs to be managed for such a system is the speaker (singular The orientation of the microphone(s) and microphone(s) is relevant. In this embodiment, the speaker may face forward and point in the same direction towards the user. However, the microphone(s) should be positioned at an angle of 20 to 30 degrees relative to the room or ceiling. It may be oriented in degrees.
[0122] Therefore, to optimize sound pickup for distance acoustic sensing, the application The configuration of the room is based on learning the topology of the room and the possible reflection paths. A structure that generates a probing sequence via a speaker for enabling Therefore, the setup process allows for Multiple people or motion sources when the sources are at different distances from the smart speaker to calibrate distance measurements of low-frequency ultrasonic echoes to aid in the simultaneous monitoring of the source From a signal processing point of view, the reverberation floor (reflection The time it takes for the energy of the incident acoustic wave to dissipate varies with different sensing signals. This is 2-3 times faster for CW (continuous wave) and about 5 times faster for FMCW (i.e., (i.e., for FMCW, the frequency range, duration, and repetition There may still be reverb fading depending on the shape of the sequence).
[0123] Separating microphones can lead to difficulties, for example, microphones in a common housing. If a microphone array is available on the smart speaker, An exemplary separation distance may be 71 mm. If the wavelength is 20 mm, this is One microphone may be in the trough while the other is in the peak region. (For example, as the user moves towards the fixed speaker, the SNR between microphones decreases.) (The SNR of the microphones changes depending on the location.) The preferred structure is to place two microphones at specific locations within the area. It can be configured with an audio sensing wavelength-related interval of 19~20mm. If such distances are unknown in a given system, the calibration process can be performed e.g. This distance can be detected as part of the setup process. The calibration process involves generating time-synchronized calibration sounds through one or more speakers to calibrate each speaker. The time of flight from the speaker to each microphone can be calculated or estimated, The difference between smartphones can be estimated based on these calculations. When detecting the distance between two or more microphones, the distance between the microphones is taken into consideration. can be put into
[0124] Other devices, such as active soundbars (i.e., containing a microphone), and and mobile smart devices may also be implemented with the sensing operations described herein.
[0125] At least one speaker and at least one microphone (or By providing a transducer (which can be configured to perform the biometric sensing function), Intelligence is being applied to these devices using active low-frequency ultrasound and its echo processing. As mentioned above, it is possible to, for example, Reproduce acoustic signals (above 8 kHz) and within known or determined system capabilities This can be realized by transmitting (for example, if the sampling rate is 4 If it is 8kHz, it can be below 24kHz, but is usually below 25 or 30kHz In contrast, medical ultrasound is typically performed at much higher frequencies (e.g., 1–18 MHz). ) and require special equipment for these operations. Smart speaker systems are already available in most homes thanks to technology. (including smartphones) without the need to purchase any expensive equipment. Non-contact measurement becomes possible.
[0126] [Coexistence of different sensing devices / applications]
[0127] Coded or uncoded ultrasound signals are used by devices and systems to identify and other data exchange purposes. For example, a mobile phone application can detect other nearby sensors. Identifying itself to devices / systems (e.g., smart speakers) where it is more useful configured to cause such signals to be generated for communication purposes, to These types of signals can be used instead of short-range radio frequency communications for identification. (e.g., if Bluetooth is unavailable or disabled) The devices in the system may detect the presence of other processing devices in their perceptible vicinity (e.g. automatically determine whether the The parameters of the generated sensing signal can be determined in a non-interfering sensing mode (e.g. For example, by using different frequency bands and / or frequency bands that are non-overlapping in time. These parameters can be adjusted to allow for:
[0128] [Low frequency ultrasonic (sonar) detection]
[0129] In many places, they emit sounds in the low-frequency ultrasonic range, just above the human hearing threshold. audio devices capable of transmitting and recording (e.g., infotainment Such devices and systems use low-frequency ultrasonic technology to detect nearby It can be adapted to perform physiological sensing of a human being. Such sensing can be performed in a standard audio system. This can be done without affecting the original intended function of the system. Such a sensing function can be realized through software updates (i.e. (i.e., it can provide additional useful functionality without increasing the cost of goods.) In some cases, one or more of the transducers in a new device or system may be It can be specified to support audio frequency range for frequency ultrasonic detection, and Further tests are performed during production to ensure that this specification is met.
[0130] Such acoustic sensing technologies (audible or non-audible) can be used for a wide variety of purposes. (e.g., proactive health management, medical devices, and security) function).
[0131] Low-frequency ultrasound systems, operating up to approximately 25 kHz, can be used with mobile smart devices or This can be realized on smart speaker devices. is transmitted to one or more subjects by one or more transducers on the electronic device, the transducers To generate sound energy over a range of frequencies, including frequencies below kHz The speakers are smartphones, smart speakers, sound bars, portable devices, a television screen, or many other devices containing transducers capable of assisting in low frequency ultrasonic sensing and processing. The speaker system may be configured to control a computer. When implemented in this way, a smart speaker system is effectively created.
[0132] Audible sounds (e.g., breathing, coughing, snoring during sleep, shortness of breath, wheezing, speaking, sniffling, These sounds are then extracted from the reflected sensor signals (detected for motion detection). can be extracted and classified from nearby sensed audio signals (to separate Some of these sounds (e.g., coughing) can cause a loss of the sensed signal (especially very low When operating at audio pressure levels, the signal may be masked, which is undesirable. However, such sounds may be detectable, and so it is important to distinguish them from other environmental sounds. (e.g. honking, motor noise, street noise, wind, slamming doors) Respiratory sounds typically have poor signal quality in quiet environments. Active sensing approaches (e.g., sonar or radar) are becoming more common. (primarily detecting body and limb movements), or camera / infrared systems ) is a good second estimate of inhalation / exhalation time (and therefore respiratory rate) In other words, the system again extracts information about the characteristics of the voice. In the case of extremely loud noises, when the quality of the relevant signal falls below an acceptable threshold, The system can skip small portions of the sensed signal.
[0133] In sonar systems, air movement due to inhalation or exhalation (detected signals are reverberated) of the acoustic mode setup in the sensing environment if it continues long enough to experience It is also possible to detect it by a method of tracking the moving wavefronts that are obtained (due to external disturbances). Detecting snoring directly from the audible signature is a relatively loud process. This detection is easier because it can be done using, for example, the average maximum decibel level. Classify snoring as mild (40-50db), moderate (50-60db) or severe (>60db) This is done by classifying the
[0134] Thus, in some cases, the processing device 100 may use motion detection for respiration detection. (e.g., sonar) technology may be used. However, in some cases, microphones The acoustic analysis of the audible respiratory signal in the It is possible.
[0135] [RF (radar) detection]
[0136] Some systems use a single port for simple internal motion detection for security purposes. These may include a Doppler radar module (updated software) The motion detection can be localized to a specific area in the vicinity. capable of localizing (especially capable of detecting and distinguishing between individual seats / occupants) The module may be replaced. Improvements in the sensor technology (e.g., ultra-wideband (UWB) sensing) signal or frequency modulated continuous wave (FMCW) sensing signal or other coding scheme (e.g., OFDM, PSK, FSK) in the generated sensing signal. These can be achieved by sensors with high accuracy ranging capabilities (1 cm or less). Such sensors may sense within a defined area (e.g., Antennas that can be configured in proximity to have a sensing direction directed toward a particular sheet. In some cases, multiple antennas may be used for a particular sensing and a distance sensing differential associated with the different antennas. It can be used with beamforming technology. Multiple sensors can detect whether a person (or pet) is inside. It is used in an area that covers multiple areas that can be used (e.g., a sensor for each seat). It is possible.
[0137] [Multi-mode data processing]
[0138] Sonar, RF or infrared sensing (i.e., infrared emitter for IR transmission and reception) When using a sensor and a detector, the processing device 100 may be affected by nearby equipment (e.g., Further data or signals generated (for example, for the purpose of estimating the degree of accuracy) may be received, which may result in This makes it possible to perform bio-motion detection based on data from such devices. For example, a seat / bed load sensor that detects whether a person is sitting on a given seat or in a bed. The sensor performs bio-motion sensing for sensing that may be associated with a particular seat or bed. Providing information to the biomotion processing device 100 for determining the timing to start Infrared systems can be used with camera systems that can track the movements of the human eye, for example. The system may optionally be equipped for drowsiness detection, for example.
[0139] The processing device may also include a processor for processing distance information for evaluation of the relevant range / distance for detection of bio-motion characteristics. For example, the processing device 100 may be configured with information about the interior of a neighborhood (e.g., a room). Such a map can be used in the design stage. Optionally, the processing device may be initially provided to specify an initial sensing configuration. The sensing system under control dynamically updates the map when used by one or more people. The initial configuration can be updated (or detected) based on, for example, the seat position and the most likely The seat configuration can be captured / detected, i.e. if the seat is movable, the sensor ,can report the current settings to the system and update ,sensing parameters (e.g., The position of the seated person may change (when the seat is slid back or forward or folded down) (which may move relative to the sensing loudspeaker if detected).
[0140] [Biometric feature detection: respiration, heart, movement and distance]
[0141] [Sensor signal processing]
[0142] The system includes a specific processing device 100, for example, where demodulation is optionally performed by the processing device. If the demodulated signal is not provided by the sensor (e.g., sonar, RF / resonator), The processing device 100 may then receive the target component. signals (e.g., direct current signals DC and very low frequency VLF (e.g., airflow), breathing, heart rate, This signal can be processed by separating the signal from the whole body motion signal. Searching for bins in a Fast Fourier Transform (FFT) window and traversing the window / or via direct peak / trough or zero-crossing analysis of the time-domain signal at specified distances tracking (e.g., over a specified distance range extracted using complex FFT analysis of the demodulated signal) This can be done by using a time domain signal (a "time domain" signal). Since the FFT as described in / EP2017 / 073613 is performed, "2D" Also called (two-dimensional) processing.
[0143] In the case of sonar detection, significant other information can be found in the audio band, Such information can be picked up by the infotainment system. audio (music, radio, TV, movies), telephone or videophone ringtones ( including human speech), ambient noise, and other internal and external sounds (e.g., Most of these audio components are interference components. can be considered as an object and suppressed from the biometric parameter estimation (e.g., a filter (May be ringed).
[0144] In the case of radar sensing, signal components from other RF sources may be suppressed.
[0145] In the case of infrared sensing (e.g., when physiological sensing is used in addition to eye tracking), temperature changes and interference may occur due to the position of the sun, which can be taken into account. A temperature sensor (e.g., from a thermostat temperature sensor) and time evaluation of the sensed signal This can be done in the processing of the signal.
[0146] Regardless of the specific sensing technology used (RF, IR, sonar), the received time domain reflection The reflected signal can be further processed (e.g., filtered by a bandpass filter). Processing by end-pass filtering, evaluation by an envelope detector, and subsequent peak detection Envelope detection is performed using the Hilbert transform or respiration data. Squaring the data, sending the squared data through a low-pass filter, and calculating the square root of the resulting signal In some examples, the respiration data may be calculated using peak and trough respiration data. It can be normalized and sent through a rough detection (or alternatively zero crossing) process. The detection process allows separation of the inhalation and exhalation sites, and in some cases and calibrating the detection process to detect the user's inhalation and exhalation sites in this case. It is possible.
[0147] Respiratory activity is typically between 0.1 and 0.7 Hz (e.g., from regular, deep breathing) Breaths / min to 42 breaths / min (typically a fast adult breathing rate) Cardiac activity is reflected in the signal at higher frequencies, with a passband range of 0.7-4 Hz ( Filtering by a bandpass filter (48 beats per minute to 240 beats per minute) This allows access to this activity. Activities resulting from whole body movement are typically Hz to 10 Hz. Note that there may be overlap in these ranges. A clear respiratory trace can lead to strong harmonics, confusing respiratory overtones with cardiac signals. Tracking is necessary to prevent this. The greater the distance from the transducer (e.g., several meters), the The relatively small cardiac mechanical signals can be extremely difficult to detect, and The estimation is performed within 1 meter of the smart speaker (e.g., on a chair / couch). It is more suited to quiet lying down settings (or in bed).
[0148] After absence / presence is determined as "presence," breathing estimation, cardiac estimation, and movement / activity are performed. moving signals (and their relative position and velocity (if moving, e.g., in or out of proximity) The estimation of (if moving in)) for one or more people in the sensor's field A system that provides ranging information is designed to measure the resting respiratory rate of multiple individuals. Even if the two people are in the same household (which is often the case with young couples), the biometrics of multiple people It can be seen that the block data can be separated.
[0149] Based on these parameters, various statistical measures (e.g., mean, median, third moment) can be calculated. After preparing the waveform (morphological treatment), the can be fed into a validation system (e.g., a simple classification or logistic regression or more complex machine learning systems using neural networks or artificial intelligence systems. The purpose of this processing is to obtain further information from the collected biometric data. The purpose is to gain insight.
[0150] [Sleep staging analysis]
[0151] Absence / Presence / Awake / (NREM) Sleep Stage 1 / Sleep Stage 2 / Sleep Stage 3 ( Slow wave sleep (SWS) / deep sleep / REM are the fundamental sleep architecture that defines the sleep cycle. Since there is a sequence related to this, we will treat this as a sequence problem rather than a non-sequence problem. It can be useful to view it as a cognitive problem (i.e., how often a person is in one state over a period of time). (Reflecting typical sleep cycles in which the individual remains asleep throughout the night). This provides a clear sequence of observations ("sleep").
[0152] On some systems, the ripeness becomes deeper (more pronounced) as the evening progresses. The "correct" sleep pattern is that the sleep cycle progresses from sleep to sleep (SWS) and then to deeper REM sleep as the night progresses. Knowledge about "normal" sleep patterns can also be utilized. This prior knowledge can be used to Weighting the classification system for sleepers (e.g., prior probabilities of these states over time) Although these assumptions are available for the purpose of adjusting for the population norms, Abnormal sleepers or those who have a habit of taking daytime naps or have poor sleep hygiene People with poor sleep habits (e.g., large variations in "going to bed" and "wake-up" times) Note that this may not be the case.
[0153] Traditionally, sleep stages have been defined in the literature [Rechtschaffen & Kales guideline In (Rechtschaffen and Kales, 1968) (a manu al of standardized terminology, techniqu es and scoring system for sleep stages o f human subjects. U.S. Public Health Service, U.S. Government Printing Office, Washington. DC1968)] have been studied in 30-second "epochs." When you look at G, the paper speed is 10mm / s (1 page is equal to 30 seconds), so Alpha and The paper states that 30 seconds is the ideal interval for visually checking the spindle. Of course, the actual physiological processes of sleep and wakefulness (and absence / presence) The time is not divided evenly into 30-second blocks, so longer or shorter times may be used. The system outlined here uses a 1 second (1 Hz) sleep state. The page output is used preferentially, but longer data blocks are used in an overlapping manner (basic Updates every 1 second (1 Hz) with an associated delay related to the size of the processing block This 1-second output allows for more precise detection of subtle changes / transitions in the sleep cycle. It is used to indicate.
[0154] [Manual vs. automatic generation of sleep features]
[0155] The sensed signal (which represents distance versus time (motion) information) can be used to detect various characteristics (e.g., sleepiness, These features are then used to calculate the user's physiological Information about the state can be derived.
[0156] For feature generation, several approaches can be implemented. For example, a human expert may Considering respiratory and other physiological data and their distributions based on personal experience; By understanding the physiological basis of certain changes and using trial and error, Features can be manually generated from processed or raw signals. Alternatively, machines can be made to "learn" with some human supervision ("machine learning"). core concepts in the field), labeled data is provided along with expected results. This can be done with some human assistance or in a fully automated manner. In the automatic case, you may provide some or none of the labeled data. There are also.
[0157] Deep learning can be broadly considered in the following broad categories: In other words, deep neural networks (DNNs), convolutional neural networks Neural Networks (CNNs), Recurrent Neural Networks (RNNs) and other types. Deep Belief Networks (DBN), Multilayer Perceptrons (MLP) and One can consider a stacked autoencoder (SAE).
[0158] Deep belief networks (DBNs) are used to automatically generate features from input data. Another approach to this end is fuzzy There is C-means clustering (FCM). Fuzzy C-means clustering (FCM) is It is a form of unsupervised learning that helps discover inherent structures in pre-processed data.
[0159] Applying digital signal processing techniques to sensed motion data to form hand-crafted features. In the ideal case, the respiratory signal is described as inspiration and expiration occur. The complete set of two amplitudes (deep or shallow) and constant frequency (constant breathing rate) In the real world, the respiratory signal may be far from a sinusoidal shape. (especially when detected from the torso area via acoustic or radio frequency sensing approaches) For example, inspiration may be more rapid and faster than expiration, causing breathing to momentarily stop. In this case, notches may appear on the waveform. The inspiratory and expiratory amplitudes and the respiratory frequency vary. For some extraction methods, after peak and trough detection, the Higher quality detection (e.g., detecting local peaks and discarding troughs) This method focuses on the inspiratory and expiratory times and volumes (e.g., time Calculated by integrating the area signal with a calculated reference baseline To estimate both the peak and trough times, the respiratory rate is This may be sufficient for estimation.
[0160] Estimation of any of these characteristics (e.g., respiratory and / or heart rate or amplitude) A variety of methods can be used to assist in the determination of
[0161] For example, peak and trough candidate signals can be extracted by recovering the respiratory waveform from noise. (There can be a variety of out-of-band and in-band noise, usually lower frequency noise) noise can dominate, complicating accurate detection of lower respiratory rates (e.g., 4 ~8 breaths / min (this is rare in spontaneous breathing, but may occur if the user tries to breathe more slowly). This can occur if a low-pass filter is required. Maximum and minimum detection after sampling and adjustment across multiple respiratory blocks By using an adaptive threshold (which allows for the detection of deep and shallow breathing) Optionally, the signal is subjected to low pass filtering and differentiation (e.g. , derivative) can be taken. Then, find the peak in the differentiated signal associated with the maximum rate of change. In this way, a constant noise can be used to obtain an indication of a respiratory event. The reference points of the respiratory waveform, which is modeled as a sinusoidal waveform containing the High frequency noise is removed. Then, differentiation is performed and peaks are detected. This allows the maximum rate of change points of the original signal (rather than the peaks and troughs of the original signal) to be detected. This is because the respiratory waveform is most clearly defined (e.g., in the broad peaks). This is because the maximum rate of change (e.g., when breathing is stopped for a short period of time) is observed. A potentially more robust method is to use a fixed baseline or There is a way to detect zero crossings (around an adaptive baseline). This is because field crossings are not directly affected by local variations in signal amplitude.
[0162] The respiratory signal varies over time (depending on the distance and angle of the chest from the sensor(s)). While easily visible in the regional signal, cardiac motion is significantly different from respiration. Higher respiratory harmonics (e.g., waveform related) are often very small signals when In the presence of respiratory harmonics, cardiac signal extraction can be complicated by rejecting or It needs to be detected and eliminated.
[0163] Frequency domain methods may be applied to, for example, respiratory data. A block of data (e.g., a data stream that is repeatedly shifted by 1 second) 30 blocks of data) or non-overlapping (e.g., data stream is 30 (combat spectral leakage) using This may include using detected peaks within an FFT band (which may be windowed for filtering). Power spectral density PS using Welch's method or parametric model (autoregression) D may be used, followed by a peak search. As the sinusoidal waveform of the respiratory signal weakens, the The spectral peaks tend to become larger (broader), and the shape tends to have sharp peaks and sharp transitions. Alternatively, the signal may contain harmonics if a ripple or notch occurs. One method is to use autocorrelation (which describes the similarity between the original and the new version of the original). The method assumes that the underlying respiratory waveform is relatively stable over a certain period of time. Tracking and filtering periodic local maxima in the autocorrelation for respiratory rate estimation. Filtering is done with local maxima (e.g., noise-free) as the most likely candidates. Autocorrelation can be performed in the time domain or by FFT in the frequency domain. Time-frequency approaches (e.g., wavelets) are also useful and offer powerful Noise removal can be performed, as well as peak detection (all at once) on the time scale of interest. The final sinusoidal waveform (i.e., within the target respiratory rate range) is A suitable solution is selected (e.g., Simlet, Daubechies).
[0164] Applying a Kalman filter (a recursive algorithm) to the time domain signal to calculate the system state Using this approach, the It provides a way to predict unknown future system states. In addition to filtering, signal decomposition This allows for separation of movements (e.g., large movements, breathing, and cardiac movements).
[0165] [Noise contamination (e.g., for the detection of physiological movements in noisy environments)] (inspection)]
[0166] If the subject stops breathing (e.g., apnea) or exhibits very shallow breathing (e.g., For example, when detecting hypopnea, peaks and troughs of respiration, e.g., when the subject is large, If you make any sudden movements (for example, rolling around in bed or moving while driving), , it is necessary to be aware of the possibility of potential confounding effects. Location tracking is possible The use of novel sensing methods provides a useful means of separating these effects. The movement can be seen as both a high frequency movement and a change of position in space. Subsequent breathing amplitude may be higher or lower, but "healthy" breathing is maintained. In other words, the detected changes in amplitude are not due to changes in the person's breathing (dynamic This may be due to changes in the extracted received respiratory signal strength (e.g., after unconversion). This allows for a novel calibration approach, where the detected distance is used to calculate the signal strength and breathing rate. It will be appreciated that the depth (and thus approximate tidal volume) can be correlated. If no such movement or displacement is observed (e.g., chest and abdominal during an obstruction event), Any decrease, cessation, or change in the range of a specified duration (due to inconsistent activity on the Abnormalities in breathing (e.g., apnea-hypopnea events) may be identified.
[0167] A practical and robust cardiorespiratory estimation system can be based on parameter localization. It is understood that there are multiple ways to achieve this. If the signal quality is high, In estimating the frequency of breathing between the stimuli, the respiratory rate is a likely estimate of local respiratory variability. The localized waveform can then be used to extract the fine peak and trough times, and the inspiratory and expiratory volumes can be calculated. Range calibration is performed for the estimation of volume (a useful feature for sleep stages). Such signal quality metrics are expected to change over time. In some cases, processing can be performed over different time scales (e.g., 30 seconds, 6 Filtering averages or medians over 0, 90, 120, and 150 seconds ).
[0168] In the case of sonar, for example (to provide further information for RF sensing systems) Other additional sensors for respiratory rate estimation (for using sonar for respiratory rate estimation and vice versa) When a known signal is realized, the envelope of the raw received waveform (e.g., the envelope of an acoustic FMCW signal) envelope) can be treated as the primary or secondary input. This is based on the detection of real disturbances in the air of a person's exhaled breath (e.g., an open window, a nearby Ensure that there are no other strong air currents (from adjacent air conditioning units, nearby heaters) in the cabin, room or nearby. To show that there is no airflow, if there is such an airflow, the effect of the airflow on the measurement should be discarded or Alternatively, such effects can be used to detect changes in airflow within the environment.
[0169] When low-frequency motion crosses an area (i.e., when a perturbation flows across an area) , tend to detect large airflows. When detecting waveforms with a lot of reverberation (e.g., Energy at one frequency accumulates in the room and in associated room modes This becomes clearer if
[0170] The general population (i.e., users with normal health status), users with diverse health status (e.g., sleep apnea, COPD, and heart problems) When considering a sleep stage system, it is possible to significantly change the baseline of breathing rate and heart rate. It is understood that it is possible to determine the age, sex, and body mass index (BMI). Consider the difference in respiratory rate (BMI) between men and women of similar age and BMI. The baseline number may be slightly higher (although a recent study of children aged 4-16 years (If the BMI is high, there is no statistical difference.) Children tend to breathe faster than adults. A normal breathing rate is much higher in children than in adults.
[0171] Therefore, in some versions, regardless of the type of sensor, e.g., the processing device The system used with 100 is configured in a hybrid implementation. obtain (e.g., initial signal processing and some hand-crafted features) After forming the model, we apply a deep belief network (DBN). The basic structure is derived from digital signal processing (DSP) by hand. Initial supervised training (which requires a mix of machine-learned features and Training is provided by expert trainers from multiple locations worldwide, either in a sleep study lab or at home with PSG. Scoring was performed using a core polysomnography (PSG) nighttime dataset. The scoring shall be conducted by at least one scorer using a designated scoring method without further supervision. The training is based on the data collected by one or more of the selected sensing methods. As a result, the data will reflect new and more diverse data outside of the sleep laboratory. The system can be developed.
[0172] Handcrafted features (i.e., designed, selected, and For each of these (i.e., the respiratory signal is extracted along with the associated signal quality level) and a specific pair of Elephants are characterized by variability in respiratory rate over different time scales and inspiratory and expiratory times. A baseline estimate of the individual's breathing rate for wakefulness and sleep is formed. For example, short-term changes in respiratory rate variability during wakefulness may be associated with mood and mood changes. These changes during sleep may be related to changes in sleep stages, whereas the changes in sleep may be related to changes in sleep stages. For example, respiratory rate variability increases during REM sleep. Longer-term changes in behavior may be related to changes in mental state (e.g., mental health These effects are particularly relevant when compared and over longer timescales. and when compared to population norms, this may have a more significant impact on the user's sleep.
[0173] The measured respiration rate variability can be correlated with the user's state (asleep / awake) or sleep stage (RE) M, N1, then N2, then SWS (lowest point of sleep). For example, normalized respiration over a period of time (e.g., 15 minutes) in a normal healthy person When looking at variability in numbers, we can see that variability is greatest during wakefulness. This variability decreases across all sleep states, with the second highest in REM sleep. (becomes lower than during wakefulness), then decreases further in the order of N1 and N2, resulting in SWS sleep. As an aside, the air pressure caused by breathing increases during REM sleep. Such an increase can have an effect on the detected acoustic signal, and may be more pronounced in quieter or quieter environments. This can be an extraneous feature that can be detected in quiet time.
[0174] These normalized respiratory rate values are obtained for healthy individuals in different positions (supine, prone, etc.). However, the relative position of the correct tidal volume should not change significantly between the two positions. It should be understood that it may be desirable to perform a calibration. The average breathing rate during sleep may be, for example, 13.2 breaths per minute (BR / MIN), while The average person may be 17.5BR / MIN, so the system can be normalized overnight. Both rates show similar variability in sleep stages. The rate difference is a function of sleep state variability. This system only masks possible changes in the class. , comparisons over time or indeed with those of similar demographics) Obstructive Sleep Apnea (OSA) In humans, respiratory variability increases in the supine position (lying on one's back), It is expected that this may be useful as an indicator of respiratory health.
[0175] In subjects with mixed apnea or central apnea, respiratory fluctuations during wakefulness The sensitivity tends to be greater (than in normal subjects (which is a useful biomarker)). Subjects with obstructive apnea also have changes relative to normal values when awake, but these changes is less clear (but is present in many cases).
[0176] A person's specific sleep patterns (e.g., breathing variability) are learned by the system over time. Therefore, systems capable of unsupervised learning have been developed in this field. If so, that would be highly desirable.
[0177] These patterns occur when breathing stops partially or completely (or when airway closure occurs). During the night (i.e., sleep) This can change (during a session) and can be affected by sleep apnea. One way to address the issue is to consider detected apneas (and It turns out that there are ways to suppress the period of breathing (and associated fluctuations in breathing rate): Instead of attempting to classify the sleep stages at that time, the system analyzes the sleep patterns of apneas and micro-arousals. Possible periodic breathing patterns (e.g., Cheyne-Stokes) can be flagged. In the case of CSR (Coronavirus Respiratory Syndrome), strong patterns of fluctuation emerge. These are associated with the pre-sleep processing stage. CSR can occur during any sleep stage, but it can also be detected during non-relapse. Pauses become more regular in REM sleep and more irregular in REM sleep (C information that the system can use to refine the sleep stages of the SR subject).
[0178] Similarly, a process that suppresses all harmonics related to the morphology of the respiratory waveform. This step allows the cardiac signal to be extracted. apnea (synchronized or central apnea) along with any associated recovery breaths and breathlessness-related movements The cardiac signal is used to calculate beat-to-beat "heart rate variability" based on physiologically relevant heart rate values. The human heart rate (HRV) signal is estimated. Spectral HRV metrics can be calculated (e.g., mean Log output of average respiratory frequency, LF / HF (low frequency / high frequency) ratio, and log of normalized HF.
[0179] The HF spectrum of the beat-to-beat time (HRV waveform) is expressed in powers in the range 0.15–0.4 Hz. There is a 2.5-7 second parasympathetic or vagus nerve activity (respiratory sinus arrhythmia (RSA)) )) rhythm and is also called the "breathing band."
[0180] The LF band is 0.04 to 0.15 Hz and is thought to reflect baroreceptor activity at rest. (Some studies suggest that this may be related to cardiac sympathetic innervation.) (It has been done).
[0181] VLF (Very Low Frequency) HRV output is 0.0033~0.04Hz (300~25 seconds) Low levels are associated with cardiac arrhythmias and post-traumatic stress disorder (PTSD).
[0182] HRV parameters can also be extracted using time-domain methods (e.g., SDN N (standard deviation of normal interbeat intervals to obtain longer-term variability) and RM SSD (Root Mean Square of Successive Beat-to-Beat Interval Differences to Obtain Short-Term Variability). RMSSD is the It can be used to screen for irregular heartbeat behavior such as that seen in atrial fibrillation. It is also possible.
[0183] For HRV, a shift in the LF / HF ratio as calculated indicates detectable non-REM sleep. The characteristic of REM sleep is the "sympathetic" HF dominance (which is the sympathetic / parasympathetic nervous system (which may be related to the balance of
[0184] More generally, there is typically a greater increase in HRV during REM sleep.
[0185] Longer term averages or medians of respiratory rate and heart rate signals are important, especially after any intervention (e.g. medication, treatment, healing of illness (physical or mental), changes in health level, changes in sleep habits If there is a change over time, it is important for a particular person when conducting a longitudinal analysis. Slightly less useful for close comparisons (unless for very similar groupings) Therefore, the characteristics of respiratory variability and cardiac variability are Normalization (e.g., mean removal, median removal, etc., as appropriate for metrics) can be used to obtain a generalized distribution across the population. This is useful because it can lead to improved efficiencies.
[0186] For further analysis of the extracted features, we use a deep belief network (DBN). Such networks are called Restricted Boltzmann Machines (RBMs), Consists of autoencoder and / or perceptron building blocks. DB DBNs are particularly useful for learning from these extracted features. available and subsequently labeled data (i.e., with the input of human experts) It is trained on data that has been validated by the
[0187] Extracted through human-crafted "learning by example" Exemplary features that can be sent onto the DBN can include: type of apnea and location, respiratory rate and its variability over different time scales, breathing, inspiration time and and expiratory time, inhalation and exhalation depth, heart rate and its variability over different time scales. Sex, ballistocardiogram heartbeat shape / morphology movement and activity type (e.g., whole body movement), P LM / RLS, signal quality (completeness of measurements over time), user information (e.g., age, height, Other statistical parameters can also be calculated (weight, sex, health status, occupation). For example, the skewness, kurtosis, and entropy of a signal). DBNs determine some characteristics of the ("learning" these features) so that the DBN understands exactly what it is representing. It can be difficult to achieve this, but they often do the job better than humans. There are points where it may end up in a bad local optimum. After "learning" the features, the system: These can be fine-tuned using certain labeling data (e.g., human expertise). Data entry by experts allows for scoring of features (by one expert or by several By expert consensus).
[0188] DBNs can also learn novel features directly from input parameters (e.g., breathing From waveforms, activity levels, cardiac waveforms, raw audio samples (in sonar), I / Q Biomotion data (in case of sonar or radar), intensity level and color level (e.g. e.g., from infrared camera data).
[0189] Machine learning approaches that use only handcrafted features are called "shallow learning" approach, which tends to reach a plateau in performance levels. In contrast, "deep learning" approaches tend to scale with data size. In the above approach (in the case of DBN), Using deep learning to generate novel features for classical machine learning (e.g., novel feature acquisition, feature selection based on feature performance (winnowing), IC A (Independent Component Analysis) or PCA (Principal Component Analysis) (i.e., Whitening by the original degeneration method and decision tree-based approaches (e.g. , random forests, or support vector machines (SVMs).
[0190] In the case of a full deep learning approach such as the one used here, The selection step is avoided, allowing the system to capture the vast range of variation found in human populations. This can be seen as an advantage in that it is not used in the system. Signatures can be learned from unlabeled data.
[0191] One approach for these multimode signals is deep belief nets. After training the network on each signal, we train it on the concatenated data. A given data stream may simply not be valid for a given period of time ( For example, if the cardiac signal quality is below the acceptable threshold, but there is good respiration, movement, and activity, Audio feature signals are available, in which case they are learned or derived from cardiac data. The feature becomes meaningless for that period.
[0192] For classification, sequence-based approaches (e.g., Hidden Markov Models (HM) M) can be applied to such HMMs. Staged "sleep art" of the output sleep graph (such as may be provided via a laboratory PSG system) Mapping to the "architecture" and abnormal sleep stage switching However, sleep is a gradual physiological process. If you recognize this, you may prefer not to force the system into a small number of sleep stages, The system acquires gradual changes (i.e., has more "intervening" sleep states) ) becomes possible.
[0193] A simpler state machine approach without hidden layers is possible, but each The problem is to generalize to a large population of sleepers with unique human physiological characteristics and behaviors. Problems can eventually arise. Another approach is to use conditional random fields. Random Fields (CRF) or its variants (e.g., Hidden States) e) CRF, Latent Dynamic CRF, or Conditional Neural Function CNF or Latent Dynamic CNF). The LSTM is especially useful for sequence spanning (more typical in normal, healthy sleepers). Note that it can have high discriminatory power when applied to turn recognition.
[0194] Semi-supervised learning uses recurrent neural networks (RNNs). R can be useful in discovering structure in unlabeled data. NN is a standard neural network structure with input, hidden layers and output. have sequence inputs / outputs (i.e., the next input depends on the previous output (i.e., ,Hidden units transmit information using graph unrolling and parameter sharing techniques. LSTM RNNs are used in natural language processing applications. It is well known that LSTM is used to address the exploding gradient and vanishing gradient problems. .
[0195] In detecting the onset of sleep, when the speech recognition service is running, the voice input by the user is Mands can be used as a second determinant of "wakefulness" (this is because they are not just meaningless sleep talking). (This should not be confused with personal smart devices.) (if used), this can also be used as a wake determinant for other sleep / wake sensing service extensions.
[0196] [Sleep architecture, sleep scores, and other movement characteristics]
[0197] The above sensing operation using low frequency ultrasonic systems and technology detects the presence / absence of people, human movement, and can be implemented for the detection of multiple biometric characteristics. can be determined (e.g., respiratory rate, relative amplitude of breathing (e.g., shallow, deep), heart rate, heart rate and heart rate variability, intensity and duration of movement, and activity factor). One or more of the parameters are used to determine whether the subject is awake or asleep. If you are asleep, your sleep stage (light N1 or N2 sleep, deep sleep or RE) It is possible to determine the sleep stage (M sleep) and predict the next likely sleep stage. Such characterization methods can be used, for example, in the methods for processing and / or generating motion signals described below. The present invention can be implemented in accordance with the techniques set forth in International Patent Application PCT / US2014 / 0 45814 (Application date: July 8, 2014, International Patent Application PCT / EP2017 / 073 613 (filing date: September 19, 2017) and international patent application PCT / EP2016 / 0 80267 (filing date: December 8, 2016) and the patent applications mentioned above in this specification. Either one.
[0198] The system enables fully automatic and seamless detection of sleep and the ability to synchronize two or more people. It has the ability to detect the sleep of two or more people from one device. For example, different ranges from a smart speaker or processing device can be used to detect With different objects within the device, the movement is evaluated according to the device's individual sensing range. For example, as described in PCT / EP2017 / 073613, The data can be processed to detect different ranges from the processing device. Peak processing allows the evaluation of different ranges of motion characteristics to be monitored at different time dynamic ranges. It can be arranged to do so in the scheme (described in PCT / EP2017 / 070773) This automatically and periodically changes the detection range to allow sensing at different ranges. Optionally, the sensing waveform is Coding schemes can be used that allow simultaneous sensing at multiple ranges. .
[0199] High-precision sleep onset detection allows home speakers to be used for home automation / mono Especially when interfaced to Internet of Things (IoT) platforms. For example, by generating one or more service control signals, a range of services may be For example, if a user falls asleep, the lights in their home can be dimmed or changed in color ( For example, from white to red) and automatically close curtains or blinds. You can adjust the thermostat settings to control the temperature of your sleeping environment, and enjoy music The playback volume can be reduced and turned off over time. Detects the user's whole body movements. The detector's ability to do this can serve as the basis for automated instrument control. For example, If the user wakes up and walks around during the night, this movement is recorded by the generated sound and the sensed sound. The smart speaker can then detect the presence of the signal (by processing) and then can control automated lighting setting changes based on that specific motion detection. For example, this allows access to the restroom / toilet but does not unduly disrupt sleep. It is possible to control the lighting so that it is dimly lit so that it does not cause a disturbance (for example, In this regard, the equipment control response will turn on and The device can be set to on and off, or the device can be set to a step level change (e.g. The device can provide different activity or sleep-related signals (high vs. low light). The system may be configured to allow the user to preselect the desired system control behavior in response to associated detections. In addition, it may have a configurable setup process.
[0200] The system can automatically capture a person's entire sleep session, where Each stage, such as bedtime, bedtime, actual sleep time, wake, and final wake, is captured. Sleep fragmentation and sleep efficiency can be estimated for the individual. The sonar sensor described is an RF sensor described in PCT / US2014 / 045814. Although slightly different from the sleep analysis, the sleep analysis is performed after determining movement and / or measuring sleep-related parameters. The score can be estimated by the method described above. Typical sleep score input parameters include: total sleep time, deep sleep (Deep sleep) time, REM sleep time, light sleep time, wake-onset (WASO) time, and sleep Onset (sleep onset) (time it takes to fall asleep). Sleep score is based on the person's age and gender. This can be used to provide a normalized score versus its population norm (normative value). Historical parameter values may also be utilized in the calculation.
[0201] [Masking sound]
[0202] Using a smart speaker, you can play masking sounds through your own speaker(s). For example, smart speakers can generate conventional masking sounds. These masking sounds can cause other annoyances. Some people prefer to sleep with masking sounds because they can mask the environmental noise they are hearing. For example, by generating a sound that is inverted in phase compared to the perceived noise, The masking sound can be a sound-canceling noise. The system can be used to assist with sleep assistance and respiratory monitoring for adults or infants. In some cases, for example, systems may utilize Ultra Wideband (UWB) schemes. In order to produce lower ultrasonic acoustic detection, the masking noise itself can be made into the detected acoustic signal. This can be done.
[0203] [Sleep Start Service]
[0204] The processor or microcontroller may detect typical past sleep patterns, e.g., by acoustic sensing. Determination of stage or recorded sleep stages (historic sleep stage) and / or Based on the timing of your sleep stage cycle, your device will automatically Such predictions may be configured to predict the expected length of the page. For example, it is possible to control the change of audio sources and lighting sources (i.e., smart The speaker or processing device may control the equipment, for example, by wireless control signals. , you can adjust the device (as described above). For example, a smart speaker can It can be configured with continuous sensing, i.e., when a person is present (day and night). When a person is asleep (whether during a full sleep session or a nap), they are awake. It will be possible to automatically identify periods of awake and absent time. The monitor can identify when the user begins to fall asleep and when the user begins to wake up. With such detection, the system can detect when someone falls asleep near the device. In response to this detection, the audio content (speaker(s), TV, etc.) The playback volume of the audio may then be reduced, indicating that the person has entered a deeper sleep phase. If the device detects this, it will turn off the volume and TV after 5-10 minutes. Such detection can include automatic dimming, setting automated burglary alarms, and adjusting heating / air conditioning settings. Adjustment of associated smart device(s) (e.g., smartphone) It can also function as a control decision for activating the "Please Do Not Enter" feature above. can.
[0205] When sleep is detected by this process, the smart speaker or processing device 100 Voice assistant prompts can be disabled, preventing people from accidentally Such prompts can help prevent situations that may lead to sudden awakening. A defined wake-up window (e.g., waking a person from a preselected sleep stage) It acts as a "smart alarm" and prompts you to minimize sleep inertia as much as possible. can be disabled until the
[0206] The processing device makes such control decisions in response to the number of people detected in the sensing space. For example, if there are two people in a bed, one of them is asleep and the other is awake, When the device detects that you are using a voice assistant, the system may, for example, adjust the volume of the voice assistant ( Adjust your volume settings to the lowest volume possible (while still allowing you to use the Assistant). As a result, an awake person can hear and control the device. This minimizes the risk of waking up a sleeping person due to the When a person is detected while awake and at least one person is detected while asleep, If content is being played (e.g., music or a movie), the device will For those who are listening to music, do not lower the volume of the media or turn the volume of the media up a little. Then, if the device detects that the remaining person(s) have also fallen asleep, , the device may further reduce and turn off the media content.
[0207] In some cases, the processing device may be configured to receive a notification of a user or users in the processing device's room. Based on location and / or presence detection, volume levels (audible audio content) For example, the volume may be reduced if there is no detection of a user's presence. The device may, for example, reduce or increase the sound level with increasing distance from the device. By increasing the volume and / or decreasing the volume with decreasing distance from the device, Produced location (e.g., distance from a processing device or a specific location (e.g., a bed)) Volume control may also be controlled based on or as a function of this detected position. For example, the processing device may lower the volume if it detects that the user is in bed. or increase the volume if it detects that the user is away from the bed. stomach.
[0208] [Awakening Start Service]
[0209] The system can provide services for several wake-up scenarios, such as: For example, if a user wakes up during the night, the user can choose to turn on, for example, a meditation program in their personal settings. It can be used with a computer to receive voice assistance to help users get back to sleep. Instead of or in addition to voice assistance, the system can By controlling the parameters of the device and changing the settings of the connected device, the person can be put to sleep again. Examples of this include changing the temperature setting, lighting, etc. Bright setting, e.g. TV, image projection onto a display panel or projected image generation by a projector Configurations include, for example, controlling the activation of bed vibrations by connected / smart bed shakers. For example, in the case of infant sleep, changes in sleep stages (e.g., from deep to light sleep stages and by detecting increased body movement) that the infant is about to awaken or After the acoustic motion sensing device detects that the child is awake, the rocking cradle is activated. Optionally, audio content can be played (e.g., children's songs). or music playing on the smart device's speaker) or automated This can be used to improve sleep time or parental attention to the infant. Increasing the amount of time that attendance can be delayed can be helpful, especially for restless infants or young children. It gives parents of children a much-needed break.
[0210] If the user sets an alarm time window, the processing device It listens to your sleep stages to find the best sleep stage for you (usually a light sleep stage). In the absence of such a stage, the processing device Low light and sound can be actively introduced to induce the user into different stages of sleep (e.g., deep sleep stages). or REM sleep stage) to lighter sleep and then wakefulness. Such smart alarm processing devices are also becoming increasingly important in connected home automation. The device may also be configured to control device functions, for example, when a user's awakening is detected: The processing device communicates with the appliances to, for example, open automated curtains / blinds. to increase alertness, to have an automated toaster oven heat up your breakfast, It is possible to have an automated coffee machine start making coffee.
[0211] It is also possible to configure a nap program using biometric sensing. I have an appointment at around 3:30pm so I need to be alert by then. It is possible to take a nap knowing that the child will be able to do so (i.e., using light and sound stimuli to stimulate the child during the day). Avoid waking up from deep sleep too quickly or not entering deep sleep at all. (This will prevent you from waking up feeling groggy.)
[0212] [Sleep improvement service]
[0213] User's recent trends and population standards of people with similar age, gender, and lifestyle Based on this, advice is sent to the user to improve their sleep behavior (sleep hygiene). can be achieved.
[0214] The described system further enables the collection of feedback from users. Feedback includes how the user is currently feeling, how well they slept last night, and the advice they received. This may relate to whether the use of certain medications or exercises was beneficial. Such feedback can be collected via an input device. This could be a smartphone keyboard. A smart speaker could be a personal audio If it includes an assistant application function, it will guide user feedback and Processing can be done via voice (e.g., answering questions about the user's condition) , asking users how they feel, providing personalized feedback to users (e.g., (Providing data on the user's sleepiness and fatigue state). The processing device may be configured with software modules for natural language processing. The software module can be used as a conversational interface, e.g., and guiding a conversation with the system to obtain information from the processing device that may be based on an evaluation of the motion signals. Provides a type sequence (e.g., an audible linguistic command / query). The management module or conversational interface is based on the Google Cloud platform. Team, DialogFlow Enterprise development suite It can be embodied by te.
[0215] As an example, if there are one or more processing devices 100 in a multi-room or multi-floor building, Consider the case where a processing device 100 using acoustic-based sensing is placed in a kitchen. One can be placed in a bedroom, and when preparing breakfast, a person can say, "OK Google, By asking verbally, "How was your sleep?", the system generates a query about the user's own sleep information. In response, the processing device may The night before can be detected by an acoustic-based motion sensing application (i.e., in the bedroom). Sleep parameters can be retrieved from the current sleep session (and trends from the previous night). Such queries to devices such as Fitbit, Garmin watches, and Res MedS+ and any other web sites (e.g., containing queried sleep-related data) It also applies to retrieving recorded session data from a site or networked server. It can be used.
[0216] Therefore, people may choose when and in what form they receive advice about their sleep. In this regard, the sleep management system provides sleep advice. Snuggets can be delivered or presented interactively through audio content. Device 1, in which audio content is processed through question and response linguistic inquiries 00 microphone(s) and speaker(s) Sleep management systems (e.g., the sleep management system described in PCT / US2014 / 045814) are shown. Advice from the sleep management system may be delivered via the processing device 100, and sleep habits may be improved. This may relate to a change in habits, a lifestyle change or an actual new product recommendation.
[0217] For example, an exemplary query using the speaker and microphone of processing device 100: The / answer query session can produce the following output: Redmond. Your recent sleep breakdown, your purchase history and daytime discomfort index. According to reports about you, you may benefit from a new mattress. There is a possibility. Premium new internal coil spring mattresses are on special sale. Sold at a discounted price. The memory foam option is not to my liking. So, I recommend this. Would you like to order it? and generating one or more searches of the Internet and recording user data and processing This may be based on sleep state detection performed by the device.
[0218] The system may therefore provide, for example, acoustic sleep detection and interactive verbal communication. Based on (a conversation between a user and a voice assistant application on the processing device 100) This allows for feedback and further action based on the The processing device accesses environmental data (e.g., sound data, temperature data, light data) ), so that searches or other detections by environmentally related systems and sensors This may include generating control signals for controlling the operation of environmental systems. The generated output of such a platform could lead to the development of related products to address sleep conditions, for example. It will also be possible to propose and sell products to people. By deploying the system on a platform, directed conversation with the system becomes possible. Then, by interacting with the processing device 100 (rather than just sending sleep-related messages), For example, a user may input a sleep signal to a processing device. On the other hand, if you ask, "OK Google, how did I sleep last night?", it will say, "Hello Redmond. This is the ResMed Sleep Score application. My score was 38. Hmm, that seems low." "I wonder what happened?" "So... We've discovered some issues with your sleep environment and your breathing patterns. Learn more Would you like to know?","Of course!","The temperature in the bedroom was 77 degrees Fahrenheit (25 degrees Celsius). This is too hot. You can turn on the air conditioning an hour before you go to bed tonight. "Is that okay?", "Yes, please.", "Can you teach me about breathing?" "You were detected snoring loudly last night. Dreaming R During EM sleep, gaps in breathing were observed. Can you explain this? "," "Yes, please. What does that mean?"," "Did you snore loudly last night?" There were several interruptions in the breathing pattern, which is the reason for the low score. "Do you often feel tired?" "Yes, I do." "I'll call my doctor." Would you like to call your health plan and discuss this condition? 0 minutes are free."
[0219] For such interactions, it works in conjunction with the processing of a support server on the network, for example. The processing device 100 may incorporate artificial intelligence (AI) processes for increased flexibility and scope. Such improvements may include, for example, the combination of the following information: For example, information from detected daytime activity (e.g., detected exercise or stair information) information (e.g., recorded user data) as well as wearable GPS, accelerometers, heart rate sensors, This information can also be detected through pulse sensors and motion sensors. ) heart rate data from other activity monitors and personalized services (e.g., combining detected data) Detected sleep for delivery of sleep-related advice and product offers based on the user's sleep habits It is information.
[0220] [Biometric Recognition Services (see also Security Sensing)]
[0221] Human recognition, control of devices in the living environment as outlined above, and cloud services. For data exchange, biometric features can be used. The analysis of motion-related characteristics derived from the signal(s) of the biomechanical movement Such a motion-based biometric method can realize trick recognition. , PCT / EP2016 / 058789. Such biometric identification is derived from the acoustic-based sensing methods described herein. The motion signal may be based on the motion signal.
[0222] By performing human recognition using ultrasonic sensing, the processing device 100 can The device distinguishes between noise and acoustic sensing signals and is used to verify the presence of a person near the device. Therefore, in some versions, the processing device 100 performs the beamforming process where different channels (different speaker signals) By learning the appropriate delay, the timing of the reception of the acoustic signal is adjusted to match the proximity and The system can adjust for differences related to the position and location of the speakers. To further localize the detection area (e.g., for detailed range detection and object tracking) The presence of an object is detected in the vicinity when the object is detected. When using a 3D printer, such a process can be performed as a setup or adjustment procedure. do.
[0223] As mentioned above, low frequency ultrasound provides an advanced security sensing platform. For example, the processing device 100 may detect the absence / presence of an object within the detection range. If and / or until such absence / presence is detected, biometric signal processing Detect and identify people based on
[0224] In some cases, the processing device 100 further adapts its sensed signals to at least Both systems offer two modes: one is to detect when people are entering the monitored area, and the other is to It may include a general monitor mode intended to detect when an intrusion occurs, and may include a different detection scheme. Systems (e.g., UWB, CW, FMCW, etc.) can be used for presence detection. After detecting the presence of a person within the monitor area, the system performs extended ranging and Selective cardiac sensing, respiratory sensing modes, etc. (e.g., dual-tone FMCW or (A) Then, for example, the person's "signature" parameters can be transferred to the processing device. device (e.g., record data) or another device (e.g., network) that cooperates with the processing device. If the detected parameters (e.g., cardiac parameters, breathing parameters, whole body movement parameters, etc.) to identify people. Alternatively, if no such signature is recognized in connection with the recorded data, processing may be suspended. The management device (or collaborative server) may determine that another person has entered its perceived proximity.
[0225] Thus, the system can detect unknown or unrecognized objects by sensing the area around the device. It can detect undetected bio-motion signals or unrecognized gestures. Based on such detection, the system (e.g., processing device) may detect potentially unwanted A system may initiate or communicate a warning or other message about the presence of a The smartphone will send an alert. The phone detects that someone other than the designated owner is using or attempting to use the phone. If so, the smartphone may determine that it has been forgotten / abandoned.
[0226] In another example, a Google Home smart speaker or other smart speaker The camera or processing device may monitor the apartment, room or house using acoustic sensing. This can be done by using a remote control, and can function as a highly intelligent intruder alarm. For example, The processing device may detect movement within the apartment, for example after a period of absence of movement. In response, the processing device may control one or more automated instruments. For example, by a control signal generated or initiated by a processing device based on this detection. This process allows you to, for example, ask the person to say their name, and then turn on the secondary lighting. The inquiry can also be initiated via the speaker by asking the Or, the person can be audibly greeted. The processing device attempts to identify a language signature associated with a particular known person. performing (e.g., comparing this audible response with a pre-recorded audible response (e.g., The evaluation of this response can be carried out via a microphone by means of sound wave analysis. If the response is not available, the processing device may terminate the challenge / response interrogation. If not, or if the processing device does not recognize the language signature, the processing device may It may communicate with other services through its communication interface, e.g., a processing device over a network or the Internet (e.g., email, text message, (by controlling the generation of messages, recorded voice messages, etc.) Another person or the owner can be notified. In other words, the processing device can optionally and, typically in conjunction with one or more further system processing devices or servers, i.e., acoustic detection, audible challenge, audible authentication and authorization or alarm or is the communication of a warning message. Furthermore, optionally, instead of voice recognition, e.g. In addition to speech recognition, the processing device may also perform the above-mentioned functions, for example by detecting breathing and heart rate. This allows for motion-based biometric authentication such as Provides a means of authentication or secondary authentication of people in the vicinity of a device.
[0227] The processing device may, for example, initiate further processing if the person's language signature is recognized. In that case, for example, the processing device may be configured to perform a further interactive query. The processing device may be linked to the person after the light is turned on, for example. For example, before going out, the person in question (Redmond) may have been going to the gym. You can even tell your processing device that you are in the forest, or you can use your own mobile Integrates with a tracking service on your phone or exercise device, allowing you to track your location. It can also be shared with the processing device (e.g. via GPS). When the device acoustically detects the presence of a person (and optionally recognizes the user / Redmond) The processing device may initiate a conversation such as: "Redmond, how's Jim?" How was it? Were you able to update your personal record? The processing device is a wireless device. For example, the processing device may detect the arrival of a Wirelessly enabled movement / health trackers (e.g., Bluetooth Fitbit) It can detect the number of steps, heart rate or other activity parameters and download them. Such data can also serve as proof of authenticity of the authentication / recognition of a particular person. Further conversations generated by the device may be based on the downloaded data ( providing an audible response via a speaker); other parameters (e.g., camera-based facial or body recognition) in multimodal identification (e.g., by a processing device) The method may also be initiated by a device (including a camera device that can be controlled via the camera).
[0228] In the room, if the non-intruder alarm is configured or if the person in question has already been identified The detected motion can act as the basis for controlling lighting changes. When darkness is detected, the time is recorded by the time clock or connected to the processing device. The processing device may be used for home automation or to turn on or increase illumination via the processing device 100's own light source. A control signal for the
[0229] The CW acoustic sensing mode detects drafts (air currents) in the room and can be used to monitor, for example, the Such reminders can serve as the basis for the creation of reminders. If you leave the door open while you go out or if the fan is running, and audible conversational messages to people via speakerphone. Automated fan control by a processing device based on the detection of the presence or absence of adequate ventilation. The device can be activated or deactivated.
[0230] [Range Detection]
[0231] The speaker of the processing device 100 (e.g., using the acoustic sensing methods described herein) By detecting the distance of one or more people from the or through other home automation services (e.g., devices) For example, the processing device may be configured to process, for example, a voice assistant application. For this option, you can adjust the volume level of the processing device's speaker(s) by can be set to an appropriate level based on the detected distance (e.g., (The volume is increased if the distance is small, and decreased if the distance is small.) The volume can be automatically changed based on the determined user distance.
[0232] In some versions, other detections may also function as distance detection enhancements. For example, range information can be obtained from other sensors (e.g., infrared components, e.g., By combining distance detection, range (or This allows for increased accuracy of the estimation of the speed (or indeed velocity).
[0233] Using data from near-field and far-field microphones (both Sonar echo (especially when multiple microphones are on the same hardware) This also allows for improved localization of the
[0234] [Gesture Control]
[0235] The processing device bases gesture recognition on the acoustic sensing methods described herein and It can also be configured to do this using processing of physiological movement signals. For example, a processing device The process involves the use of characteristic gestures (e.g., as described in PCT / EP2016 / 058806). configured to recognize a specific hand movement, arm movement, or leg movement from a motion signal. (e.g., when physiological movement signals are derived by the acoustic sensing methods herein, Therefore, the functions of the processing device that can be controlled by the processing device and / or the Adjust the functionality of your mobile device based on the detected gestures (e.g., increasing the speaker volume). a clockwise circular movement to increase volume, a counterclockwise movement to decrease volume, and optionally Turn off the device (e.g., processing device 100 or any of the equipment described herein) The gesture can be performed by following the steps below: may be used for controlling (e.g., causing a processing device to initiate communication (e.g., When sending notifications using a computer's management device (e.g., voice or emergency calls, text messages, etc.), Message, email, etc.
[0236] [Respiratory care and feedback services]
[0237] The processing device 100 may be configured to provide daytime applications, such as for stress and blood pressure reduction. This daytime application can be implemented with meditation outputs (e.g., deep breathing exercises). exercises) based on acoustically derived movement signals, e.g., breathing movements. The signal is processed / evaluated by the processing device 100 to generate timing cues. These timing cues allow for customized mindfulness. Control breathing exercises (e.g., "belly breathing" at a rate of less than 8 breaths / min) output to speaker(s) and / or associated equipment (e.g., lighting control) The system can assist in providing respiratory entrainment. The use may, for example, allow the device to provide audible and / or visual feedback to the user. In the case of control, the detected respiratory curve and the respiratory stimulus / cues delivered by the system are compared. By linking the values to the target values, it becomes possible to achieve the target values.
[0238] The respiratory cue can be an audible stimulus, and the system plays music at a target breathing cadence. After pre-filtering, the sensed signals are overlaid and then reconstructed. Selecting music of a type with a tempo close to the user's detected breathing rate and / or heart rate It can also be done as follows.
[0239] When anxious or stressed, a user's breathing pattern may become shallow and rapid, In this case, the upper chest and neck muscles are used to breathe rather than the abdominal muscles. Some benefits include reducing stress hormones through health promotion without drug intervention. This helps to "balance" the autonomic nervous system (parasympathetic and sympathetic) and stimulates alpha brain waves. Increases blood flow and improves diaphragm and abdominal muscle function (especially in those with respiratory problems) in the case of).
[0240] In an exemplary relaxation game provided by a processing device, The patient is asked to breathe deeply and the processing device detects breathing movements through acoustic motion sensing. A shape may be presented on a display device controlled by the processing device based on Controlling the ball or balloon for automatic sensing by the acoustic sensing method described herein The ball or balloon appears on the display in a manner synchronized with the user's chest movement. Inflate and deflate the target (e.g., steady breathing) for 3 minutes. When the cursor reaches a certain point (i.e., without any movement), the processing device changes the display. The processing device can indicate completion by raising a button, indicating that the user has "won" If a cardiac signal is also detected by the processing device, a notification may be sent via a speaker. Increased consistency between breathing and heart rate will award a few more points in the game Instead of shape, light whose intensity changes in sync with breathing or a strong rhythmic component can be used. The reward for "winning" the relaxation game may be a light color. Or it could be a change in musical sequence.
[0241] Other uses of the system include pacing (via light or display devices). Adjusted illumination and / or increased user breathing rate and modulation of user inhalation / exhalation times There are wakeful breathing exercises with one or more of the special audio sequences for Biofeedback from low frequency ultrasound biometric detection is optionally used. can be done.
[0242] [Improved privacy]
[0243] Many virtual assistants use speech recognition services (e.g., "Okay, go!"). Continuously listens for keywords / phrases such as "Guru," "Alexa," or "Hey, Siri" (microphone(s)) (The keyword is always on). In this respect, it is usually active. Anything said after the code / keyphrase can be sent to the internet. Some people find it unpleasant to have their conversations constantly monitored in this way. One advantage of motion detection is that it keeps the processing device running until the user is ready to use it. At some point, the device can disable the microphone. The voice assist operation of the device may be configured not to be activated under normal circumstances. The device may perform acoustic motion detection (e.g., using detection of specific movements or motion gestures). The method is applied to microphone-based speech recognition (rather than the typical speech keyword recognition using a microphone). You can start the Crophone Monitor, which will constantly monitor your conversations. This eliminates the need for a processing device, improving user privacy. Although it is possible to do this in close proximity to the processing device, In this way, standard movement gestures or user gestures can be detected. Customized movement gestures (e.g., learned during the setup process) The audio signal (which is generated by the audio signal) can be used to trigger a processing device to activate the speaker. This may initiate a natural language session with the processing device (e.g., via a speaker). triggering a verbal response from the processing device). However, optionally, the processing device After activation, the system may wait an additional period to detect an audible keyword via a microphone. The language keywords are then used to initiate a natural language conversation session with the processing device 100. This can act as a trigger for
[0244] Such a method for activating a processing device or initiating a session is more cost-effective than conventional methods. It may also be preferable to have the processing device assist when the user is standing still and talking. The app will not listen for language keywords until motion is detected. This feature allows the processing device to automatically listen to speech (or other information in the room) if no one is in the room. However, it can mean that the user is not listening to the ("I'm leaving the room now, so please monitor the room.") Active listening mode can be selectively turned on for more continuous operation.
[0245] Another privacy benefit is the need for recording cameras in the bedroom / living room. No need for detailed sleep tracking or daytime absence / presence, breathing and / or heart rate tracking There are some aspects of Noh that are unique to it.
[0246] [Other uses]
[0247] Another application of this technology is where a speaker and microphone are already available. Or it may be possible to easily retrofit it. For example, in elevators. occupancy sensors (used for emergency communications) and existing speakers and microphones can be used to sense breathing signals or other movement signals in space. Cut.
[0248] For range sensing using FMCW or large area sensing using CW, It is also possible to use a Pair Address (PA) system. In some cases, a 20 kHz or similar tone is already used to ensure proper operation. and ultrasonic biomedical applications (e.g., for use in train stations, libraries, shops, etc.). Upgrades are available to include trick detection.
[0249] [Power nap]
[0250] The processing device 100 of the system is a smart nap assist device. The bed can be programmed to have a nap function, which is activated when a person is lying down in bed. whether you are sitting, on the couch, or in any area / nearby of the processing device For example, a user may say to the device, "I'm going to take a nap now." The interactive audio device can then play a This can be audibly assisted by sleep and / or time monitoring. Based on the predicted available time, current sleep deprivation estimate, time of day and user request. For example, the duration of a nap can be adjusted to a target duration. for a maximum of 20 minutes, 30 minutes, 60 minutes, or 90 minutes (a full sleep cycle) A 60-minute nap is designed to optimize deep sleep and can be completed with While allowing some time for recovery from any sleep inertia, the 20-minute and 3-minute A target time of 0 minutes will either wake the user up when they are still in a light sleep stage or It is optimized to wake you up within 1-2 minutes of entering deep sleep. In addition, the time spent awake before sleep (nap) is also recorded.
[0251] If your recent sleep patterns are normal, a 20-25 minute nap is better than a 90 minute nap. This may be preferable to a full sleep cycle because the longer duration may affect sleep quality for the night. This is because it may affect your sleep.
[0252] [Multi-user proximity sensing]
[0253] In some versions, the processing device may transmit the audio to one or more speakers (single or multiple speakers). Two or more people can be simultaneously monitored by one or more microphones. For example, the processing device may be configured to monitor different users on different frequencies. It is possible to generate multiple sensing signals at different sensing frequencies for sensing multiple signals. In this case, the processing device may use an interface to sense different users at different times. The generation of the left sense signal (e.g., different sense signals at different times) can be controlled. In some cases, the processing device may detect signals at different distances (e.g., in parallel sensing). The distance gating for sensing at different times can be adjusted sequentially.
[0254] In some cases, there are several conditions that can help maximize signal quality: For example, a processing device (e.g., an RF sensor or a sonar-enabled smart In this case, the first person's body can be placed on the bedside locker. This can result in a "shadowing" effect whereby a large portion of the sensed signal can be blocked (in which case the distance Benefits from remote gating (detecting only one person in bed). , the processing device detects two (or more) different sensing signals to detect two (or more) people. Or even generate a single FMCW (triangle, dual ramp or other) sensing signal. The isolation of the FMCW range allows for One sensing signal is sufficient (as if monitoring two people simultaneously). In order to achieve this, users should place the processing device in a high position to absorb most of the acoustic energy. so that the message reaches both users (e.g., the first user and the user's sheet / duvet) (Ensure that the sensing signal is not significantly blocked by the comforter.) Even if there are two microphones (e.g., near peak on one and low on the other), (In particular, the smaller amplitude received signal is more likely to be received by a target at a greater distance.) This can be advantageous when there is significant constructive / destructive interference.
[0255] It may be more advantageous (and possibly more costly) to have the sonar / smartphone closer to the second person. (A higher quality signal can be achieved.) The second person uses their smartphone / processing The first person can use their smartphone / processing device to The generation of the sensing signals is time- and / or frequency-dependent to avoid mutual interference. These devices can listen before transmitting (to avoid overlapping frequencies). and / or by processing signals received from the user This allows for efficient sensing, without interfering with existing sensing signals from other nearby devices. Thus, the sensing signal modulation / technique can be selectively selected.
[0256] In some cases, the audio pressure levels of the transmit and receive sensitivities are different between the two smartphones. This is so that if two phones have the exact same transmitted sense signal, they will not cause interference. If so, a mixture of air damping and acoustically absorbing surfaces (fabrics, carpets, bedspreads) The second source may fall below the interference threshold due to the effect of the second source. When used by each user, the device automatically reduces its output. , whereby the biometrics of the subject being monitored can be The signal can be detected sufficiently with, for example, the minimum required power. This allows each device to detect motion while avoiding interference with other devices. become.
[0257] [Noise / Interference Avoidance]
[0258] In some cases, the processing device may transmit content sounds (e.g., music, speech) to a speaker. When playing on a computer, the processing device applies a low pass filter to this sound content. This allows us to (for example) remove all content above 18kHz. The processing device may overlay the low frequency ultrasonic signal for sensing. The stem removes (or significantly attenuates) components of the "music" that may interfere with perception. However, in some cases, such filtering may be omitted. Such filtering can lead to cloudy sound content (high-pitched sounds) in high-quality Hi-Fi. (or the user may feel that way (rather than the actual sound)) (This can cause the perception of a low-pass filter to the music / sound content.) No filtering (or (in case of playback on separate and nearby speakers) (without actually providing low-pass filtering capabilities for music / sound content) Therefore, unwanted components may be included in the sensing signal. The method provides feedback on real-time and recent (in-time) "in-band" interference. There are ways to adapt the sensing signal by implementing a Depending on the processing resources (which may need to function extremely fast (low latency)), The device may transmit two sensing signals simultaneously (in the case of FMCW signals, two co-located (It could be a triangular lamp placed in front of the camera or actually two sets of dual lamps.) The device may include a voting process in which the obtained biometric motion signals are The signal is evaluated and a better signal is selected (in terms of signal quality) based on the detected motion in the signal. This process is dynamic throughout the sensing process (time of use). It can continue.
[0259] 5.2 Other Notes A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner reserves the right to modify, revise, or otherwise modify this patent document. If any person reproduces this patent document or this patent disclosure by facsimile, the Patent Office's patent file If it is something that is recorded in the mail or record, there is no objection if it is for a specific purpose, but if it is for any other purpose, All rights reserved.
[0260] Unless otherwise clearly indicated by the context and unless a range of values is provided, the lower limit 1 / 10 of a unit, between the upper and lower limits of the range, and any other stated value in the stated range It is understood that each intervening value for any of the above is encompassed by the present technology. The upper and lower limits of these intervention ranges, which are independently included in the intervention range, are included in the stated range. If the scope of the description falls within these limits, it is also included in the present technology. Any extent exceeding either or both of these stated limits, including one or both, This technology is encompassed by the present invention.
[0261] Furthermore, when a value or values are provided herein and implemented as part of the present technology, other Unless otherwise specified, such values may be approximated and may vary as practical technical implementations allow or require. It is understood that such values may be used to any suitable degree of significance. .
[0262] Unless otherwise defined, all technical and scientific terms used herein belong to the art. It has the same meaning as commonly understood by those skilled in the art. and any methods and materials similar or equivalent to the materials used in the practice or testing of this technology. Although a limited number of exemplary methods and materials can be used in It will be described.
[0263] Although certain materials are described as being suitable for use in the construction of components, their properties may vary. Similar and obvious alternative materials may be used as substitutes. Insofar as any and all components described herein are understood to be manufacturable. Therefore, they can be manufactured collectively or separately.
[0264] As used herein and in the appended claims, the singular form "a" "an" and "the" are used interchangeably unless the context clearly indicates otherwise. Please note that the term "includes multiple equivalents of" and "includes multiple equivalents of".
[0265] All publications mentioned herein are incorporated by reference in their entirety for all purposes, including, but not limited to, the methods and / or methods that are the subject of these publications. or the disclosure and description of the materials are incorporated herein by reference. The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein should be construed as a substitute for any express or implied warranty that the present technology is in any way indicative of a prior art patent. Nothing herein should be construed as an admission that the disclosure of any of the disclosed publications is not ancillary to the present invention. The dates may be different from the actual publication dates, which may need to be independently confirmed.
[0266] The words "comprises" and "comprising" mean elements, constituent elements, The elements or steps described should be interpreted in a non-exclusive sense. An element, component, or step may be combined with other elements, components, or steps not specified. Indicates that they can exist, be used, or be combined together.
[0267] Headings used in the detailed description are for the convenience of the reader and are provided to assist in understanding the present disclosure or should not be used to limit what appears in the claims as a whole. These headings are provided for informational purposes only and should not be construed as limiting the scope of the claims or the limitations of the claims. should not be used in this way.
[0268] The technology herein has been described with reference to particular embodiments, but these embodiments It should be understood that these are merely illustrative of the principles and applications of the present technology. In some cases, terms and symbols may indicate specific details that are not necessary for the practice of the present technology. For example, the terms "first" and "second" (etc.) are used, unless otherwise specified. These terms are not intended to denote any order, but rather to distinguish between separate elements. Furthermore, the process steps in the method are described or illustrated in order. Although the descriptions may be presented in order, such order is not required. The order of these aspects may be changed and / or the aspects may be performed simultaneously or even synchronously. Recognize that it is possible.
[0269] It is therefore to be understood that numerous modifications may be made in the illustrative embodiments and that other arrangements may be devised without departing from the spirit and scope of the present technology. The claims of the parent application as originally filed were as follows: Claim 1: 1. A processor-readable medium having stored thereon processor-executable instructions, the processor-executable instructions, when executed by a processor of an interactive audio device, causing the processor to detect a physiological movement of a user, the processor-executable instructions comprising: instructions for controlling generation of an audio signal into the vicinity of the interactive audio device via a speaker connected to the interactive audio device; instructions for controlling sensing of reflected sound signals from the vicinity via a microphone connected to the interactive audio device; instructions for deriving a physiological movement signal using a signal indicative of at least a portion of the sensed reflected sound signal and at least a portion of the audio signal; instructions for generating an output based on an evaluation of at least a portion of the derived physiological movement signal; 1. A processor-readable medium comprising: Claim 2: 10. The processor-readable medium of claim 1, wherein at least a portion of the generated audio signal is substantially in an inaudible range. Claim 3: 3. The processor-readable medium of claim 2, wherein a portion of the generated audio signal is a low-frequency ultrasonic acoustic signal. Claim 4: 4. The processor-readable medium of claim 1, wherein the signal indicative of a portion of the audio signal comprises an internally generated oscillator signal or a direct path measurement signal. Claim 5: 5. The processor-readable medium of claim 1, wherein the derived physiological motion signals include one or more of respiratory motion, whole body motion, or cardiac motion. Claim 6: 6. The processor-readable medium of claim 1, wherein the instructions for deriving the physiological movement signal are configured to multiply an oscillator signal by a portion of the sensed reflected sound signal. Claim 7: 7. The processor-readable medium of claim 1, further comprising processor-executable instructions for evaluating the audible verbal communication sensed via the microphone connected to the interactive audio device, wherein the instructions for generating the output are configured to generate the output in response to the sensed audible verbal communication. Claim 8: 8. A processor-readable medium according to any one of claims 1 to 7, wherein the processor-executable instructions for deriving the physiological movement signal include demodulating a portion of the sensed reflected sound signal using at least a portion of the audio signal. Claim 9: 9. The processor-readable medium of claim 8, wherein the demodulation comprises multiplying a portion of the audio signal with a portion of the sensed reflected sound signal. Claim 10: 10. The processor-readable medium of claim 1, wherein the processor-executable instructions for controlling generation of the sound signal result in generation of a dual-tone frequency modulated continuous wave signal. Claim 11: 11. The processor-readable medium of claim 1, wherein the dual-tone frequency modulated continuous wave signal comprises a first sawtooth frequency variation superimposed with a second sawtooth frequency variation in a repeating waveform. Claim 12: 12. The processor-readable medium of claim 1, further comprising processor-executable instructions for generating an ultra-wideband (UWB) audio signal as audible white noise, the processor-readable medium comprising instructions for detecting user movement using the UWB audio signal. Claim 13: 13. The processor-readable medium of claim 1, further comprising processor-executable instructions for generating a probing sound sequence from the speaker in a setup process for calibration of distance measurements of low-frequency ultrasonic echoes. Claim 14: 14. The processor-readable medium of claim 1, further comprising instructions executable by a processor to generate time-synchronized calibration acoustic signals from one or more speakers, including the speaker, in a setup process for estimating a distance between the microphone and another microphone of the interactive audio device. Claim 15: 15. The processor-readable medium of any one of claims 1 to 14, further comprising processor-executable instructions for activating a beamforming process to further localize the detected region. Claim 16: 16. The processor-readable medium of claim 1, wherein the generated output based on evaluation of the portion of the derived physiological movement signal includes monitored user sleep information. Claim 17: 17. The processor-readable medium of claim 1, wherein evaluating the portion of the derived physiological movement signal comprises detecting one or more physiological parameters. Claim 18: 18. The processor-readable medium of claim 17, wherein the one or more physiological parameters include any one or more of respiratory rate, relative respiratory amplitude, heart rate, cardiac amplitude, relative cardiac amplitude, and heart rate variability. Claim 19: 19. The processor-readable medium of claim 16, wherein the monitored user sleep information includes any of a sleep score, a sleep stage, and time in a sleep stage. Claim 20: 20. The processor-readable medium of claim 1, wherein the generated output comprises an interactive query answer presentation. Claim 21: 21. The processor-readable medium of claim 20, wherein the generated interactive query answer presentation is performed via a speaker. Claim 22: 22. The processor-readable medium of claim 21, wherein the generated interactive query response presentation includes advice for improving monitored user sleep information. Claim 23: 23. The processor-readable medium of any one of claims 1 to 22, wherein the generated output based on evaluating the portion of the derived physiological movement signal is further based on accessing a server on a network and / or performing a search of a network resource. Claim 24: 24. The processor-readable medium of claim 23, wherein the search is based on recorded user data. Claim 25: 25. The processor-readable medium of any one of claims 1 to 24, wherein the generated output based on evaluation of the portion of the derived physiological movement signal comprises a control signal for control of an automated device or system. Claim 26: 26. The processor-readable medium of claim 25, further comprising processor control instructions for transmitting the control signal to the automated device or system over a network. Claim 27: 27. The processor-readable medium of claim 1, further comprising processor control instructions for generating a control signal for modifying a setting of the interactive audio device based at least in part on an evaluation of the derived physiological movement signals. Claim 28: 28. The processor-readable medium of claim 27, wherein the control signals for changing settings of the interactive audio device include a volume change based on detection of a user distance from the interactive audio device, a user state, or a user location. Claim 29: 29. The processor-readable medium of claim 1, wherein the processor-executable instructions are configured to evaluate movement characteristics of different acoustic sensing ranges for monitoring sleep characteristics of multiple users. Claim 30: 30. The processor-readable medium of claim 29, wherein the instructions for controlling the generation of the sound signals control generating simultaneous sensing signals at different sensing frequencies for sensing different users at different frequencies. Claim 31: 31. The processor-readable medium of claim 30, wherein the instructions for controlling the generation of the acoustic signals control the generation of interleaved acoustic sensing signals for sensing different users at different times. Claim 32: 33. The processor-readable medium of any one of claims 1 to 32, further comprising processor-executable instructions for detecting the presence or absence of a user based at least in part on the derived physiological movement signal. Claim 33: 33. The processor-readable medium of claim 1, further comprising processor-executable instructions for performing biometric recognition of a user based at least in part on the derived physiological movement signals. Claim 34: 34. The processor-readable medium of any one of claims 1 to 33, further comprising processor-executable instructions for generating communication over a network based on (a) a biometric assessment determined from analysis of the derived physiological movement signals and / or (b) presence detection determined from analysis of at least a portion of the derived physiological movement signals. Claim 35: 35. The processor-readable medium of any one of claims 1 to 34, further comprising processor-executable instructions configured to cause the interactive audio device to acoustically detect the presence of user movement, audibly address the user, audibly authenticate the user, and communicate an authorization or alarm message about the user. Claim 36: 36. The processor-readable medium of claim 35, wherein the processor-executable instructions for audibly authenticating the user are configured to compare sound waves of the user's words sensed by the microphone with sound waves of pre-recorded words. Claim 37: 37. The processor-readable medium of any one of claims 35-36, wherein the processor-executable instructions are configured to authenticate the user using motion-based biometric sensing. Claim 38: processor-executable instructions that cause the interactive audio device to detect a movement gesture based at least in part on the derived physiological movement signal; processor-executable instructions for generating a control signal, sending a notification, or controlling a change to the operation of (a) an automated appliance and / or (b) the interactive audio device based on the detected movement gesture; 38. The processor-readable medium of any one of claims 1 to 37, further comprising: Claim 39: 39. The processor-readable medium of claim 38, wherein the control signal for controlling changes to the operation of the interactive audio device includes activating microphone sensing to initiate interactive voice assistant operation of the interactive audio device, thereby engaging an inactive interactive voice assistant process. Claim 40: 40. The processor-readable medium of claim 39, wherein the initiation of the interactive voice assistant operation is further based on detection of a linguistic keyword by the microphone sensing. Claim 41: 41. The processor-readable medium of any one of claims 1 to 40, further comprising processor-executable instructions for causing the interactive audio device to detect a user's respiratory movement based at least in part on the derived physiological movement signal, and processor-executable instructions for generating an output cue that triggers the user to regulate their breathing. Claim 42: 42. The processor-readable medium of any one of claims 1 to 41, further comprising processor-executable instructions for causing the interactive audio device to receive communications from another interactive audio device controlled by another processing device via inaudible sound waves sensed by the microphone of the interactive audio device. Claim 43: 43. The processor-readable medium of claim 42, wherein the instructions for controlling generation of the sound signal in the vicinity of the interactive audio device adjust parameters for generating the sound signal based at least in part on the communication. Claim 44: 44. The processor-readable medium of claim 43, wherein the adjusted parameters reduce interference between the interactive audio device and the other interactive audio device. Claim 45: 45. The processor-readable medium of claim 1, wherein evaluating the portion of the derived physiological movement signal further includes detecting sleep onset or wake onset, and wherein the output based on the evaluation includes a service control signal. Claim 46: 46. The processor-readable medium of claim 45, wherein the service control signals include one or more of a lighting control, an appliance control, a volume control, a thermostat control, and a window covering control. Claim 47: 47. The processor-readable medium of any one of claims 1-46, further comprising processor-executable instructions for prompting a user to collect feedback and, in response thereto, generating advice based on at least one of a portion of the derived physiological movement signals and the feedback. Claim 48: 48. The processor-readable medium of claim 47, further comprising processor-executable instructions for determining environmental data, wherein the advice is further based on the determined environmental data. Claim 49: 49. The processor-readable medium of claim 48, further comprising processor-executable instructions for determining environmental data and for generating control signals for an environmental control system based at least in part on the determined environmental data and the derived physiological movement signals. Claim 50: 50. The processor-readable medium of any one of claims 1 to 49, further comprising instructions executable by a processor for providing a sleep improvement service, the sleep improvement service including any of (a) generating advice in response to detected sleep states and / or collected user feedback, and (b) generating control signals for controlling environmental equipment to set sleep environment conditions and / or provide sleep-related advice messages to the user. Claim 51: 51. The processor-readable medium of any one of claims 1 to 50, further comprising processor-executable instructions for detecting a gesture based at least in part on the derived physiological movement signal and for initiating microphone sensing for initiation of interactive voice assistant operation of the interactive audio device. Claim 52: 52. The processor-readable medium of any one of claims 1-51, further comprising processor-executable instructions for initiating and monitoring a nap session based at least in part on the derived physiological movement signals. Claim 53: 53. The processor-readable medium of any one of claims 1 to 52, further comprising processor-executable instructions for varying a detection range by varying at least some parameters of the generated sound signal to track a user's movement through the vicinity. Claim 54: 54. The processor-readable medium of any one of claims 1 to 53, further comprising processor-executable instructions for detecting unauthorized movement using at least a portion of the derived physiological movement signal, and processor-executable instructions for generating an alarm or communication for notification to the user or a third party. Claim 55: 55. The processor-readable medium of claim 54, wherein the communication provides an alarm on a smartphone or smartwatch. Claim 56: 56. A server having access to the processor-readable medium of any one of claims 1 to 55, the server being configured to receive instructions to download the processor-executable instructions of the processor-readable medium over a network to an interactive audio device. Claim 57: 56. An interactive audio device, comprising: one or more processors; a speaker connected to the one or more processors; a microphone connected to the one or more processors; and a processor-readable medium according to any one of claims 1 to 55, wherein the one or more processors are configured to access instructions executable by the processors using a server according to claim 56. Claim 58: 58. The interactive audio device of claim 57, wherein the interactive audio device is a portable device. Claim 59: 59. The interactive audio device according to claim 57, wherein the interactive audio device comprises a mobile phone, a smart watch, a tablet computer, or a smart speaker. Claim 60: 56. A method of a server having access to the processor-readable medium of any one of claims 1 to 55, the method comprising: receiving at the server a request to download the processor-executable instructions of the processor-readable medium to an interactive audio device over a network; and transmitting the processor-executable instructions to the interactive audio device in response to the request. Claim 61: 1. A method of a processor of an interactive audio device, comprising: accessing the processor-readable medium of any one of claims 1 to 55 by a processor; executing, on the processor, the processor-executable instructions of the processor-readable medium; A method comprising: Claim 62: 1. A method for detecting physiological movements of a user of an interactive audio device, comprising: generating an audio signal in the vicinity of the interactive audio device via a speaker connected to the interactive audio device; sensing reflected audio signals from said vicinity via a microphone connected to an interactive audio device; deriving, in a processor, a physiological movement signal using a signal indicative of at least a portion of the sensed reflected sound signal and at least a portion of the audio signal; generating an output from a processor based on evaluating at least a portion of the derived physiological movement signals; A method comprising: Claim 63: 63. The method of claim 62, wherein at least a portion of the generated audio signal is substantially in an inaudible range. Claim 64: 64. The method of claim 63, wherein some of the generated audio signals are low frequency ultrasonic acoustic signals. Claim 65: 65. The method of any one of claims 62 to 64, wherein the signal indicative of a portion of the audio signal comprises an internally generated oscillator signal or a direct path measurement signal. Claim 66: 66. The method of any one of claims 62 to 65, wherein the derived physiological motion signals include one or more of respiratory motion, whole body motion or cardiac motion. Claim 67: 67. The method of any one of claims 62 to 66, comprising deriving the physiological movement signal by multiplying an oscillator signal by a portion of the sensed reflected sound signal. Claim 68: 68. The method of any one of claims 62 to 67, further comprising evaluating the sensed audible verbal communication in a processor via the microphone connected to the interactive audio device, and wherein generating the output is performed in response to the sensed audible verbal communication. Claim 69: 69. The method of any one of claims 62 to 68, wherein deriving the physiological movement signal comprises demodulating a portion of the sensed reflected sound signal using at least a portion of the audio signal. Claim 70: 70. The method of claim 69, wherein the demodulation comprises multiplying a portion of the audio signal by a portion of the sensed reflected sound signal. Claim 71: 71. The method according to claim 62, wherein the generating of the audio signal comprises generating a dual-tone frequency modulated continuous wave signal. Claim 72: 72. The method of any one of claims 62 to 71, wherein the dual-tone frequency modulated continuous wave signal comprises a first sawtooth frequency variation superimposed with a second sawtooth frequency variation in a repeating waveform. Claim 73: 73. The method of any one of claims 62 to 72, further comprising generating an ultra-wideband (UWB) audio signal as audible white noise, and wherein user movement is detected using the UWB audio signal. Claim 74: 74. The method of any one of claims 62 to 73, further comprising generating a probing sound sequence from the speaker in a setup process to calibrate distance measurements of low frequency ultrasonic echoes. Claim 75: 75. The method of any one of claims 62 to 74, further comprising generating time-synchronized calibration acoustic signals from one or more speakers, including the speaker, in a setup process to estimate a distance between the microphone and another microphone of the interactive audio device. Claim 76: 76. The method of any one of claims 62 to 75, further comprising activating a beamforming process to further localize the detected region. Claim 77: 77. The method of any one of claims 62 to 76, wherein the output generated based on evaluation of the portion of the derived physiological movement signal comprises monitored user sleep information. Claim 78: 78. A method according to any one of claims 62 to 77, wherein evaluating the portion of the derived physiological movement signal comprises detecting one or more physiological parameters. Claim 79: 79. The method of claim 78, wherein the one or more physiological parameters comprise any one or more of respiratory rate, relative respiratory amplitude, heart rate, relative cardiac amplitude, cardiac amplitude, and heart rate variability. Claim 80: 80. The method of any one of claims 77 to 79, wherein the monitored user sleep information includes any of a sleep score, a sleep stage, and time in a sleep stage. Claim 81: 81. The method of any one of claims 62 to 80, wherein the generated output comprises an interactive query answer presentation. Claim 82: 82. The method of claim 81, wherein the generated interactive query answer presentation is performed via a speaker. Claim 83: 83. The method of claim 82, wherein the generated interactive query response presentation includes advice for improving monitored user sleep information. Claim 84: 84. The method of any one of claims 62 to 83, wherein the generated output based on evaluating the portion of the derived physiological movement signal is further based on accessing a server on a network and / or performing a search of a network resource. Claim 85: 85. The method of claim 84, wherein the search is based on recorded user data. Claim 86: 86. The method of any one of claims 62 to 85, wherein the output generated based on evaluation of the portion of the derived physiological movement signal comprises a control signal for control of an automated device or system. Claim 87: 87. The method of claim 86, further comprising transmitting the control signal to the automated equipment or system over a network. Claim 88: 88. The method of any one of claims 62 to 87, further comprising generating, within the processor, a control signal for modifying a setting of the interactive audio device based at least in part on an evaluation of the derived physiological movement signals. Claim 89: 90. The method of claim 88, wherein the control signals for changing settings of the interactive audio device include volume changes based on detection of a user distance, a user state, or a user location from the interactive audio device. Claim 90: 90. The method of any one of claims 62 to 89, further comprising: performing in the processor an assessment of movement characteristics of different acoustic sensing ranges to monitor sleep characteristics of a plurality of users. Claim 91: 91. The method of claim 90, further comprising controlling generation of simultaneous acoustic sensing signals at different sensing frequencies for sensing different users at different frequencies. Claim 92: 92. The method of claim 91, further comprising controlling generation of interleaved acoustic sensing signals for sensing different users at different times. Claim 93: 93. The method of any one of claims 62 to 92, further comprising detecting the presence or absence of a user based at least in part on the derived physiological movement signal. Claim 94: 94. The method of any one of claims 62 to 93, further comprising performing, by the processor, biometric recognition of the user based at least in part on the derived physiological movement signals. Claim 95: 95. The method of any one of claims 62-94, further comprising causing the processor to generate communications over a network based on (a) a biometric assessment determined from analysis of the derived physiological movement signals and / or (b) presence detection determined from analysis of at least a portion of the derived physiological movement signals. Claim 96: 96. The method of any one of claims 62 to 95, further comprising acoustically detecting the presence of user movement, audibly addressing the user, audibly authenticating the user, and authorizing the user or communicating an alarm message about the user, by the interactive audio device. Claim 97: 97. The method of claim 96, wherein audibly authenticating the user includes comparing sound waves of the user's words sensed by the microphone with sound waves of pre-recorded words. Claim 98: 98. The method of any one of claims 96-97, further comprising authenticating the user through motion-based biometric sensing. Claim 99: detecting, by the interactive audio device, a movement gesture based at least in part on the derived physiological movement signal; generating, within a processor of the interactive audio device, a control signal for sending a notification or for controlling a change to (a) an automated device and / or (b) the operation of the interactive audio device based on the detected movement gesture; The method of any one of claims 62 to 98, further comprising: Claim 100: 100. The method of claim 99, wherein the control signal for controlling changes to the operation of the interactive audio device includes activating microphone sensing to initiate interactive voice assistant operation of the interactive audio device, thereby engaging an inactive interactive voice assistant process. Claim 101: 101. The method of claim 100, wherein initiating the interactive voice assistant action is further based on detecting a language keyword by the microphone sensing. Claim 102: 102. The method of any one of claims 62 to 101, further comprising detecting a user's respiratory movement based at least in part on the derived physiological movement signal, and causing the interactive audio device to generate output cues that trigger the user to regulate their breathing. Claim 103: 103. The method of any one of claims 62 to 102, further comprising receiving, by the interactive audio device, a communication from another interactive audio device controlled by another processing device via inaudible sound waves sensed by the microphone of the interactive audio device. Claim 104: 104. The method of claim 103, further comprising adjusting parameters for generating at least a portion of the sound signal in a vicinity of the interactive audio device based on the communication. Claim 105: 105. The method of claim 104, wherein the parameters for adjusting reduce interference between the interactive audio device and other interactive audio devices. Claim 106: 106. The method of any one of claims 62 to 105, wherein evaluating the portion of the derived physiological movement signal further comprises detecting sleep onset or wake onset, and wherein the output based on the evaluation comprises a service control signal. Claim 107: 107. The method of claim 106, wherein the service control signals include one or more of a lighting control, an appliance control, a volume control, a thermostat control, and a window covering control. Claim 108: 108. The method of any one of claims 62 to 107, further comprising prompting a user to collect feedback, and in response generating advice based on at least one of a portion of the derived physiological movement signals and the feedback. Claim 109: 109. The method of claim 108, further comprising determining environmental data, wherein the advice is further based on the determined environmental data. Claim 110: 109. The method of claim 108, further comprising determining environmental data and generating a control signal for an environmental control system based at least in part on the determined environmental data and the derived physiological movement signal. Claim 111: A method according to any one of claims 62 to 110, further comprising controlling the provision of a sleep improvement service, the sleep improvement service comprising any one of (a) generating advice in response to the detected sleep state and / or collected user feedback, and (b) generating control signals for controlling environmental equipment to set sleep environment conditions and / or provide sleep-related advice messages to the user. Claim 112: 112. The method of any one of claims 62 to 111, further comprising: detecting a gesture based on the derived physiological movement signal; and initiating microphone sensing for initiation of an interactive voice assistant operation of the interactive audio device. Claim 113: 113. The method of any one of claims 62 to 112, further comprising initiating and monitoring a nap session based at least in part on the derived physiological movement signal. Claim 114: 114. The method of any one of claims 62 to 113, further comprising varying a detection range by varying parameters of at least some of the generated sound signals to track a user's movements through the vicinity. Claim 115: 115. The method of any one of claims 62 to 114, further comprising detecting unauthorized movement using at least a portion of the derived physiological movement signal, and generating an alarm or communication for notification to the user or a third party. Claim 116: 116. The method of claim 115, wherein the communication provides an alarm on a smartphone or smartwatch.
Claims
1. 1. A processor-readable medium having stored thereon processor-executable instructions, the processor-executable instructions, when executed by a processor of an interactive audio device, causing the processor to detect a physiological movement of a user, the processor-executable instructions comprising: instructions for controlling transmission of a radio frequency (RF) signal into a vicinity of the interactive audio device via a transmitter coupled to the interactive audio device; instructions for controlling sensing of reflections of the transmitted signal via a receiver connected to the interactive audio device; instructions for deriving a physiological movement signal using a signal indicative of at least a portion of the sensed reflected signal and at least a portion of the transmitted signal; instructions for generating an output based on an evaluation of at least a portion of the derived physiological movement signal; processor-executable instructions for evaluating audible verbal communications sensed via a microphone coupled to the interactive audio device, the instructions for generating the output configured to generate the output in response to the sensed audible verbal communications; wherein the generated output comprises an interactive query and response presentation.
2. The processor-readable medium of claim 1 , wherein the derived physiological motion signals include one or more of respiratory motion, whole body motion, and cardiac motion.
3. The processor-readable medium of any one of claims 1 to 2, further comprising processor-executable instructions for activating a beamforming process to further localize the detected region.
4. 4. The processor-readable medium of claim 1, wherein the generated output based on evaluation of the portion of the derived physiological movement signal includes monitored user sleep information.
5. 5. The processor-readable medium of claim 1, wherein evaluating the portion of the derived physiological movement signal comprises detecting one or more physiological parameters including any one or more of respiration rate, relative respiration amplitude, heart rate, cardiac amplitude, relative cardiac amplitude, and heart rate variability.
6. (a) the generated output includes a presentation of the interactive query and response performed via a speaker; and / or (b) the generated output includes advice for improving the user's sleep; The processor-readable medium of any one of claims 1 to 5.
7. (a) the generated output based on evaluating the portion of the derived physiological movement signal is further based on accessing a server on a network and / or performing a search of a network resource; and / or (b) the generated output based on evaluation of the portion of the derived physiological movement signal includes a control signal for control of an automated device or system, and the processor-readable medium further includes processor control instructions for transmitting the control signal to the automated device or system over a network. The processor-readable medium of any one of claims 1 to 6.
8. A processor-readable medium as described in any one of claims 1 to 6, further comprising processor control instructions for generating control signals for changing settings of the interactive audio device based on an evaluation of at least a portion of the derived physiological movement signals, wherein the control signals for changing settings of the interactive audio device include a volume change based on detection of a user distance from the interactive audio device, or a user state, or a user position.
9. 9. The processor-readable medium of claim 1, wherein the processor-executable instructions are configured to evaluate movement characteristics of different sensing ranges for monitoring sleep characteristics of multiple users.
10. 10. The processor-readable medium of claim 1, further comprising processor-executable instructions for detecting the presence or absence of a user based at least in part on the derived physiological movement signal.
11. 11. The processor-readable medium of claim 1, further comprising processor-executable instructions for performing biometric recognition of a user based at least in part on the derived physiological movement signals.
12. 12. The processor-readable medium of claim 1, further comprising processor-executable instructions for generating communication over a network based on (a) a biometric assessment determined from analysis of the derived physiological movement signals and / or (b) presence detection determined from analysis of at least a portion of the derived physiological movement signals.
13. 13. The processor-readable medium of any one of claims 1 to 12, further comprising processor-executable instructions configured to cause the interactive audio device to detect the presence of user movement, audibly address the user, audibly authenticate the user, and communicate an authorization or alarm message about the user.
14. A processor-readable medium as described in claim 13, wherein the instructions executable by the processor for audibly authenticating the user are configured to compare sound waves of the user's words sensed by a microphone connected to the processor with sound waves of pre-recorded words.
15. A processor-readable medium as described in claim 13 or 14, wherein the processor-executable instructions are configured to authenticate the user using motion-based biometric sensing.
16. processor-executable instructions that cause the interactive audio device to detect a movement gesture based at least in part on the derived physiological movement signal; processor-executable instructions for generating a control signal or sending a notification based on the detected movement gesture, or for controlling a change to the operation of (a) an automated appliance and / or (b) the interactive audio device; The processor-readable medium of any one of claims 1 to 15, further comprising:
17. A processor-readable medium as described in claim 16, wherein the control signal for controlling changes to the operation of the interactive audio device includes activating microphone sensing to initiate interactive voice assistant operation of the interactive audio device, thereby engaging an unactivated interactive voice assistant process, and wherein the initiation of the interactive voice assistant operation is further based on detection of a language keyword by the microphone sensing.
18. 18. The processor-readable medium of any one of claims 1 to 17, further comprising processor-executable instructions for causing the interactive audio device to receive communications from another interactive audio device controlled by another processing device via inaudible sound waves sensed by a microphone of the interactive audio device.
19. A processor-readable medium as described in any one of claims 1 to 18, wherein evaluation of the portion of the derived physiological movement signals further includes detection of sleep onset or wake onset, and the output based on the evaluation includes a service control signal, wherein the service control signal includes one or more of lighting control, appliance control, volume control, thermostat control, and window covering control.
20. 20. The processor-readable medium of claim 1, further comprising instructions executable by a processor for providing a sleep enhancement service, the sleep enhancement service including: (a) generating advice in response to detected sleep states and / or collected user feedback; and (b) generating control signals for controlling environmental equipment to set sleep environment conditions and / or provide sleep-related advice messages to the user.
21. 21. The processor-readable medium of claim 1, further comprising processor-executable instructions for detecting a gesture based at least in part on the derived physiological movement signal and for initiating microphone sensing for initiation of an interactive voice assistant operation of the interactive audio device.
22. 22. The processor-readable medium of any one of claims 1 to 21, further comprising processor-executable instructions for initiating and monitoring a nap session based at least in part on the derived physiological movement signals.
23. 23. The processor-readable medium of any one of claims 1 to 22, further comprising processor-executable instructions for varying a detection range by varying parameters of at least some of the transmitted RF signals to track a user's movement through the vicinity.
24. A processor-readable medium as described in any one of claims 1 to 23, further comprising instructions executable by a processor to detect unauthorized movement using at least a portion of the derived physiological movement signals and generate an alarm or communication for notification to the user or a third party, wherein the communication provides an alarm on a smartphone or smartwatch.
25. 25. A server having access to the processor-readable medium of any one of claims 1 to 24, the server being configured to receive instructions to download the processor-executable instructions of the processor-readable medium over a network to an interactive audio device.
26. An interactive audio device configured to include one or more processors, a speaker connected to the one or more processors, a microphone connected to the one or more processors, and a medium readable by the processor according to any one of claims 1 to 24.
27. 27. The interactive audio device of claim 26, wherein the interactive audio device comprises one of a mobile phone, a smart watch, a tablet computer, and a smart speaker.
28. 25. A method of a server having access to the processor-readable medium of any one of claims 1 to 24, the method comprising: receiving at the server a request to download the processor-executable instructions of the processor-readable medium to an interactive audio device over a network; and transmitting the processor-executable instructions to the interactive audio device in response to the request.
Citation Information
Patent Citations
Elastic wave detector, personal authentication device, voice output device, elastic wave detection program, and personal authentication program
JP2012239748A
Information exchange system, information exchange method and program
JP2015041372A
Authentication server, voiceprint authentication system and voiceprint authentication method
JP2016136299A
Sleep management method and system
JP2016532481A
Mobile terminal and controlling method thereof
US20170207859A1