Contactless monitoring of respiratory rate and respiratory loss using facial video

By combining facial video with motion and rPPG methods and using a machine learning model to select respiratory rate monitoring, the problem of respiratory rate monitoring equipment requiring contact with the skin and high cost of high-end equipment in the existing technology is solved, and non-contact and accurate respiratory rate monitoring is achieved.

CN120603535APending Publication Date: 2025-09-05SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480008351.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-05
Filing Date
2024-02-13
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In existing technologies, respiratory rate monitoring equipment requires direct contact with the skin and is not suitable for people with sensitive skin. Wireless signal measurement is also limited. High-end infrared and depth camera equipment is expensive, the resolution of cameras available to consumers is low, and camera-based chest movement signal extraction is difficult.

Method used

By utilizing facial videos, combined with motion-based and remote photoplethysmography (rPPG) methods, respiratory rate (RR) monitoring is selected through a machine learning model. Facial videos are captured using a camera, facial feature points are identified, and RR estimation is performed by combining motion and color changes to overcome the limitations of the respective methods.

Benefits of technology

It achieves non-contact, accurate respiratory rate and respiratory loss monitoring, is suitable for various environments, reduces equipment costs, and is suitable for consumer electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120603535A_ABST
    Figure CN120603535A_ABST
Patent Text Reader

Abstract

A method includes acquiring a video using a camera. The method further includes determining a motion-based respiratory rate (RR) and a motion-based respiratory signal based on a face of the person, the face of the person being identified based on the video. The method further includes determining a remote photoplethysmography (rPPG)-based RR and an rPPG-based respiratory signal based on a face of the person, the face of the person being identified based on the video. The method further includes selecting one of the motion-based RR or the rPPG-based RR by inputting the motion-based respiratory signal and the rPPG-based respiratory signal as inputs into a trained machine learning model. In addition, the method includes presenting a selected one of a motion-based RR or an rPPG-based RR based on the prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to machine learning systems and processes, and more particularly to contactless monitoring of respiratory rate and breathing absence using facial video. Background Art

[0002] Respiratory rate (RR) is a vital sign that indicates overall respiratory system function and health. Among other things, it is a reliable predictor of intensive care admission or death. It is also valuable information for patient care, particularly for those with asthma, congestive heart failure, cardiac arrest, and respiratory distress due to infection. Furthermore, respiratory rate information can be useful in understanding fatigue, emotional state, or exercise progress. Summary of the Invention

[0003] Solution to the problem

[0004] The present disclosure relates to contactless monitoring of respiration rate and lack of respiration using facial video.

[0005] In a first embodiment, a method includes acquiring a video using a camera. The method also includes determining a motion-based respiration rate (RR) and a motion-based respiration signal based on a person's face, the person's face being identified based on the video. The method also includes determining a remote photoplethysmography (rPPG)-based RR and an rPPG-based respiration signal based on the person's face, the person's face being identified based on the video. The method also includes selecting one of the motion-based RR or the rPPG-based RR by inputting the motion-based respiration signal and the rPPG-based respiration signal into a trained machine learning model. Furthermore, the method includes presenting the selected one of the motion-based RR or the rPPG-based RR based on a prediction.

[0006] In a second embodiment, an electronic device includes a camera. The electronic device also includes at least one processing device. The electronic device also includes a memory storing instructions. The instructions, when executed by at least a portion of the at least one processing device, cause the electronic device to capture a video using the camera. The instructions, when executed by at least a portion of the at least one processing device, cause the electronic device to determine a motion-based RR and a motion-based respiration signal based on a person's face, the person's face being identified based on the video. The instructions, when executed by at least a portion of the at least one processing device, cause the electronic device to determine an rPPG-based RR and an rPPG-based respiration signal based on the person's face being identified based on the video. The instructions, when executed by at least a portion of the at least one processing device, cause the electronic device to select one of the motion-based RR and the rPPG-based RR by inputting the motion-based respiration signal and the rPPG-based respiration signal into a trained machine learning model. The electronic device also includes a memory storing instructions. The instructions, when executed by at least a portion of the at least one processing device, cause the electronic device to present the selected one of the motion-based RR and the rPPG-based RR based on a prediction.

[0007] In a third embodiment, a non-transitory machine-readable medium contains instructions that, when executed, cause an electronic device to capture a video using a camera. The non-transitory machine-readable medium also contains instructions that, when executed, cause the electronic device to determine a motion-based RR and a motion-based respiration signal based on a person's face, the person's face being identified based on the video. The non-transitory machine-readable medium also contains instructions that, when executed, cause the electronic device to determine an rPPG-based RR and an rPPG-based respiration signal based on the person's face being identified based on the video. The non-transitory machine-readable medium also contains instructions that, when executed, cause the electronic device to select one of a motion-based RR or an rPPG-based RR by inputting the motion-based respiration signal and the rPPG-based respiration signal into a trained machine learning model. Furthermore, the non-transitory machine-readable medium contains instructions that, when executed, cause the electronic device to present the selected one of the motion-based RR or the rPPG-based RR based on a prediction.

[0008] Other technical features may be apparent to those skilled in the art from the following drawings, descriptions, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which like reference numerals represent like parts:

[0010] Figure 1An example network configuration including electronic devices according to the present disclosure is shown;

[0011] Figure 2 An example process for contactless monitoring of respiration rate using facial video according to the present disclosure is shown;

[0012] Figure 3 shows an example video frame in which facial regions have been identified according to the present disclosure;

[0013] Figure 4A and Figure 4B shows an example graph illustrating feature extraction from a motion-based breathing signal and an rPPG-based breathing signal for machine learning-based breathing rate selection in accordance with the present disclosure;

[0014] Figure 5 An example process for detection of lack of breathing using facial video according to the present disclosure is shown;

[0015] Figure 6 An example method for contactless monitoring of respiration rate and lack of respiration using facial video according to the present disclosure is shown;

[0016] Figure 7 An example process for contactless monitoring of respiration rate using facial video according to the present disclosure is shown; and

[0017] Figure 8 An example process for contactless monitoring of respiration rate using facial video according to the present disclosure is shown. DETAILED DESCRIPTION

[0018] It may be helpful to set forth definitions of certain words and phrases used throughout this patent document. The terms "send," "receive," and "communicate," and their derivatives, encompass both direct and indirect communications. The terms "include," "comprise," and their derivatives, mean to include, but are not limited to. The term "or" is inclusive, meaning and / or. The phrase "associated with," and its derivatives, means to include, be included within, be interconnected with, contain, be contained within, be connected to or connected with, be coupled to or coupled with, be communicable with, cooperate with, be interwoven, be juxtaposed, be proximate to, be bound to or bound with, have, have the property of, have a relationship to or with, and the like.

[0019] Furthermore, the various functions described below may be implemented or supported by one or more computer programs, each of which is formed of computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, related data, or portions thereof, suitable for implementation in suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as read-only memory (ROM), random-access memory (RAM), hard drives, compact disks (CDs), digital video disks (DVDs), or any other type of memory. "Non-transitory" computer-readable media excludes wired, wireless, optical, or other communication links that transmit transitory electrical or other signals. Non-transitory computer-readable media includes media in which data can be permanently stored, as well as media in which data can be stored and later rewritten, such as rewritable optical disks or erasable memory devices.

[0020] As used herein, terms and phrases such as "having", "may have", "include", or "may include" a feature (such as a number, function, operation, or component such as a part) indicate the presence of the feature and do not exclude the presence of other features. In addition, as used herein, the phrases "A or B", "at least one of A and / or B", or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and B", and "at least one of A or B" may indicate all of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. In addition, as used herein, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate user devices that are different from each other, regardless of the order or importance of the devices. Without departing from the scope of the present disclosure, a first component may be represented as a second component, and vice versa.

[0021] It will be understood that when an element (such as a first element) is referred to as being (operably or communicatively) “coupled” / “coupled to” another element (such as a second element) or “connected” / “connected to” another element (such as a second element), it may be coupled or connected to / coupled or connected to the other element directly or via a third element. In contrast, it will be understood that when an element (such as a first element) is referred to as being “directly coupled” / “directly coupled to” another element (such as a second element) or “directly connected” / “directly connected to” another element (such as a second element), no other elements (such as a third element) are interposed between the element and the other element.

[0022] As used herein, the phrase "configured (or configured) to" may be used interchangeably with the phrases "suitable for," "capable of," "designed to," "adapted to," "manufactured to," or "capable of," depending on the circumstances. The phrase "configured (or configured) to" does not inherently mean "designed specifically in hardware to." Rather, the phrase "configured to" may mean that a device can perform an operation together with another device or part. For example, the phrase "a processor configured (or configured) to perform A, B, and C" may mean a general-purpose processor (such as a CPU or application processor) that can perform operations by executing one or more software programs stored in a memory device, or a dedicated processor (such as an embedded processor) for performing operations.

[0023] The terms and phrases used herein are provided only to describe some embodiments of the present disclosure, but do not limit the scope of other embodiments of the present disclosure. It will be understood that, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include plural references. All terms and phrases used herein (including technical and scientific terms and phrases) have the same meaning as those generally understood by those of ordinary skill in the art to which the embodiments of the present disclosure belong. It will be further understood that terms and phrases (such as those defined in commonly used dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and will not be interpreted in an idealized or overly formal sense unless explicitly defined as such herein. In some cases, the terms and phrases defined herein may be interpreted to exclude embodiments of the present disclosure.

[0024] Examples of "electronic devices" according to embodiments of the present disclosure may include at least one of the following: a smartphone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head-mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, or a smart watch). Other examples of electronic devices include smart home appliances. Examples of smart home appliances may include at least one of the following: a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave oven, a washing machine, a dryer, an air purifier, a set-top box, a home automation control panel, a security control panel, a TV box (such as SAMSUNG HOMESYNC, APPLETV, or GOOGLE TV), a speaker or smart speaker with an integrated digital assistant (such as SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZONECHO), a game console (such as XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camera, or an electronic photo frame. Other examples of electronic devices include at least one of the following: various medical devices (such as various portable medical measuring devices (such as blood glucose measuring devices, heart rate measuring devices, or body temperature measuring devices), magnetic resource angiography (MRA) devices, magnetic resource imaging (MRI) devices, computed tomography (CT) devices, imaging devices, or ultrasound devices), navigation devices, global positioning system (GPS) receivers, event data recorders (EDRs), flight data recorders (FDRs), car infotainment devices, navigation electronic devices (such as navigation navigation devices or gyrocompasses), avionics equipment, security equipment, vehicle-mounted head units, industrial or household robots, automated teller machines (ATMs), point-of-sale (POS) devices, or Internet of Things (IoT) devices (such as light bulbs, various sensors, electricity meters or gas meters, sprinklers, fire alarms, thermostats, streetlights, ovens, fitness equipment, hot water tanks, heaters, or boilers). Other examples of electronic devices include at least a portion of a piece of furniture or a building / structure, an electronic board, an electronic signature receiving device, a projector, or various measuring devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). Note that according to various embodiments of the present disclosure, the electronic device may be one or a combination of the devices listed above. According to some embodiments of the present disclosure, the electronic device may be a flexible electronic device. The electronic devices disclosed herein are not limited to the devices listed above and may include new electronic devices depending on technological developments.

[0025] In the following description, an electronic device is described with reference to the accompanying drawings according to various embodiments of the present disclosure. As used herein, the term "user" may refer to a person using an electronic device or another device (such as an artificial intelligence electronic device).

[0026] Definitions for certain other words and phrases may be provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most instances, such definitions apply to prior, as well as future uses of such defined words and phrases.

[0027] None of the description in this application should be read as implying that any particular element, step, or function is essential to be included in the claims scope.

[0028] The following discussion is described with reference to the accompanying drawings. Figures 1 to 8 However, it should be understood that the present disclosure is not limited to these embodiments, and all changes and / or equivalents or replacements thereto also fall within the scope of the present disclosure.

[0029] As discussed above, respiratory rate (RR) is a vital sign that indicates overall respiratory system function and health. Among other things, respiratory rate is a reliable predictor of intensive care admission or death. It is also valuable information for patient care, particularly for those with asthma, congestive heart failure, cardiac arrest, and respiratory distress due to infection. Furthermore, respiratory rate information can be useful in understanding fatigue, emotional state, or exercise progress.

[0030] Many conventional RR monitoring devices require direct contact with human skin. It is desirable for wearable sensors to be attached directly to or in contact with an individual's body (such as the face, torso, wrist, or finger). Available commercial devices for respiratory monitoring include chest straps, smartwatches, face masks, pulse oximeters, nasal sensors, and wristbands. Chest straps measure chest movement using capacitive sensors. Optical sensors on smartwatches or pulse oximeters can measure RR based on photoplethysmography (PPG) and / or electrocardiography (ECG). Recently, inertial measurement unit (IMU) sensors on earbuds have been used to measure RR. However, contact-based measurements are not suitable for people with sensitive skin, such as premature newborns and the elderly. They are also cumbersome for patients who need to wear on-body sensors for long-term monitoring. Furthermore, sharing contaminated sensors poses an extreme risk of disease transmission in hospitals and assisted living facilities.

[0031] Non-contact RR measurements can be obtained using wireless signals (such as acoustic or radio frequency signals). For example, a person's breathing state can be identified by using continuously propagating waves, which are affected by repetitive chest movements during breathing. As a specific example, ultra-wideband (UWB) radar-based systems have been used to detect the breathing patterns of multiple individuals. However, using wireless signals to estimate RR generally has limitations. For example, the signal transmitter must be positioned close to the person, and measurements are primarily optimized for indoor settings.

[0032] Camera-based respiratory monitoring is gaining increasing interest as a contactless method and is being developed to take advantage of recent advances in camera and image processing technologies. Infrared thermal imaging (also known as thermography) is one method of camera-based respiratory monitoring. Infrared thermal imaging captures the radiation naturally emitted from human skin. Some studies have used far-infrared (FIR) cameras to extract respiratory signs by observing changes in thermal airflow at a person's nostrils. Furthermore, depth cameras can be used to estimate respiratory rate during sleep by recording chest movements. Both infrared and depth cameras require no light source, but they are high-end and prohibitively expensive. Consumer-accessible cameras suffer from low pixel resolution and low sampling rates and are generally not available on personal consumer-grade devices.

[0033] Visually capturing the respiration-induced motion of a person's chest cavity is another direct method for observing respiratory status. Various camera-based RR estimation methods attempt to obtain motion signals from the chest region. However, the chest region is not always available in facial videos. Extracting chest motion signals from videos is challenging because the chest region lacks unique feature points to discern when covered by various clothing items. Therefore, chest recognition typically relies on facial detection.

[0034] In addition to RR estimation, the detection of apnea is an important feature for monitoring respiratory activity. Apnea is a pause in the respiratory rhythm, and there are two types of sleep apnea. Obstructive sleep apnea occurs when the upper airway is obstructed, while central sleep apnea occurs when there is a loss of respiratory motor output from the brainstem. The main difference between these two types of apneic events is that obstructive sleep apnea persists with respiratory movements of the trunk. In contrast, central sleep apnea does not involve any respiratory movements. The human head and neck system, which is biomechanically connected to the trunk, is also affected by respiratory movements. When respiratory-induced trunk movements are reduced, unconstrained head movements, which are a sequence of trunk movements, are also reduced. Therefore, both types of apnea can be observed by the reduction or cessation of respiratory-induced head movements.

[0035] The present disclosure provides various techniques for contactless monitoring of respiratory rate and lack of breathing using facial video. As described in more detail below, the disclosed embodiments can determine motion-based RR based on a video of a person's face captured using a camera. The disclosed embodiments can also determine remote photoplethysmography (rPPG)-based RR based on a video of a person's face. A pre-trained machine learning model can select between motion-based RR or rPPG-based RR to maintain accuracy in various measurement scenarios. Note that while some of the embodiments discussed below are described in the context of use in consumer electronic devices (such as smartphones), this is merely an example. It will be understood that the principles of the present disclosure can be implemented in any number of other suitable contexts and using any suitable device.

[0036] Figure 1 An example network configuration 100 including electronic devices according to the present disclosure is shown. Figure 1 The embodiment of the network configuration 100 shown is for illustration only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.

[0037] According to an embodiment of the present disclosure, an electronic device 101 is included in a network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, or a sensor 180. In some embodiments, the electronic device 101 may exclude at least one of these components, or may add at least one other component. The bus 110 includes circuitry for connecting the components 120-180 to each other and for transmitting communications (such as control messages and / or data) between the components.

[0038] Processor 120 includes one or more processing devices, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). In some embodiments, processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), a graphics processor unit (GPU), or a neural processing unit (NPU). Processor 120 is capable of controlling at least one of the other components of electronic device 101 and / or performing operations or data processing related to communication or other functions. As described in more detail below, processor 120 can perform one or more operations for contactless monitoring of respiratory rate and apnea using facial video.

[0039] Memory 130 may include volatile and / or non-volatile memory. For example, memory 130 may store commands or data related to at least one other component of electronic device 101. According to an embodiment of the present disclosure, memory 130 may store software and / or programs 140. Programs 140 include, for example, a kernel 141, middleware 143, application programming interfaces (APIs) 145, and / or application programs (or "apps") 147. At least a portion of kernel 141, middleware 143, or API 145 may be represented as an operating system (OS).

[0040] The kernel 141 can control or manage system resources (such as the bus 110, processor 120, or memory 130) used to perform operations or functions implemented in other programs (such as middleware 143, API 145, or application 147). The kernel 141 provides an interface that allows the middleware 143, API 145, or application 147 to access various components of the electronic device 101 to control or manage system resources. Application 147 can support one or more functions for contactless monitoring of respiratory rate and apnea using facial video, as discussed below. These functions can be performed by a single application or by multiple applications, each performing one or more of these functions. For example, the middleware 143 can act as a relay to allow the API 145 or application 147 to communicate data with the kernel 141. Multiple applications 147 can be provided. The middleware 143 can control work requests received from the applications 147, such as by assigning priority to the use of the electronic device 101's system resources (such as the bus 110, processor 120, or memory 130) among at least one of the multiple applications 147. The API 145 is an interface that allows the application 147 to control functions provided from the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (such as a command) for file control, window control, image processing, or text control.

[0041] The I / O interface 150 serves as an interface that can, for example, transmit commands or data input from a user or other external device to other component(s) of the electronic device 101. The I / O interface 150 can also output commands or data received from other component(s) of the electronic device 101 to the user or other external devices.

[0042] Display 160 includes, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a quantum dot light-emitting diode (QLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. Display 160 may also be a depth-sensing display, such as a multifocal display. Display 160 is capable of displaying various content (such as text, images, videos, icons, or symbols) to the user. Display 160 may include a touch screen and may receive touch, gesture, proximity, or hovering input, for example, using an electronic pen or a part of the user's body.

[0043] For example, the communication interface 170 can establish communication between the electronic device 101 and an external electronic device (such as the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 can connect to the network 162 or 164 via wireless or wired communication to communicate with the external electronic device. The communication interface 170 can be a wired or wireless transceiver or any other component for transmitting and receiving signals.

[0044] Wireless communication can use, for example, at least one of WiFi, Long Term Evolution (LTE), Long Term Evolution-Advanced (LTE-A), fifth-generation wireless systems (5G), millimeter wave or 60 GHz wireless communication, wireless USB, code division multiple access (CDMA), wideband code division multiple access (WCDMA), universal mobile telecommunications system (UMTS), wireless broadband (WiBro), or global system for mobile communications (GSM) as a communication protocol. Wired connections can include, for example, at least one of universal serial bus (USB), high-definition multimedia interface (HDMI), recommended standard 232 (RS-232), or plain old telephone service (POTS). Network 162 or 164 includes at least one communication network, such as a computer network (e.g., a local area network (LAN) or a wide area network (WAN)), the Internet, or a telephone network.

[0045] Electronic device 101 also includes one or more sensors 180, which can measure physical quantities or detect the activation state of electronic device 101 and convert the measured or detected information into electrical signals. For example, one or more sensors 180 may include one or more cameras or other imaging sensors for capturing images of a scene. Sensor(s) 180 may also include one or more buttons for touch input, gesture sensors, gyroscopes or gyroscopic sensors, air pressure sensors, magnetic sensors or magnetometers, acceleration sensors or accelerometers, grip sensors, proximity sensors, color sensors (such as red, green, and blue (RGB) sensors), biophysical sensors, temperature sensors, humidity sensors, illuminance sensors, ultraviolet (UV) sensors, electromyography (EMG) sensors, electroencephalography (EEG) sensors, electrocardiography (ECG) sensors, infrared (IR) sensors, ultrasonic sensors, iris sensors, or fingerprint sensors. Sensor(s) 180 may also include an inertial measurement unit (IMU), which may include one or more accelerometers, gyroscopes, and other components. Furthermore, sensor(s) 180 may include control circuitry for controlling at least one of the sensors included herein. Any of these sensor(s) 180 may be located within the electronic device 101 .

[0046] In some embodiments, the electronic device 101 may be a wearable device or a wearable device (such as an HMD) that can be mounted on an electronic device. For example, the electronic device 101 may represent an AR wearable device, such as a headset with a display panel or smart glasses. In other embodiments, the first external electronic device 102 or the second external electronic device 104 may be a wearable device or a wearable device (such as an HMD) that can be mounted on an electronic device. In those other embodiments, when the electronic device 101 is mounted on the electronic device 102 (such as the HMD), the electronic device 101 may communicate with the electronic device 102 through the communication interface 170. The electronic device 101 may be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network.

[0047] The first external electronic device 102, the second external electronic device 104, and the server 106 may each be a device of the same or different type as the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. In addition, according to certain embodiments of the present disclosure, all or some operations performed on the electronic device 101 may be performed on another one or more other electronic devices (such as the electronic devices 102 and 104 or the server 106). In addition, according to certain embodiments of the present disclosure, when the electronic device 101 should automatically or upon request perform some functions or services, the electronic device 101 may request another device (such as the electronic devices 102 and 104 or the server 106) to perform at least some functions associated therewith, instead of performing the function or service itself, or performing the function or service in addition. The other electronic devices (such as the electronic devices 102 and 104 or the server 106) are capable of performing the requested function or additional function and transmitting the result of the execution to the electronic device 101. The electronic device 101 may provide the requested function or service by processing the received result as is or in addition. To this end, for example, cloud computing, distributed computing, or client-server computing technology may be used. Although Figure 1 The electronic device 101 is shown to include a communication interface 170 to communicate with the external electronic device 104 or the server 106 via the network 162 or 164 , but according to some embodiments of the present disclosure, the electronic device 101 may operate independently without a separate communication function.

[0048] Server 106 may include components 110-180 (or a suitable subset thereof) that are the same as or similar to electronic device 101. Server 106 may support driving electronic device 101 by performing at least one of the operations (or functions) implemented on electronic device 101. For example, server 106 may include a processing module or processor that may support processor 120 implemented in electronic device 101. As described in more detail below, server 106 may perform one or more operations to support technology for contactless monitoring of respiration rate and lack of respiration using facial video.

[0049] although Figure 1 An example of a network configuration 100 including an electronic device 101 is shown, but the Figure 1 Various changes may be made. For example, network configuration 100 may include any number of each component in any suitable arrangement. In general, computing and communication systems have a wide variety of configurations, and Figure 1 The scope of this disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document may be used, but these features may be used in any other suitable system.

[0050] Figure 2 An example process 200 for contactless monitoring of respiration rate using facial video according to the present disclosure is shown. For ease of explanation, the process 200 is described as using the above-mentioned Figure 1 1. However, this is merely one example, and process 200 may be implemented using any other suitable device(s), such as server 106, and in any other suitable system(s).

[0051] like Figure 2 As shown, process 200 illustrates two camera-based methods that can be used to monitor respiratory information: for example, rPPG-based RR measurement (which tracks skin color changes) and motion-based RR measurement (which tracks body movement). First, rPPG is proportional to the amount of blood flowing through a person's blood vessels. This can be observed as subtle, transient changes in skin color seen by an RGB camera or other camera. In some cases, the rPPG signal can be derived from the temporal changes in the RGB values ​​of skin pixels in a video. A respiratory component can be extracted from pulsatile activity, as heart rate increases with inspiration and decreases with expiration, known as the respiratory sinus arrhythmia (RSA) relationship. Note that obtaining a clean rPPG signal can involve overcoming motion artifacts, various lighting spectra, and varying skin tones. Furthermore, skin tissue typically needs to be visible to the camera to collect the rPPG.

[0052] Secondly, motion-based RR can be measured by observing small, repetitive movements of the respiratory system (e.g., a person's lungs, nose, trachea, and respiratory muscles). Because RR is obtained by tracking the movement of selected pixels, detecting skin tissue in the video may not be necessary. Therefore, when, for example, a hat or mask covers a person's face, motion-based methods can better estimate RR than rPPG-based methods. Note that motion artifacts unrelated to respiration-induced motion can negatively impact the measurement accuracy of motion-derived RR estimates. Furthermore, the lack of respiratory motion can lead to inaccurate RR estimates.

[0053] The process 200 combines rPPG-based methods and motion-based methods to overcome the limitations of each modality and improve the overall performance. Thus, the process 200 provides a novel multimodal method to monitor respiratory activity using color changes and movements of the face observed by a camera. Figure 2As shown, process 200 includes a video capture operation 205 in which electronic device 101 captures a video of a person's face 210. Video capture operation 205 may be performed in response to an event, such as a user actuating a video capture control of electronic device 101. Video capture operation 205 may be performed continuously, intermittently, repeatedly, on demand for a selected period of time, or at any other suitable frequency and duration.

[0054] In some embodiments, video 210 may be an RGB video captured using one imaging sensor 180 of electronic device 101 (such as a camera with an RGB sensor). In other embodiments, video 210 may be captured using multiple imaging sensors 180 of electronic device 101. Furthermore, in some embodiments, one or more imaging sensors 180 are positioned at a distance of approximately 50 centimeters in front of a person's face, although other distances and placements are possible. Furthermore, in some embodiments, the frame rate of video 210 is 30 or 60 frames per second (fps), although other frame rates are possible and within the scope of the present disclosure.

[0055] After capturing video 210, electronic device 101 performs face and landmark detection operation 215. In operation 215, electronic device 101 searches frames of video 210, such as by starting from an initial frame, for a rectangle or other region that depicts a person's face. Any suitable technique can be used to detect a person's face, such as a deep learning face detection algorithm or the Viola-Jones algorithm. If a face region is not found in the first frame, electronic device 101 can move to successive frames until a frame with a face region is found. Figure 3 An example video frame 300 is shown in which a facial region 305 has been identified in accordance with the present disclosure. Any background next to the facial region 305 may be removed for further processing and privacy protection.

[0056] Once the facial region 305 is identified, the electronic device 101 selects a plurality of facial landmarks 315 within the facial region 305. In some embodiments, the electronic device 101 selects ten facial landmarks 315 in the forehead region of the person and seven facial landmarks 315 in the nose region of the person, although other numbers of landmarks may be used in each region. Furthermore, in some embodiments, the facial landmarks 315 may be selected from a database of predetermined facial landmarks, although facial landmarks may be identified in any other suitable manner.

[0057] The electronic device 101 also selects multiple regions of interest (ROIs) within the facial region 305 based on the selected landmark points 315. In some embodiments, the electronic device 101 selects two rectangular or other ROIs: (i) a first ROI 310 corresponding to the person's nose region and (ii) a second ROI 310 corresponding to the person's forehead region. The electronic device 101 may also select additional ROIs 310 for use in rPPG-based RR estimation. For example, the electronic device 101 may employ a Gaussian mixture model to identify skin pixels in the detected facial region 305 and use the skin likelihood scores to select multiple (e.g., 32) ROIs 310.

[0058] Once the electronic device 101 has detected the facial region 305, ROI 310, and facial landmarks 315 in the video 210, the electronic device 101 performs two separate RR estimation techniques: (i) motion-based RR estimation 220 and (ii) rPPG-based RR estimation 250. The motion-based RR estimation 220 includes a motion extraction operation 225, in which the electronic device 101 extracts facial motion signals by tracking the facial landmarks 315 over time. In some embodiments, the electronic device 101 uses a motion tracking algorithm to track the horizontal (X-axis) and vertical (Y-axis) movement of the facial landmarks 315 by detecting the X and Y coordinates of the center point of each facial landmark 315 in each frame of the video 210. In certain embodiments, the electronic device 101 only utilizes position changes along the Y-axis because a person's respiratory motion is highly correlated with vertical head movement during an upright posture. Any suitable technique can be used for motion tracking, such as the Lucas-Kanade-Tomasi (LKT) optical flow algorithm. The electronic device 101 may also use an overlapping sliding window method to estimate the RR per second. Accordingly, the motion signal may be buffered into a sliding window with a specified length (such as forty seconds) and a step size of one second.

[0059] Typically, facial motion signals can be susceptible to noise or motion artifacts due to sudden, active or inactive movements of a person during the recording of video 210. Therefore, after motion extraction operation 225, electronic device 101 performs motion artifact removal operation 230 to remove motion artifacts from the motion signal. In motion artifact removal operation 230, electronic device 101 smoothes the motion signal, such as by using a moving average. Electronic device 101 also determines a motion velocity signal by calculating the difference between consecutive values ​​in the motion signal. Finally, electronic device 101 uses the absolute value of the motion velocity signal to define a threshold for motion artifact removal. Sudden motion artifacts have higher velocities than the movement of the head and chest caused by breathing. Therefore, artifacts appear as outliers in the distribution of the motion velocity signal. Electronic device 101 can use kurtosis or other techniques to determine whether the motion signal within a thirty-second or other window contains sudden motion artifacts. Kurtosis-based motion artifact removal sets the noise component to zero based on a dynamic threshold. If the kurtosis increases, the probability distribution becomes thinner and more concentrated around the mean. Therefore, when the kurtosis is greater than a selected value (such as three), the motion signal has more outliers.

[0060] After the electronic device 101 identifies the presence of motion artifacts, the electronic device 101 may determine outliers, such as based on a static or dynamic threshold. In some embodiments, a value of 0.35 may be selected as a static threshold based on an observation of the distribution of amplitude signal values. Of course, other values ​​are possible and within the scope of the present disclosure. The first ten percent or other portion of the distribution of the absolute velocity signal on the Y-axis may become the dynamic threshold in each window. In some cases, only the velocity signal on the Y-axis may be used because breathing primarily affects vertical movement of the face or chest. Any movement on the X-axis is more likely to be noise during active movement. Therefore, Y-axis velocity values ​​that exceed the threshold may be considered outliers and may be replaced with zero, which is similar to replacing sudden movements with holding your breath.

[0061] After removing motion artifacts, electronic device 101 uses spectral analysis 235 to determine motion-based respiration signal 240 and estimate instantaneous motion-based RR 245. For example, electronic device 101 may remove the linear trend of the clean velocity signal and smooth the signal using a moving average technique. In some embodiments, a second-order Savitzky-Golay filter with a two-second subset window or other window may be applied to further smooth the signal. Electronic device 101 may use a filter (such as a Butterworth filter using a Hamming window with cutoff frequencies fc1 = 0.05 Hz and fc2 = 0.75 Hz) to extract signals within the spectrum related to respiration. The filtered signal corresponds to motion-based respiration signal 240 within a forty-second window or other window. Motion-based respiration signal 240 may be normalized, such as using the Frobenius norm, and transformed, such as using a discrete Fourier transform (DFT) with zero padding. Electronic device 101 can estimate the respiratory rate (RR) from a frequency domain signal, such as 3 to 45 breaths per minute (BPM), to avoid excessively inaccurate estimates. By observing the DFT signal, the frequency component with the highest peak value can correspond to the instantaneous respiratory rate (RR). The instantaneous respiratory rate (RR) can be measured accordingly for all landmark points. The signal-to-noise ratio (SNR) can determine the signal waveform that is highly correlated with respiration. Therefore, electronic device 101 can select the RR with the highest SNR among the RRs measured from multiple landmark points as the motion-based RR 245.

[0062] In the rPPG-based RR estimation 250, the electronic device 101 performs an rPPG extraction operation 255, in which an rPPG signal is extracted from the ROIs 310 of the video 210. Any suitable technique may be used to extract the rPPG signal. In some embodiments, the electronic device 101 may use a chrominance (CHROM) method to extract the rPPG signal from each ROI 310.

[0063] After extracting the rPPG signal from each ROI 310, the electronic device 101 performs an artifact removal operation 260. In some ROIs 310, camera artifacts (such as those produced by a smartphone camera) may be stronger than the heartbeat of the person being videoed. In other ROIs 310, the camera artifacts are weaker and barely noticeable. To remove camera artifacts, the electronic device 101 may examine the rPPG signals from the ROIs 310 for the presence of strong harmonics. If the power spectral density (PSD) of the second harmonic (such as at 2 Hz) is higher than the dominant PSD (such as at 1 Hz) multiplied by a factor, the rPPG signal may be classified as containing strong camera artifacts and may be discarded. After artifact removal, the rPPG signals from multiple ROIs may be combined into a weighted rPPG signal, such as by using a weighted average based on the SNR.

[0064] The electronic device 101 also performs a signal filtering operation 265. In some embodiments, if the cardiac activity is not pulsating around 1 Hz, the electronic device 101 applies a filter (such as a comb notch filter) to further suppress the weighted rPPG signal having a 1 Hz fundamental frequency. The electronic device 101 can also apply a narrower filter (such as an "HR-RR tuned filter") with a bandwidth that uses the coarse heart rate and respiration rate to the weighted rPPG signal.

[0065] Electronic device 101 uses the weighted rPPG signal to perform inter-beat interval (IBI) extraction 270 to generate an IBI signal. IBI is defined as the distance between consecutive heartbeats in the rPPG, such as in milliseconds. One of the main fluctuations in heart rate is caused by respiratory sinus arrhythmia (RSA). The IBI value decreases with inspiration and increases with expiration. The IBI signal is considered a respiratory signal that can be used to calculate rPPG-based RR 280. In some embodiments, electronic device 101 may use peak detection to generate the IBI signal.

[0066] Because the IBI signal provides a more explicit RSA relationship than the filtered rPPG signal, electronic device 101 selects the interpolated IBI signal as rPPG-based respiration signal 275 and estimates rPPG-based RR 280. Linear trends in the IBI signal may be removed to reduce low-frequency noise. In some embodiments, electronic device 101 may employ linear interpolation so that rPPG-based respiration signal 275 has the same sample size as motion-based respiration signal 240. Electronic device 101 may normalize rPPG-based respiration signal 275, such as by using the Frobenius norm, and transform rPPG-based respiration signal 275, such as by using a DFT with zero padding. Electronic device 101 may estimate rPPG-based RR 280 from a frequency-domain signal, such as from 3 to 45 BPM, to avoid excessively inaccurate estimates.

[0067] The results of motion-based RR estimation 220 and rPPG-based RR estimation 250 include two independent breathing signals (motion-based breathing signal 240 and rPPG-based breathing signal 275) and two RR values ​​(motion-based RR 245 and rPPG-based RR 280). Electronic device 101 may perform a breathing rate selection operation 285 to predict whether motion-based RR 245 or rPPG-based RR 280 is more likely to be accurate, and may select the more accurate frequency, which electronic device 101 may output, display, or otherwise present as RR output 290. In some embodiments, electronic device 101 uses a trained machine learning model (such as a lightweight machine learning classifier) ​​to select between motion-based RR 245 and rPPG-based RR 280. Electronic device 101 may input motion-based breathing signal 240 and rPPG-based breathing signal into the trained machine learning model and may obtain inference results provided by the trained machine learning model. For example, if the absolute difference between the two RR values ​​(motion-based RR 245 and rPPG-based RR 280) is greater than a specified value (such as 2 BPM) and the sample size of the IBI signal is greater than another specified value (such as 19), the electronic device 101 may apply the trained ML model. Otherwise, the rPPG signal quality may be considered insufficient, and the electronic device 101 may select motion-based RR 245 as the default choice. For post-processing of continuous RR estimates, seven-point median smoothing or other smoothing operations may be employed to reduce random noise before finalizing the RR.

[0068] As input to the ML model, the electronic device 101 may extract a plurality of features, such as SNR, number of peaks, and skewness, from each of the windowed respiration signals 240 and 275. These features may represent the signal quality of the respiration signals 240 and 275. For example, Figure 4A and Figure 4B Example graphs 401 and 402 are shown illustrating feature extraction from motion-based and rPPG-based respiratory signals for machine learning-based RR selection, in accordance with the present disclosure. In particular, Figure 4A Graph 401 in FIG. 4 depicts an example motion-based respiration signal 240 over a forty second time window, and Figure 4B Graph 402 in depicts an example rPPG-based respiratory signal 275 over a forty second window.

[0069] like Figure 4A and Figure 4B As shown, the signal-to-noise ratio (SNR), number of peaks, and skewness can be identified from signals 240 and 275. The SNR determines the signal waveform associated with respiration height and can be calculated based on the PSD of each respiration signal 240 and 275. The number of peaks in a periodic respiration signal can be directly correlated with the respiration rate (RR). In some embodiments, electronic device 101 can apply the same peak detection algorithm used for IBI detection. Skewness is a measure of the asymmetry of a probability distribution. Among the eight signal quality indicators (SQIs) for PPG signals, the skewness indicator outperforms the others. The shape of the individual waveforms of respiration signals 240 and 275 differs from that of PPG signals, but the skewness indicator can determine whether distortion exists on the windowed signal. The skewness indicator may increase when the windowed signal has a weak or irregular waveform. The number of peaks and skewness can be calculated in the time domain.

[0070] According to other embodiments, an ML model can be trained to receive at least a portion of a video as input and provide an output indicating whether motion-based RR 245 or rPPG-based RR 280 is more likely to be accurate, or an output indicating whether motion-based RR or rPPG-based RR is more accurate. In this embodiment, the electronic device 101 can obtain two different types of RR from the video, such as motion-based RR and rPPG-based RR, and can input the video into the trained model to select one of the two different types of RR. According to another embodiment, before obtaining two different types of RR (e.g., motion-based RR and rPPG-based RR), the electronic device can select one of the two different types of RR by using an RR signal or by using the video. After selecting one of the two different types of RR, the electronic device 101 can obtain the selected type of RR and not obtain the unselected type of RR.

[0071] As discussed above, in some embodiments, the ML model can be a binary classification model, but there is no limitation on the type of ML model. The classification model can be trained to determine the final output between two calculated RRs. To train the ML model, electronic device 101 (or server 106 or other device) can access a dataset comprising multiple training samples. In some embodiments, each training sample includes a motion-based breathing signal, an rPPG-based breathing signal, and a label indicating whether the motion-based RR or the rPPG-based RR is closer to the ground truth RR for that training sample. Furthermore, in some embodiments, the label for each training sample is the name of the modality with the smaller error in the calculated RR. Furthermore, in some embodiments, electronic device 101, server 106, or other device can partition the dataset into a training set and a test set, such as in a 2:1 ratio. Therefore, only a subset of the entire dataset can be used for training to avoid overfitting.

[0072] For each training sample in the training set, electronic device 101, server 106, or other device performs training. Specifically, electronic device 101, server 106, or other device extracts features from the motion-based and rPPG-based respiration signals and provides the features as input to an ML model, which predicts whether the motion-based or rPPG-based RR is more likely to be closer to the ground truth RR. The ML classifier can be trained using any suitable set of features. In some embodiments, the features may include SNR, number of peaks, and skewness. Based on the comparison of the labels with the predictions, electronic device 101, server 106, or other device updates one or more parameters or weights of the ML model. In some cases, a 9:1 class weighting for rPPG-derived and motion-derived RRs can be applied to the decision tree to address any class imbalance in the feature set. As discussed, training the ML model can be performed by at least one of electronic device 101, server 106, or other device. Furthermore, inference of the ML model can be performed by at least one of electronic device 101, server 106, or other device. The electronic device may request inference of the ML model to the server 106 by sending an input value of the ML model, and may receive a result of the inference from the server 106 .

[0073] although Figures 2 to 4B One example and related details of a process 200 for contactless monitoring of respiration rate using facial video are shown, but may be used for Figures 2 to 4B For example, although process 200 is described as involving a particular sequence of operations, Figure 2 The various operations described may overlap, occur in parallel, occur in a different order, or occur any number of times (including zero). Furthermore, Figure 2 The specific operations shown in are examples only and may be performed using other techniques. Figure 2 Each operation shown in . As a specific example, instead of kurtosis, skewness can be used to determine signal distortion because skewness measures the asymmetry of the distribution values ​​in the windowed signal. Thus, for example, the top ten percent of the distribution of the absolute velocity signal on the Y-axis can be used as a dynamic threshold in each window. Y-axis velocity values ​​that exceed the threshold can be considered outliers and can be replaced with zero, which can be similar to replacing any sudden movement with holding your breath. Furthermore, instead of using a second-order Savitzky-Golay filter, a Butterworth filter can be used to smooth the motion signal.

[0074] It should be noted that the use of spectral analysis to estimate RR (such as Figure 2 The described method may fail to detect breath-hold events because the peak of the power spectrum cannot be zero if there is any motion noise in the signal. Therefore, an ML-based breath-missing detector algorithm can be used to identify apnea events and improve the overall RR estimation accuracy.

[0075] Figure 5 An example process 500 for detecting lack of breathing using facial video according to the present disclosure is shown. For ease of explanation, the process 500 is described as using the above Figure 1 1. However, this is merely an example, and process 500 may be implemented using any other suitable device(s), such as server 106, and in any other suitable system(s).

[0076] like Figure 5 As shown, process 500 includes Figure 2 In some embodiments, process 500 and process 200 may be performed together sequentially or in parallel to provide a more robust respiratory assessment solution. In process 500, electronic device 101 captures video 510 showing a person's face.

[0077] The electronic device 101 performs a face and landmark detection operation 515 on the video 510 to detect a face region of a person, a plurality of ROIs, and a plurality of facial landmarks. Operation 515 may be performed with Figure 2The facial and landmark detection operation 215 is the same or similar to that of FIG. In some embodiments, the electronic device 101 may implement an ML model to detect facial regions and facial landmarks. A facial detection algorithm may be used to analyze each frame of the video 510. When a facial region is detected, the background may be removed to reduce image processing costs and potential incorrect facial detection. The average position of a set of landmarks in the forehead region and a set of landmarks in the nose region may be determined in each frame.

[0078] The electronic device 101 tracks facial landmarks over time to generate a motion tracking signal 520 representing head movement. Robust motion tracking signal 520 is useful for obtaining respiration-related information from the video 510. In some embodiments, the electronic device 101 may determine the positional changes of the landmarks in XY coordinates on a frame-by-frame basis to generate the motion tracking signal 520. If necessary, if the detected face moves out of the frame, the face and landmark detection operation 515 may be performed again.

[0079] The electronic device 101 also performs loss of breathing detection 525 using a sliding window of the motion tracking signal 520. In some embodiments, the electronic device 101 may use a seven-second sliding window method with a one-second interval. However, note that other window sizes (such as six or eight seconds) and other intervals (such as two or three seconds) may be possible. Loss of breathing detection 525 includes a feature extraction operation 530. In the feature extraction operation 530, the electronic device 101 generates multiple signals from the motion tracking signal 520, such as a normalized signal, a filtered signal, and a velocity signal. The raw motion tracking signal 520 for each window may be normalized by removing the linear trend of the signal, thereby generating a normalized signal. The electronic device 101 may create the filtered signal using a filter (such as a second-order Butterworth filter with cutoff frequencies of 0.05 and 0.75). The velocity signal may represent the difference between consecutive values ​​of the smoothed normalized signal using a moving average.

[0080] The electronic device 101 extracts statistical features from the normalized signal, the filtered signal, and the speed signal in the time domain. The statistical features represent characteristics of the signal, such as mean, variance, standard deviation, minimum, maximum, absolute maximum, average quadratic power, range, median, root mean square, crest factor, skewness, kurtosis, or any combination thereof. The electronic device 101 also extends the normalized signal using zero padding and transforms the normalized signal, such as using a fast Fourier transform (FFT), to obtain features in the frequency domain. The electronic device 101 can calculate the same statistical features from the power spectrum, such as the frequency range between 3 BPM and 45 BPM.

[0081] Once electronic device 101 has obtained the various features, it feeds the extracted features into a random forest classifier model 535 trained for apnea detection. In some embodiments, random forest classifier model 535 uses an average of multiple decision tree classifiers trained on various subsamples of the training dataset. In some embodiments, an apnea event is defined as a pause in respiratory activity exceeding a predetermined duration (such as 9 seconds, 10 seconds, 11 seconds, or other duration). Consecutive breath-hold classification results can be aggregated to detect apnea episodes.

[0082] Electronic device 101 also performs respiration signal extraction 540 using a sliding window of motion tracking signal 520. Respiration signal extraction 540 may include motion artifact removal 545 (which may be the same or similar to motion artifact removal 230) and spectral analysis 550 (which may be the same or similar to spectral analysis 235). Motion artifact removal 545 may be used to determine whether motion tracking signal 520 exhibits any active head movement. When the kurtosis of the velocity signal is greater than a specified value (such as three), the windowed signal may be excluded. The results of spectral analysis 550 may be used to calculate the respiration rate (RR). A final RR output 590 may be determined by combining the RR with the results from loss of respiration detection 525.

[0083] The random forest classifier model 535 can be trained using a dataset of training videos. In some embodiments, the dataset can be collected while the subject is video recorded while performing various tasks. These tasks can include breath holding, in which the subject holds his or her breath for a period of time (such as up to one minute) and has natural breathing for another period of time (such as ten seconds). These tasks can also include controlled breathing, in which the subject watches a guided breathing video in order to perform controlled breathing at a target rate (such as 5, 10, 15, 20, and 25 breaths per minute). These tasks can further include spontaneous breathing at lower light levels, so that a video of the face breathing spontaneously is recorded at low lighting levels. The video can be captured using a commercially available RGB camera (such as a smartphone camera) or other imaging device(s). In some embodiments, to avoid any overfitting problems, the dataset can be divided into a training set and a test set, for example with a ratio of 2:1.

[0084] although Figure 5 One example of a process 500 for detecting absence of breathing using facial video is shown, but may be used for Figure 5 For example, although process 500 is described as involving a specific sequence of operations, Figure 5 The various operations described may overlap, occur in parallel, occur in a different order, or occur any number of times (including zero). Furthermore, Figure 5The specific operations shown in are examples only and may be performed using other techniques. Figure 5 Each operation shown in .

[0085] Figure 6 An example method 600 for contactless monitoring of respiration rate and lack of respiration using facial video according to the present disclosure is shown. For ease of explanation, Figure 6 The method 600 shown is described using Figure 1 The electronic device 101 and Figure 2 However, Figure 6 The method 600 shown in FIG. 6 may be used with any other suitable device(s) or system(s), and may be used to perform any other suitable process, such as Figure 5 The process 500 shown in FIG.

[0086] like Figure 6 As shown, at step 601, a video of a person's face is captured using a camera. This may include, for example, the electronic device 101 capturing a video 210 of a person's face, such as Figure 2 At step 603, a motion-based RR and a motion-based respiration signal are determined based on the video of the person's face. This may include, for example, the electronic device 101 performing a motion-based RR estimation 220 to determine a motion-based respiration signal 240 and a motion-based RR 245, such as Figure 2 As shown in .

[0087] At step 605, an rPPG-based RR and an rPPG-based breathing signal are determined based on the video of the person's face. This may include, for example, the electronic device 101 performing an rPPG-based RR estimation 250 to determine an rPPG-based RR 280 and an rPPG-based breathing signal 275, such as Figure 2 At step 607, the trained ML model is used to predict whether the motion-based RR or the rPPG-based RR is more likely to be accurate. The ML model receives the motion-based breathing signal and the rPPG-based breathing signal as input. This may include, for example, the electronic device 101 performing an ML-based breathing rate selection operation 285, such as Figure 2 At step 609, the motion-based RR or rPPG-based RR is presented based on the prediction. This may include, for example, the electronic device 101 displaying, transmitting, or otherwise outputting the RR output 290, such as Figure 2 As shown in .

[0088] although Figure 6 One example of a method 600 for contactless monitoring of respiratory rate and absence of breathing using facial video is shown, but may be used for Figure 6For example, although shown as a series of steps, Figure 6 The various steps in may overlap, occur in parallel, occur in a different order, or occur any number of times (including zero).

[0089] The disclosed embodiments are applicable to a wide variety of use cases. For example, the disclosed embodiments enable any suitable consumer electronic device (such as a person's smartphone, smart TV, tablet computer, etc.) to monitor a person's vital signs in real time. Because the user does not have to wear any sensors, vital signs can be monitored in a contactless manner. Vital signs can be monitored during home workouts, during video calls (such as with a healthcare provider), or while sleeping. As a specific example, an infant's vital signs can be monitored as part of a neonatal or infant monitoring application.

[0090] Notice, Figures 2 to 6 Shown in or about Figures 2 to 6 The described operations and functions may be implemented in any suitable manner in the electronic device 101, 102, 104, server 106, or other device(s). For example, in some embodiments, the described operations and functions may be implemented or supported using one or more software applications or other software instructions executed by the processor 120 of the electronic device 101, 102, 104, server 106, or other device(s). Figures 2 to 6 Shown in or about Figures 2 to 6 In other embodiments, dedicated hardware components may be used to implement or support the operations and functions described. Figures 2 to 6 Shown in or about Figures 2 to 6 At least some of the operations and functions described. Generally, any suitable hardware or any suitable combination of hardware and software / firmware instructions may be used to perform Figures 2 to 6 Shown in or about Figures 2 to 6 Describes the operation and functionality.

[0091] Figure 7 An example method 700 for contactless monitoring of respiration rate and lack of respiration using facial video according to the present disclosure is shown. For ease of explanation, Figure 7 The method 700 shown is described as using Figure 1 However, Figure 7 The method 700 shown in FIG. 7 may be used with any other suitable device(s) or system(s) and may be used to perform any other suitable process.

[0092] like Figure 7As shown, at step 701, a video may be captured using a camera. At least a portion of the video (i.e., at least a portion of an image comprised of the video) may include a person's face. At step 703, the electronic device 101 may capture first data (e.g., a motion-based breathing signal) from at least a portion of the video and determine a first type of respiration rate (RR) by applying a first scheme (e.g., a motion-based RR based on the first data). At step 705, the electronic device 101 may capture second data (e.g., an rPPG-based breathing signal) from at least a portion of the video and determine a second type of respiration rate (RR) by applying a second scheme (e.g., an rPPG-based RR based on the second data). At step 707, the electronic device 101 may select one of the first type of RR and the second type of RR by inputting the first and second data as inputs into a trained machine learning model. As discussed, the trained machine learning model may receive the first data (e.g., the motion-based breathing signal) and the second data (e.g., the rPPG-based breathing signal) and provide an inference result indicating one of the first type of RR and the second type of RR. At step 709, the electronic device 101 may present the one of the first type of RR and the second type of RR. In other embodiments, the electronic device 101 may select one of the first type of RR and the second type of RR by inputting at least a portion of the video into a trained machine learning model instead of the first data and the second data as input.

[0093] Figure 8 An example method 700 for contactless monitoring of respiration rate and lack of respiration using facial video according to the present disclosure is shown. For ease of explanation, Figure 8 The method 800 shown in FIG. 8 is described as using Figure 1 However, Figure 8 The method 800 shown in FIG. 8 may be used with any other suitable device(s) or system(s) and may be used to perform any other suitable process.

[0094] like Figure 8As shown, at step 801, a video may be captured using a camera. At least a portion of the video (i.e., at least a portion of an image formed by the video) may include a person's face. At step 803, the electronic device 101 may capture first data, such as a motion-based respiration signal, from at least a portion of the video. At step 805, the electronic device 101 may capture second data, such as an rPPG-based respiration signal, from at least a portion of the video. At step 807, before capturing at least one of a first type of respiration rate (RR) (e.g., a motion-based RR) or a second type of respiration rate (RR) (e.g., an rPPG-based RR), the electronic device 101 may select one of the first type of RR and the second type of RR, such as the first type of RR, by inputting the first and second data as inputs into a trained machine learning model. As discussed above, the trained machine learning model may receive the first data (e.g., the motion-based respiration signal) and the second data (e.g., the rPPG-based respiration signal), and provide an inference result indicating one of the first type of RR and the second type of RR before at least one of the first type of RR or the second type of RR is captured. At step 809, the electronic device 101 may obtain the first type of RR based on the first date, while avoiding obtaining the second type of RR. In other embodiments, the electronic device 101 may select one of the first type of RR and the second type of RR by inputting at least a portion of the video instead of the first data and the second data into a trained machine learning model.

[0095] Although the present disclosure has been described with reference to various exemplary embodiments, various changes and modifications may be suggested to one skilled in the art. The present disclosure is intended to encompass such changes and modifications as fall within the scope of the appended claims.

Claims

1. A method comprising: Use the camera to capture video; determining a motion-based respiration rate (RR) and a motion-based respiration signal based on a face of a person identified based on the video; determining a remote photoplethysmography (rPPG)-based RR and an rPPG-based respiration signal based on a face of the person, the face of the person being identified based on the video; selecting one of the motion-based RR or the rPPG-based RR by inputting the motion-based respiration signal and the rPPG-based respiration signal as input into a trained machine learning model; as well as A selected one of the motion-based RR or the rPPG-based RR is presented.

2. The method according to claim 1, wherein Determining the motion-based RR and the motion-based respiration signal based on the person's face includes: identifying facial landmarks on the face of the person based on the video, wherein the facial landmarks are on the forehead of the person and the nose of the person; generating a motion signal based on vertical position changes of the facial landmarks in the video; and The motion-based respiration signal is extracted based on the motion signal using spectral analysis.

3. The method according to claim 1 , wherein: Determining the motion-based RR and the motion-based respiration signal based on the person's face further comprises: removing artifacts from the motion signal using a kurtosis-based motion artifact detection technique; and A filter is used to smooth the motion signal.

4. The method according to claim 1 , wherein: Determining the rPPG-based RR and the rPPG-based respiration signal based on the face of the person includes: identifying a region of interest on the person's face; extracting rPPG signals from each region of interest based on the video; extracting an inter-beat interval (IBI) signal based on a weighted combination of the rPPG signals corresponding to the region of interest; as well as The rPPG-based respiration signal is extracted based on the IBI signal.

5. The method according to claim 1 , wherein: The machine learning model is a binary classifier model trained by the following operations: accessing a training dataset comprising a plurality of training samples, each training sample comprising a motion-based respiration signal, an rPPG-based respiration signal, and a label indicating whether the motion-based RR or the rPPG-based RR is closer to a ground truth RR of the training sample; as well as For each training example: extracting features of the motion-based respiratory signal and the rPPG-based respiratory signal; providing the features as input to the machine learning model, the machine learning model predicting whether the motion-based RR or the rPPG-based RR is more likely to be closer to the baseline true RR; as well as Parameters of the machine learning model are updated based on the comparison of the labels to the predictions.

6. The method according to claim 1 , wherein: The characteristics include one or more of: signal-to-noise ratio, number of peaks, and skewness.

7. The method according to claim 1 , wherein: The camera is coupled to a mobile device, a computer, or a television.

8. An electronic device comprising: Camera; and at least one processing device; a memory storing instructions that, when executed by at least a portion of the at least one processing device, cause the electronic device to: acquiring a video using said camera; determining a motion-based respiration rate (RR) and a motion-based respiration signal based on a face of a person identified based on the video of the face of the person; determining a remote photoplethysmography (rPPG)-based RR and an rPPG-based respiration signal based on a face of the person identified based on the video of the face of the person; selecting one of the motion-based RR or the rPPG-based RR by inputting the motion-based respiration signal and the rPPG-based respiration signal as input into a trained machine learning model; as well as A selected one of the motion-based RR or the rPPG-based RR is presented.

9. The electronic device according to claim 8, wherein: To determine the motion-based RR and the motion-based respiration signal based on the person's face, the instructions, when executed by at least a portion of the at least one processing device, cause the electronic device to: identifying facial landmarks on the face of the person based on the video, wherein the facial landmarks are on the forehead of the person and the nose of the person; generating a motion signal based on vertical position changes of the facial landmarks in the video; and The motion-based respiration signal is extracted using spectral analysis based on the motion signal.

10. The electronic device according to claim 8, wherein: To determine the motion-based RR and the motion-based respiration signal based on the person's face, the instructions, when executed by at least a portion of the at least one processing device, cause the electronic device to: removing artifacts from the motion signal using a kurtosis-based motion artifact detection technique; as well as A filter is used to smooth the motion signal.

11. The electronic device according to claim 8, wherein: To determine the rPPG-based RR and the rPPG-based respiration signal based on the person's face, the instructions, when executed by at least a portion of the at least one processing device, cause the electronic device to: identifying a region of interest on the person's face; extracting rPPG signals from each region of interest based on the video; extracting an inter-beat interval (IBI) signal based on a weighted combination of the rPPG signals corresponding to the region of interest; as well as The rPPG-based respiration signal is extracted based on the IBI signal.

12. The electronic device according to claim 8 , wherein: The machine learning model is a binary classifier model; and To train the machine learning model, the instructions, when executed by at least a portion of the at least one processing device, cause the electronic device to: accessing a training dataset comprising a plurality of training samples, each training sample comprising a motion-based respiration signal, an rPPG-based respiration signal, and a label indicating whether the motion-based RR or the rPPG-based RR is closer to a ground truth RR of the training sample; as well as For each training example: extracting features of the motion-based respiratory signal and the rPPG-based respiratory signal; providing the features as input to the machine learning model, the machine learning model predicting whether the motion-based RR or the rPPG-based RR is more likely to be closer to the baseline true RR; as well as Parameters of the machine learning model are updated based on the comparison of the labels to the predictions.

13. The electronic device according to claim 8, wherein: The characteristics include one or more of: signal-to-noise ratio, number of peaks, and skewness.

14. The electronic device according to claim 8, wherein: The electronic device includes a mobile device, a computer or a television.

15. A non-transitory machine-readable medium containing instructions that, when executed, cause an electronic device to: Use the camera to capture video; determining a motion-based respiration rate (RR) and a motion-based respiration signal based on a face of a person identified based on the video; determining a remote photoplethysmography (rPPG)-based RR and an rPPG-based respiration signal based on a face of the person, the face of the person being identified based on the video; selecting one of the motion-based RR or the rPPG-based RR by inputting the motion-based respiration signal and the rPPG-based respiration signal as input into a trained machine learning model; as well as The one of the motion-based RR or the rPPG-based RR is presented.