Enhanced vision-based vital sign monitoring using multi-modal soft tags

By using multimodal soft tag technology and enhancing deep learning models with physiological parameter values ​​from contact sensors, the problem of accurate measurement of rPPG in uncontrolled environments is solved, achieving high-precision vital sign monitoring, adapting to changes in users and the environment, and expanding the measurement range of physiological parameters.

CN121816149APending Publication Date: 2026-04-07SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing vision-based vital sign monitoring methods, especially remote photoplethysmography (rPPG), struggle to achieve accurate measurements in uncontrolled environments. They are limited by single video modalities and noise factors, and training robust deep learning models requires synchronous and diverse datasets, leading to performance degradation when users and environments change.

Method used

By employing multimodal soft-label technology, physiological parameter values ​​from contact sensors are used as soft labels to enhance the training process of deep learning models. Through video enhancement and cost functions, the generalization ability of the models is improved, adapting to changes in users and the environment, and improving the accuracy of physiological parameter measurements.

Benefits of technology

It enables high-precision measurement of vital signs in uncontrolled environments, expands the measurement range of physiological parameters, improves the robustness and adaptability of the model, and enables continuous monitoring of vital signs under non-laboratory conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121816149A_ABST
    Figure CN121816149A_ABST
Patent Text Reader

Abstract

A method performed by at least one processor includes obtaining an image of an object; preprocessing the image of the object; inputting the pre-processed image into a machine learning model trained according to a first frequency distribution, the first frequency distribution corresponding to first truth values obtained from one or more sensors performing vital sign measurements on one or more test subjects; and obtaining, from the machine learning model, an estimate of the signal corresponding to the vital sign measurement of the subject.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to enhanced vision-based vital sign monitoring using soft tags. Background Technology

[0002] Heart rate (HR) and heart rate variability (HRV) are important vital signs, biomarkers, and physiological parameters used to assess a person's cardiac function. Most devices that measure the heart's pulse (such as fingertip pulse oximeters (PPG) or electrocardiogram (ECG) patches) require physical contact with the subject. Additionally, these devices can be prohibitively expensive, limiting measurements to visits to medical facilities.

[0003] This paper introduces a vision-based method for non-contact measurement of blood volume pulse using a camera. This vision-based method is called remote photoplethysmography (rPPG). rPPG enables low-cost and widespread health monitoring using low-cost cameras readily available in mobile phones, computers, tablets, etc. rPPG signals can be analyzed to extract multiple physiological parameters, including but not limited to HR, HRV, RR (respiratory rate), SpO2 (oxygen saturation), or BP (blood pressure). Despite the obvious benefits, implementing an accurate rPPG system remains challenging in practice.

[0004] rPPG allows for non-contact measurement of blood volume pulse using a conventional camera. Most studies have evaluated the robustness of rPPG systems by frequency over short time windows (e.g., pulse rate in bpm). As systems improve, it will be beneficial to support more challenging measurement configurations.

[0005] While camera-based vital sign measurement has improved in recent years, traditional rPPG methods follow a stepwise transformation from a single input video to a temporal signal (rPPG) representing the pulse. Commonly used methods include color transformation, blind source separation, and signal processing. These methods do not always handle environmental noise factors (e.g., motion) well. To create the most robust rPPG algorithms, researchers have begun exploring data-driven approaches (such as deep learning using convolutional neural networks (CNNs) or transformers) to predict the rPPG temporal signal solely from the video. Supervised learning frameworks are used to train the neural networks, where ground truth PPG or ECG signals are used as target labels during backpropagation.

[0006] However, for deep learning systems to be believable and generalizable, current solutions require large training datasets with diverse sets of data covering skin color, lighting, camera sensors, motion, and physiological range. Collecting such diverse data is challenging because it requires capturing physiological ground truth simultaneously. Many modern deep learning frameworks for rPPG even require time-synchronized PPG waveforms. Summary of the Invention

[0007] According to one aspect of this disclosure, a method executed by at least one processor includes: obtaining an image of an object; preprocessing the image of the object; inputting the preprocessed image into a machine learning model trained according to a first frequency distribution, wherein the first frequency distribution corresponds to a first ground truth obtained from one or more sensors performing vital sign measurements on one or more test objects; and obtaining an estimate of a signal corresponding to the vital sign measurements of the object from the machine learning model.

[0008] According to one aspect of this disclosure, an apparatus includes: a memory; and processing circuitry incorporated into the memory, the processing circuitry being configured to: acquire an image of an object; preprocess the image of the object; input the preprocessed image into a machine learning model trained according to a first frequency distribution, wherein the first frequency distribution corresponds to a first ground truth obtained from one or more sensors performing vital sign measurements on one or more test objects; and obtain an estimate of a signal corresponding to the vital sign measurements of the object from the machine learning model.

[0009] According to one aspect of this disclosure, a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method comprising: obtaining an image of an object; preprocessing the image of the object; inputting the preprocessed image into a machine learning model trained according to a first frequency distribution, wherein the first frequency distribution corresponds to a first ground truth obtained from one or more sensors performing vital sign measurements on one or more test objects; and obtaining an estimate of a signal corresponding to the vital sign measurements of the object from the machine learning model. Attached Figure Description

[0010] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which: Figure 1 This is a diagram illustrating an environment in which the methods, apparatus, and systems described herein can be implemented according to embodiments of this disclosure.

[0011] Figure 2 According to embodiments of this disclosure Figure 1 A block diagram of example components of one or more devices.

[0012] Figure 3An example system for vision-based vital sign monitoring using multimodal soft tags according to embodiments of the present disclosure is shown.

[0013] Figure 4 An example block diagram for determining a loss function according to an embodiment of the present disclosure is shown.

[0014] Figure 5 An example block diagram of a system for training using multimodal soft labels according to an embodiment of the present disclosure is shown.

[0015] Figure 6 An example system for vision-based vital sign monitoring using multimodal soft tags according to embodiments of the present disclosure is shown.

[0016] Figure 7 An example system for vision-based vital sign monitoring using multimodal soft tags according to embodiments of the present disclosure is shown.

[0017] Figure 8 An example system for vision-based vital sign monitoring using multimodal soft tags according to embodiments of the present disclosure is shown. Detailed Implementation

[0018] The following detailed description of the exemplary embodiments is with reference to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0019] The foregoing disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. In view of the above disclosure, modifications and variations are possible, or modifications and variations may be obtained from practice of the embodiments. Furthermore, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, in the flowcharts and descriptions of operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be interchanged.

[0020] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware or firmware. The actual dedicated control hardware used to implement these systems and / or methods does not limit the implementation.

[0021] Even if specific combinations of features are listed in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features can be combined in ways not specifically listed in the claims and / or not disclosed in the specification. Although each dependent claim listed below may be directly subordinated to only one claim, the disclosure of possible embodiments includes a combination of each dependent claim in the claim set with each of the other claims.

[0022] Unless explicitly stated otherwise, elements, actions, or instructions used herein should not be construed as critical or essential. Furthermore, as used herein, articles are intended to include one or more items and are used interchangeably with “one or more.” Where intended to include only one item, the term “a” or similar language is used. Furthermore, as used herein, terms such as “having,” “including,” etc., are intended to be open-ended terms. Furthermore, unless explicitly stated otherwise, the phrase “based on” is intended to mean “at least partially based on.” Additionally, expressions such as “at least one of [A] and [B]” or “at least one of [A] or [B]” should be understood to include only A, only B, or both A and B.

[0023] Throughout this specification, references to "an embodiment," "embodiment," or similar language mean that a particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of this solution. Therefore, the phrases "in one embodiment," "in an embodiment," and similar language throughout this specification may, but not necessarily all, refer to the same embodiment.

[0024] Furthermore, the features, advantages, and characteristics described herein may be combined in any suitable manner in one or more embodiments. In view of the description herein, those skilled in the art will recognize that this disclosure may be practiced without one or more specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not be present in all embodiments of this disclosure may be recognized in particular embodiments.

[0025] Accurately estimating blood volume pulse from video is challenging. The main difficulty lies in the fact that the pulse signal is extremely small compared to other dynamic data in the captured video. If the object moves, the rPPG signal can be very difficult to determine. Furthermore, noisy camera sensors can cause pixel variations over time, which do not represent changes in the observed environment. Another challenge is the variability between objects, where different objects may have higher or lower peripheral rPPG signal amplitudes due to physical differences in melanin concentration, microvascular system, or hair covering the skin. Creating algorithms that can reliably determine the pulse in the presence of all these factors is an ongoing research problem.

[0026] Systems relying solely on video input may have limitations in accurately measuring a wide range of vital signs in uncontrolled environments. More detailed context of the input signal (e.g., in the time or frequency domain) is required to distinguish subtle physiological changes from external noise without sacrificing the operational range of physiological parameters. For example, these limitations include measuring thermal rate (HR) or respiratory rate (RR) only within the nominal / resting range of the subject at rest, or insufficient signal quality for measuring more complex vital signs such as oxygen saturation (SpO2) or blood pressure (BP), where measurements depend on more subtle changes in signal amplitude.

[0027] Current systems utilize data-driven methods, such as supervised deep learning, which requires large amounts of video data with target labels. For rPPG, the target is the temporal signal of blood volume pulse. In the past, target pulse or ECG signals have been collected by contact-based PPG sensors or ECG patches synchronized with the camera recording the video. However, setting up a single device for video and PPG collection is prone to failure, requires expertise, and is limited to laboratory settings with computers. This setup severely limits the environmental diversity in the training data necessary to train robust deep learning models. Furthermore, relying on a single modality from the video limits the context provided for estimating physiological parameters. For example, relying on a single modality from the video to infer rPPG signals and parameters may limit the performance of vital sign monitoring. Moreover, obtaining synchronized PPG / ECG signals for training robust rPPG models is highly improbable. However, “soft labels” inferred from PPG / ECG signals that are not necessarily synchronized with the video are particularly common in smart wearable devices.

[0028] Embodiments of this disclosure pertain to systems and methods for vision-based, contactless monitoring of physiological parameters. The embodiments propose a novel framework that enables rPPG models to utilize soft vital sign labels from modalities other than cameras to achieve: (i) higher measurement accuracy over a wider range of physiological values; (ii) an expanded list of vital signs (such as blood pressure); and (iii) an adaptive model for the user or environment.

[0029] According to one or more embodiments, soft tags of physiological parameter values ​​(HR, RR, SpO2, ...) can be utilized during the generation of the rPPG time signal. Soft tags from other modalities (such as contact sensors (PPG / ECG)) can be used to improve the system in a variety of ways, including but not limited to: (i) rPPG prediction model enhancement using soft physiological tags (e.g., soft tag-based rPPG models); (ii) a physiologically sensed cost function (rPPG cost function) for rPPG model training; (iii) physiologically sensed video enhancement for robust rPPG prediction models (e.g., rPPG video enhancement); (iv) rPPG prediction models adapted to the user and environment (e.g., adaptive rPPG models); and (v) multimodal rPPG prediction models utilizing soft physiological tags (e.g., multimodal rPPG models). One or a combination of these novel components can be used to develop contactless vital sign monitoring based on multimodal soft tags.

[0030] Embodiments of this disclosure relate to a method for training an enhanced deep learning (DL) model architecture for predicting time-series rPPG signals from video. Instead of actual synchronized time-series signal labels (PPG / ECG), embodiments of this disclosure may employ soft labels of physiological parameter values ​​provided by other sources or modalities for training. According to one or more embodiments, the soft labels themselves may guide the training of the model to capture time-series signals (e.g., rPPG) corresponding to target physiological parameters (e.g., HR).

[0031] Embodiments of this disclosure implement a novel set of target cost functions that take into account the inherent properties of physiological signals and extracted parameters to facilitate deep learning training, thereby achieving higher accuracy. The cost functions can formulate error functions that process generated rPPG signals within the target's physiological parameter range for comparison with soft labels of the physiological parameters. These cost functions enable training of predictive models based on soft labels rather than ground-value time-series signals.

[0032] Embodiments of this disclosure pertain to a set of video enhancement methods with soft labels corresponding to physiologically rich videos for training more robust deep learning models. Automatic video enhancement, which considers the inherent correlation between physiological parameters and the characteristics of the input video, advantageously prevents overfitting and improves the generalization of the DL model under varying conditions, such as lighting or loss of video information.

[0033] Embodiments of this disclosure relate to a method in which an rPPG prediction model is tuned and fine-tuned during the inference phase to capture novel dynamics of the target user and environment. The method utilizes soft physiological tags, for example, from other sources and modalities, to tune and correct the underlying rPPG model to adapt to novel user and environmental dynamics: skin color, lighting conditions, distance, or camera settings.

[0034] Embodiments of this disclosure relate to a method and model architecture that utilizes physiological tags from other modalities as well as video input to generate enhanced rPPG signals. The additional context from the physiological tags of another modality helps the model better distinguish physiological signals in the video from external noise, rather than relying solely on the video. The enhanced multimodal rPPG model and enhanced rPPG signals can contribute to higher accuracy in vital sign monitoring and enable the capture of an expanded list of biometrics (such as continuous BP or SpO2) in uncontrolled environments.

[0035] Figure 1 This is a diagram illustrating an environment 100 in which the methods, apparatus, and systems described herein can be implemented according to embodiments. Figure 1 As shown, environment 100 may include user device 110, platform 120, and network 130. Devices in environment 100 may be interconnected via wired connection, wireless connection, or a combination of wired and wireless connection.

[0036] User device 110 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with platform 120. For example, user device 110 may include computing devices (e.g., desktop computers, laptop computers, tablet computers, handheld computers, smart speakers, servers, etc.), mobile phones (e.g., smartphones, cordless phones, etc.), wearable devices (e.g., smart glasses or smartwatches), or similar devices. In some embodiments, user device 110 may receive information from platform 120 and / or send information to platform 120.

[0037] Platform 120 includes one or more devices as described elsewhere herein. In some embodiments, platform 120 may include a cloud server or a group of cloud servers. In some embodiments, platform 120 may be designed to be modular, allowing software components to be swapped in or out as needed. This allows platform 120 to be easily and / or quickly reconfigured for different purposes.

[0038] In some implementations, as shown in the figures, platform 120 may be hosted in a cloud computing environment 122. It is worth noting that while the implementations described herein depict platform 120 as being hosted in a cloud computing environment 122, in some implementations, platform 120 may not be cloud-based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud-based.

[0039] The cloud computing environment 122 includes the environment of the hosting platform 120. The cloud computing environment 122 provides services such as computing, software, data access, and storage, which do not require end users (e.g., user device 110) to know the physical location and configuration of the system and / or device of the hosting platform 120. As shown in the figure, the cloud computing environment 122 may include a set of computing resources 124 (collectively referred to as "multiple computing resources 124" and individually as "computing resources 124").

[0040] Computing resource 124 includes one or more personal computers, workstations, server devices, or other types of computing and / or communication devices. In some embodiments, computing resource 124 may host platform 120. Cloud resources may include computing instances executing in computing resource 124, storage devices provided in computing resource 124, data transmission devices provided by computing resource 124, etc. In some embodiments, computing resource 124 may communicate with other computing resources 124 via wired connections, wireless connections, or a combination of wired and wireless connections.

[0041] For example Figure 1 As shown, computing resources 124 include a set of cloud resources, such as one or more applications (APP) 124-1, one or more virtual machines (VM) 124-2, virtualized storage (VS) 124-3, one or more hypervisors (HYP) 124-4, etc.

[0042] Application 124-1 includes one or more software applications that can be provided to or accessed by user device 110 and / or platform 120. Application 124-1 eliminates the need to install and run software applications on user device 110. For example, application 124-1 may include software associated with platform 120 and / or any other software that can be provided via cloud computing environment 122. In some implementations, an application 124-1 may send information to / receive information from one or more other applications 124-1 via virtual machine 124-2.

[0043] Virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that runs programs like a physical machine. Depending on its purpose and degree of correspondence with any real machine, virtual machine 124-2 can be a system virtual machine or a process virtual machine. A system virtual machine can provide a complete system platform supporting the operation of a full operating system (OS). A process virtual machine can run a single program and can support a single process. In some implementations, virtual machine 124-2 can execute on behalf of a user (e.g., user device 110) and manage the infrastructure of cloud computing environment 122 (such as data management, synchronization, or long-duration data transfer).

[0044] Virtualized storage 124-3 includes one or more storage systems and / or devices that utilize virtualization technology within the storage system or device of computing resource 124. In some embodiments, the type of virtualization, within the context of the storage system, may include block virtualization and file virtualization. Block virtualization may refer to the abstraction (or separation) of logical storage from physical storage, enabling access to the storage system regardless of physical storage or heterogeneous architecture. This separation allows storage system administrators more flexible management of storage for end users. File virtualization eliminates the dependency between data accessed at the file level and the location where the file is physically stored. This enables optimization of storage utilization, server consolidation, and / or performance for non-disruptive file migration.

[0045] Hypervisor 124-4 provides hardware virtualization technology that allows multiple operating systems (such as "guest operating systems") to run concurrently on a host (such as computing resource 124). Hypervisor 124-4 can present a virtual operating platform to guest operating systems and manage the operation of guest operating systems. Multiple instances of various operating systems can share virtualized hardware resources.

[0046] Network 130 includes one or more wired and / or wireless networks. For example, network 130 may include cellular networks (e.g., fifth-generation (5G) networks, long-term evolution (LTE) networks, third-generation (3G) networks, code division multiple access (CDMA) networks, etc.), public land mobile networks (PLMN), local area networks (LAN), wide area networks (WAN), metropolitan area networks (MAN), telephone networks (e.g., public switched telephone network (PSTN)), private networks, self-organizing networks, intranets, the Internet, fiber-optic networks, etc., and / or combinations of these or other types of networks.

[0047] Figure 1 The number and arrangement of devices and networks shown are provided as examples. In practice, with Figure 1 Compared to those shown, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks arranged differently. Furthermore, Figure 1 The two or more devices shown can be implemented within a single device, or Figure 1 The single device shown can be implemented as multiple distributed devices. Additionally or alternatively, the set of devices in environment 100 (e.g., one or more devices) can perform one or more functions described as being performed by another set of devices in environment 100.

[0048] Figure 2 yes Figure 1A block diagram of example components of one or more devices. Device 200 may correspond to user device 110 and / or platform 120. Device 200 may be any other suitable device (such as a TV, wall panel, etc.). Figure 2 As shown, the device 200 may include a bus 210, a processor 220, a memory 230, a storage component 240, an input component 250, an output component 260, and a communication interface 270.

[0049] Bus 210 includes components that allow communication between components of device 200. Processor 220 is implemented in hardware, firmware, or a combination of hardware and software. Processor 220 is a central processing unit (CPU), graphics processing unit (GPU), accelerometer processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or another type of processing component. In some embodiments, processor 220 includes one or more processors capable of being programmed to perform functions. Memory 230 includes random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic storage, and / or optical storage) that stores information and / or instructions used by processor 220.

[0050] Storage component 240 stores information and / or software related to the operation and use of device 200. For example, storage component 240 may include hard disks (e.g., magnetic disks, optical disks, magneto-optical disks, and / or solid-state disks), compact optical disks (CDs), digital versatile optical disks (DVDs), floppy disks, cassette tapes, magnetic tapes, and / or other types of non-transitory computer-readable media and corresponding drives.

[0051] Input component 250 includes components that allow device 200 to receive information, such as via user input (e.g., a touchscreen display, keyboard, keypad, mouse, buttons, switches, and / or microphone). Additionally or optionally, input component 250 may include sensors for sensing information (e.g., a Global Positioning System (GPS) component, accelerometer, gyroscope, and / or actuator). Output component 260 includes components that provide output information from device 200 (e.g., a display, speaker, and / or one or more light-emitting diodes (LEDs)).

[0052] Communication interface 270 includes transceiver components (e.g., transceivers and / or separate receivers and transmitters) that enable device 200 to communicate with other devices, such as via wired connections, wireless connections, or a combination of wired and wireless connections. Communication interface 270 allows device 200 to receive information from and / or provide information to another device. For example, communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0053] Apparatus 200 may perform one or more processes described herein. Apparatus 200 may perform these processes in response to processor 220 executing software instructions stored in non-transitory computer-readable media, such as memory 230 and / or storage components 240. Computer-readable media are defined herein as non-transitory memory devices. Memory devices include memory space within a single physical storage device or memory space distributed across multiple physical storage devices.

[0054] Software instructions can be read from another computer-readable medium or from another device into memory 230 and / or storage component 240 via communication interface 270. When executed, the software instructions stored in memory 230 and / or storage component 240 cause processor 220 to perform one or more processes described herein. Additionally or alternatively, hard-wired circuitry may be used in place of or in combination with software instructions to perform one or more processes described herein. Therefore, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.

[0055] Figure 2 The number and arrangement of components shown are provided as an example. In practice, with... Figure 2 Compared to those shown, device 200 may include additional components, fewer components, different components, or components arranged differently. Additionally or alternatively, the set of components of device 200 (e.g., one or more components) may perform one or more functions described as being performed by another set of components of device 200.

[0056] In one or more examples, device 200 may be a controller for a smart home system that communicates with one or more sensors, cameras, smart home appliances, and / or autonomous robots. Device 200 may communicate with a cloud computing environment 122 to offload one or more tasks.

[0057] Figure 3 An example system for vision-based vital sign monitoring using multimodal soft tags, according to embodiments of the present disclosure, is shown. The system may include a training phase 300 and an inference phase 330.

[0058] During the training phase, video capture 302 can be performed on one or more test subjects. Video capture can be performed by having a user sit in front of an electronic device containing a camera (such as a smartphone, laptop, TV, etc.). Video capture 302 can be a full-body image or a portion of the test subject (e.g., from the waist or shoulders upwards). After video capture 302, a detection process can be performed to identify regions of interest. For example, a face detection process 304 can be performed on the captured image.

[0059] After performing face detection 304, rPPG video augmentation 306 can be performed. According to one or more embodiments, video augmentation is used during model training to prevent overfitting and improve the generalization of the rPPG model (e.g., in a 3D CNN). For example, video augmentation can modify the video in one or more ways to provide the model with additional data for learning.

[0060] Enhanced video can be provided in the form of one or more RGB waveforms, which are fed as input to a soft-label-based rPPG model 308. The soft-label-based rPPG model 308 estimates the rPPG time-series signal corresponding to the estimated vital signs. The output of model 308 can be provided to an rPPG cost function 310. Additionally, data from one or more sensors 312 and physiological labels 314 can be obtained, which are fed as input to the rPPG cost function. In one or more examples, the data from the one or more sensors 312 can be data from PPG or ECG sensors monitoring one or more vital signs of the test subject. The rPPG cost function can calculate a loss function that indicates the degree of correspondence between the estimated time-series signal (e.g., the output from model 308) and the data from the one or more sensors 312.

[0061] The inference phase 330 can be used to estimate the vital signs of a target user. In one or more examples, the target user may be a patient undergoing a telemedicine consultation. During the inference phase, video capture 332 and face detection 334 can be performed in a manner similar to video capture 302 and face detection 304, respectively. One or more RGB signals can be input into a soft-label-based rPPG model 338, which may correspond to a trained version of model 308. In one or more examples, after model 308 is trained, the model can be downloaded to the target user's device, or the user can download an application that provides access to the trained model 308 stored on a server. A cost function 340 can be performed based on the output of model 338 and soft labels 344 obtained from data from one or more sensors 342 monitoring the target user. Model 338 can be further fine-tuned based on the cost function 340. Estimates of the vital signs of the target user 350 can be obtained based on the output of model 338.

[0062] In one or more examples, video augmentation involves adjusting the brightness of the video to account for lighting noise. By brightening or darkening the training video, the model, utilizing soft labels, learns to handle different lighting conditions in the physical environment. Since some environments may lack sufficient lighting, this video augmentation enables the model to learn to find rPPG signals with lower dynamic range. In one or more examples, given an input video X, the augmented video X' can be determined as follows: Equation (1):

[0063] In one or more examples, video enhancement includes Gaussian pixel-level noise. For example, a camera sensor can naturally add noise to the collected video. This noise can be independently modeled as a Gaussian distribution for each pixel location and channel. Models that can still recover the rPPG signal in the presence of noise are more robust. In one or more examples, given an input video X, the enhanced video X' can be determined as follows: Equation (2):

[0064] In one or more examples, video enhancement includes horizontal flipping. For example, each frame in the video can be flipped around the y-axis. During training, this enhancement increases the diversity of the spatial arrangement of facial pixels, so the model will adapt to different facial poses.

[0065] In one or more examples, video enhancement includes random cropping. For example, the video can be randomly cropped along both the width and height, starting at, say, U(0.75,1), and then linearly interpolated to the original video width and height. This enhancement simulates potential failures in face detectors and also allows the model to isolate facial regions with strong rPPG signals. In one or more examples, if only the center-cropped video is provided to the model, the model can remember the user's location without attributing it to the actual infusion of the video.

[0066] In one or more examples, video enhancement may include speed alteration, where the speed of the video is modified (e.g., slowed down or increased). This enhancement may resemble a random cropping operation but is performed along the time dimension of the video. In one or more examples, the video may be randomly slowed down or sped up by a factor c, where c is sampled from U(0.75, 1.75). This enhancement may affect the object's baseline pulse rate. If an object in the original video has a pulse rate of 60 bpm and the video is sped up by 1.5 times, the new pulse rate will be 90 bpm. When the video speed is changed by a factor c, the baseline HR value can be scaled by c. This enhancement allows the model to learn HR values ​​beyond the range of HR values ​​present in the training data, thus enabling the model to extrapolate to different physiological characteristics.

[0067] In one or more examples, video enhancement may include video encoding. For example, the rPPG signal derived from a video may be affected by information loss due to video encoding or compression. Information loss is a common challenge in video-based vital sign monitoring. In one or more examples, various compression ratios and encoding parameters (e.g., bit rate, configuration parameters) can be used to enhance the input video. This enhancement enables the trained model to reconstruct the rPPG signal more robustly, even in the event of information loss.

[0068] According to one or more embodiments, model 308 may be a predictive rPPG model that receives RGB color signals as input and predicts time-series rPPG signals as output. The rPPG model may be trained based on a comprehensive video dataset of the test subject. In one or more examples, any suitable time-series model known to those skilled in the art may be used as the architecture for this model. For example, a series of 3D CNN models may be used to capture the temporal and spatial correlations of RGB signals at different locations within a region of interest (e.g., a user's face). In one or more examples, the 3D CNN model effectively captures these correlations to reconstruct the rPPG signal. In one or more examples, any wearable sensor (PPG or ECG patch) may be used to provide ground truth physiological parameters. The ground truth in this component may be a "soft tag" generated by a ground truth device, but does not necessarily need to be precisely synchronized with the video signal. For example, heart rate from a PPG device may be used as a "soft tag" ground truth.

[0069] According to one or more embodiments, by defining a Gaussian distribution over a frequency band, the true HR labels collected from the truth apparatus can be transformed to the same domain as the power spectral density of the rPPG model. Centered on the main lobe, where the standard deviation is a function of the main lobe. For example, given N frames of video sampled at f frames per second, the standard deviation is defined as a function of the main lobe (f / N). The following equation defines... Figure 4 The target Gaussian distribution shown in the figure: Equation (3):

[0070] like Figure 4 As shown, one of the video enhancement techniques discussed above is used to enhance a 400-bit captured video. The enhanced video is fed to a 3D CNN model to predict the rPPG signal. The power spectrum of the rPPG signal can be obtained. For example, the time signal can be converted to the frequency domain using a Fast Fourier Transform (FFT). The power spectrum can be calculated by taking the square of the absolute value of the FFT output. The following equation describes this process: Equation (4):

[0071] In one or more examples, a soft tag definition 402 may be obtained from one or more devices (e.g., PPG, ECG) that measure vital signs (such as heart rate) of a person or test subject. The power spectrum of the rPPG signal may be compared with the soft tag definition, wherein one or more objective functions 404 may be used as cost functions.

[0072] In one or more examples, the rPPG cost function can be used to process the predicted rPPG signal and compare it with the processed ground truth "soft labels" to compute the prediction error or gradient used for training. In one or more examples, Figure 4 The distribution with diagonal lines indicates soft frequency labels. An example. Once the labels and predictions are together in the frequency domain, three objective functions (404) can be used to provide feedback to the neural network for training. The first objective function can be a weakly supervised objective, which can be defined as The squared Wasserstein distance between two distributions in a frequency band is shown below: Equation (5):

[0073] In one or more examples, the second objective can be the signal-to-noise ratio (SNR), which is defined as the power centered at the peak frequency relative to the lower cutoff frequency and the upper cutoff frequency of the physiological signal (respectively...). and The ratio of the sum of the power of the two is shown below: Equation (6):

[0074] In one or more examples, the third objective function can be the independent power ratio (IPR), which is defined as the lower cutoff frequency and the upper cutoff frequency of the physiological signal (respectively...). and The ratio of the power of the two components to the total power of the signal is shown below: Equation (7):

[0075] Figure 5 An example block diagram of a system for training using multimodal soft labels according to an embodiment of the present disclosure is shown.

[0076] In the first stage 502, video of the object can be captured. In the second stage 504, preprocessing can be performed on the captured image. For example, to simplify the learning process for the network, the image can be detected, cropped, and reduced in size before each forward pass. Next, the minimum and maximum landmark positions can be identified along the x and y axes. The face can be cropped with margins along the edges (e.g., 6% margin on the sides and 22% margin on the top and bottom). After cropping the face, the image can be downsampled to 50×70 pixels using bilinear interpolation, which can be an aspect ratio similar to the face size. This process can be applied to all video frames, resulting in a tensor X∈RT×70×50×C, where T is the number of frames and C is the channel dimension (e.g., RGB).

[0077] In the third stage, 506, the pre-processed video is fed to a model. In one or more examples, this model is a neural network (e.g., a 3D convolutional neural network). The neural network model takes a segment X (e.g., the pre-processed video) as input and predicts a temporal signal with the same number of samples. In stage 506, each "block" can perform 3D convolution, batch normalization, and ReLU activation. In stage 506, the "Conv10" layer can be a 1×1 convolutional layer to be compressed into a single output channel without any activation function.

[0078] The output of stage 506 can be the rPPG time signal 508 converted to the frequency domain. In stage 510, the power spectrum of the rPPG time signal can be compared with the soft tag in the frequency domain, where the loss function can be calculated in stage 512.

[0079] Each training segment can have a corresponding frequency label for the pulse rate. Since the model predicts the waveform, and then the waveform is transformed to the frequency domain, the label can be defined as a Gaussian distribution centered on the true pulse rate. The standard deviation of the Gaussian label can correspond to the "softness". When the standard deviation is close to zero, the label becomes one-soft, and as the standard deviation increases, the label becomes more dispersed across the supported frequency range. The standard deviation can be set as the beamwidth (sampling rate divided by the number of samples) divided by an integer (e.g., from 1 to 6).

[0080] In one or more examples, the model can be fine-tuned to suit the user's specific environment. For example, such as Figure 3 As shown, the output of cost function 340 is provided to model 338 to adjust model 338 based on the target user's environment. Specifically, a general prediction model may not capture all variations in the user and environment. Many factors influence the intensity of light captured from the target user, and these are physiologically relevant. For example, the target user's skin type, ambient light intensity, light temperature, distance, etc., may affect the estimated rPPG model. The model may require a comprehensive dataset to model these variations.

[0081] According to one or more embodiments, "soft labels" are integrated with an "rPPG cost function" to perform fine-tuning or calibration of model 338 at runtime. Based on these features, a general model (e.g., 308) pre-trained offline (e.g., in the cloud) can be fine-tuned on the target user's device based on one or more samples of physiological parameter values ​​from another source (e.g., a wearable device such as a smartwatch). The rPPG cost function can evaluate the error between the "soft labels" (such as HR from a smartwatch) and the measured vital signs from rPPG signals to adjust the weights in the rPPG model.

[0082] According to one or more embodiments, the rPPG prediction model has specific layers and neurons that are tunable during the inference phase, wherein fine-tuning can be performed using soft labels from external sources given an rPPG cost function. Model fine-tuning can be performed based on one or more of the following methods.

[0083] In one or more examples, for fine-tuning, one or more parameters corresponding to one or more layers of the model in the upper layers are adjustable, while the parameters in the lower layers are fixed. Additional layers can be added for "personalization" to the target user or the target user's environment. Neurons that can be tuned or adjusted can be automatically determined by an optimization algorithm based on the input, the target "soft label," and the prediction error.

[0084] In one or more examples, for fine-tuning, parameters corresponding to one or more layers of the model responsible for capturing the influence of specific user or environmental factors on the rPPG signal can be adjustable, while parameters in other layers remain fixed. Layers and neurons that can be tunable or adjustable can be determined offline based on analysis of the rPPG model. For example, adjustment of phototemperature can be performed in the middle layers of the model, where the correlation between RGB frequency features is captured to reconstruct the rPPG signal. In one or more examples, the correlation of RGB color intensity signals in the adjustable model is captured to distinguish physiological signals from noise in the lower middle layers when calibrating user skin factors.

[0085] The model can be optimized periodically using available sparse "soft-label" samples from the truth setter. Triggering events for fine-tuning the model can depend on prediction error, changes in user / environmental factors, or be based on predetermined timing (e.g., fine-tuning the model hourly). In one or more examples, reinforcement learning methods can be used to help accelerate the calibration of personalized models.

[0086] Figure 6 An example system for vision-based vital sign monitoring using multimodal soft tags, according to embodiments of this disclosure, is shown. (The following is not repeated...) Figure 3 The description of the common components of the system is shown in the figure.

[0087] Figure 6 The system includes a training phase 600 and an inference phase 630. The training phase 600 pre-trains a multimodal rPPG model 608, and the inference phase 630 utilizes and fine-tunes a multimodal rPPG model 638. Model 638 may correspond to the pre-trained model 608 prior to the fine-tuning.

[0088] According to one or more embodiments, “soft tags” from an external source can be used as additional context in rPPG models (e.g., models 608 and 638). This additional context, as another modality, helps distinguish physiological signals from noise signals in the video. The “soft tags” can be converted into synthetic time-series physiological signals. This converted signal can be a reference signal used by the model when converting RGB channels to rPPG signals. In one or more examples, these signals can also be used as additional modalities when the external source provides raw PPG or ECG signals.

[0089] like Figure 6 As shown, in one or more examples, the same architecture as the unimodal models (e.g., models 308 and 338) can be used for multimodal rPPG models (e.g., models 608 and 638). In one or more examples, channels corresponding to physiological signals can be combined with other RGB channels and fed into the 3D CNN. If more sources are available, more channels can be appended to each other.

[0090] In one or more examples, the model training objective may be similar to the previous components. "Soft labels" can again be used to create a distribution of physiological values ​​to assess the error against rPPG-based physiological parameters.

[0091] During the inference phase 630, additional data sources can be used to generate enhanced rPPG signals for vital sign monitoring, resulting in a wider range of values ​​or an expanded list of physiological parameters.

[0092] Figure 7 An example system for vision-based vital sign monitoring using multimodal soft tags according to embodiments of the present disclosure is illustrated. In system 700, data from one or more PPG sensors or ECG patches 702 can be collected. This data may come from a target user or from data from other users besides the target user. A physiological tag 704 and raw PPG or ECG signals 706 can be determined from the data from the PPG sensors or ECG patches. Furthermore, periodic / physiological signals 708 can be synthesized from the physiological tag 704. A time-series signal 710 and RGB data 712 corresponding to video captures of the target user can be provided to a model 638. The physiological tag 704 and the output of model 638 can be provided to a cost function 340. Furthermore, the output of cost function 340 can be used to fine-tune model 638.

[0093] Figure 8 An example system for vision-based vital sign monitoring using multimodal soft tags according to embodiments of the present disclosure is shown. Figure 8 The system includes Figure 3The model consists of a training phase 300 and an inference phase 830. In the inference phase 830, before using the RGB signals in the rPPG model to predict the final rPPG signal, the RGB signals can be preprocessed using additional context from an external source. For example, another source measuring HR or RR samples can provide additional context about a narrower range of user vital signs at that time. This additional context can help filter the RGB signals to a more specific range corresponding to the physiological signal, eliminating noise signals. Without this additional context, there is no prior information about the range of physiological parameter values; therefore, the method may sacrifice measurement accuracy to find physiological signals within all possible ranges.

[0094] In one or more examples, in RGB enhancement 832, various filtering methods, bandpass filters, Kalman filters, etc., can be utilized in the preprocessing stage before feeding the RGB signal to the rPPG model 338. In one or more examples, another ML model can also be developed for the purpose of enhancing rPPG using additional physiological value sources. Combining the above steps of rPPG enhancement within a single rPPG model yields the "multimodal" rPPG model as described above.

[0095] The embodiments have been described above, and as shown in the accompanying drawings, the embodiments are illustrated in the form of blocks that perform the described functions. These blocks may be physically implemented by analog and / or digital circuitry including one or more of logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, etc., and may also be implemented or driven by software and / or firmware (configured to perform the functions or operations described herein). The circuitry may be embodied, for example, in one or more semiconductor chips, or on a substrate support (such as a printed circuit board). The circuitry included in the block may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware (for performing some functions of the block) and a processor (for performing other functions of the block). Each block of the embodiments may be physically divided into two or more interactive and discrete blocks. Similarly, the blocks of the embodiments may be physically combined into more complex blocks.

[0096] While this disclosure has described several non-limiting embodiments, variations, substitutions, and various alternative equivalents fall within the scope of this disclosure. Therefore, it will be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and are thus within its spirit and scope.

[0097] According to one or more embodiments, a method includes: acquiring a training dataset comprising a video of a test subject's face and corresponding sensor data (e.g., from a PPG or ECG sensor), wherein the video and sensor data are not time-synchronized; determining a Gaussian distribution based on the sensor data to be used as a soft label for training a machine learning model; providing the video of the test subject's face as input to the machine learning model, the machine learning model being configured to predict an rPPG signal based on the video; and training the machine learning model using a loss function comprising the difference between the Gaussian distribution based on the sensor data and the FFT of the predicted rPPG.

[0098] The method also includes: during the inference phase, receiving sensor data (e.g., heart rate measured using a smartwatch) collected while the user is capturing video for rPPG prediction; and fine-tuning the machine learning model to suit the user by updating parameters in one or more layers of the machine learning model, wherein a Gaussian distribution determined based on the sensor data is used as a soft label for the predicted rPPG.

[0099] The above disclosure also covers the following listed embodiments: (1) A method executed by at least one processor, the method comprising: obtaining an image of an object; preprocessing the image of the object; inputting the preprocessed image into a machine learning model trained according to a first frequency distribution, wherein the first frequency distribution corresponds to a first ground truth obtained from one or more sensors performing vital sign measurements on one or more test objects; and obtaining an estimate of a signal corresponding to the vital sign measurements of the object from the machine learning model.

[0100] (2) The method as described in feature (1) further includes: obtaining a second frequency distribution corresponding to a second true value of a vital sign measurement from one or more sensors performing vital sign measurements on the object; converting an estimate of a signal corresponding to a vital sign measurement of the object into a frequency domain signal; determining an error between the frequency domain signal and the second frequency distribution; and updating a machine learning model based on the determined error.

[0101] (3) The method as described in feature (1) or feature (2), wherein the first frequency distribution is a Gaussian distribution centered on the first true value and having a first standard deviation, wherein the first standard deviation is a function of N frames sampled per second f frames, wherein N and f are positive integers.

[0102] (4) The method as described in feature (2) or feature (3), wherein the first frequency distribution is a Gaussian distribution centered on a first true value and having a second standard deviation, wherein the second standard deviation is a function of N frames sampled per second f frames, wherein N and f are positive integers.

[0103] (5) The method as described in feature (3) or feature (4), wherein the operation of determining the error includes: determining the mean square error (MSE) loss between the frequency domain signal and the second frequency distribution.

[0104] (6) The method as described in feature (5), wherein the operation of determining the error further includes: determining the signal-to-noise ratio (SNR) based on the ratio of the power centered at the peak frequency of the frequency domain signal to the sum of the power between the lower cutoff frequency and the upper cutoff frequency of the frequency domain signal.

[0105] (7) The method as described in feature (6), wherein the operation of determining the error further includes: determining the irrelevant power ratio (IPR) based on the ratio of the power between the lower cutoff frequency and the upper cutoff frequency of the frequency domain signal to the total power of the frequency domain signal.

[0106] (8) The method of any one of features (1) to (7), wherein the operation of preprocessing the image of the object includes: detecting the region of interest of the image object; and adjusting the size of the region of interest of the image object.

[0107] (9) The method as described in feature (8), wherein the region of interest is at least a portion of the face of the object.

[0108] (10) The method of any one of features (1) to (9), wherein the vital signs measurement is one of pulse rate, blood pressure, and oxygen saturation level.

[0109] (11) The method described by any of features (1) to (10), wherein the machine learning model is a three-dimensional (3D) convolutional neural network (CNN).

[0110] (12) An apparatus comprising: a memory; processing circuitry incorporated into the memory, the processing circuitry being configured to: acquire an image of an object; preprocess the image of the object; input the preprocessed image into a machine learning model trained according to a first frequency distribution, wherein the first frequency distribution corresponds to a first ground truth obtained from one or more sensors performing vital sign measurements on one or more test objects; and obtain an estimate of a signal corresponding to the vital sign measurements of the object from the machine learning model.

[0111] (13) The device as described in feature (12), wherein the processing circuitry is further configured to: obtain a second frequency distribution corresponding to a second true value of a vital sign measurement from one or more sensors performing vital sign measurements on an object; convert an estimate of a signal corresponding to a vital sign measurement of the object into a frequency domain signal; determine the error between the frequency domain signal and the second frequency distribution; and update a machine learning model based on the determined error.

[0112] (14) The device as described in feature (12) or feature (13), wherein the first frequency distribution is a Gaussian distribution centered on a first true value and having a first standard deviation, wherein the first standard deviation is a function of N frames sampled per second f frames, wherein N and f are positive integers.

[0113] (15) The device as described in feature (13) or feature (14), wherein the first frequency distribution is a Gaussian distribution centered on a first true value and having a second standard deviation, wherein the second standard deviation is a function of N frames sampled per second f frames, wherein N and f are positive integers.

[0114] (16) The device as described in feature (14) or feature (15), wherein, in order to determine the error, the processing circuit is further configured to: determine the mean square error (MSE) loss between the frequency domain signal and the second frequency distribution.

[0115] (17) The device as described in feature (16), wherein, in order to determine the error, the processing circuitry is further configured to determine the signal-to-noise ratio (SNR) based on the ratio of the power centered at the peak frequency of the frequency domain signal to the sum of the power between the lower cutoff frequency and the upper cutoff frequency of the frequency domain signal.

[0116] (18) The device as described in feature (17), wherein, in order to determine the error, the processing circuit is further configured to determine the irrelevant power ratio (IPR) based on the ratio of the power between the lower cutoff frequency and the upper cutoff frequency of the frequency domain signal to the total power of the frequency domain signal.

[0117] (19) The device as described in any of features (12) to (18), wherein, in order to preprocess the image of the object, the processing circuitry is further configured to: detect the region of interest of the image object; and adjust the size of the region of interest of the image object.

[0118] (20) A non-transitory computer-readable medium storing instructions which, when executed by a processor, cause the processor to perform a method comprising: obtaining an image of an object; preprocessing the image of the object; inputting the preprocessed image into a machine learning model trained according to a first frequency distribution, wherein the first frequency distribution corresponds to a first ground truth obtained from one or more sensors performing vital sign measurements on one or more test objects; and obtaining an estimate of a signal corresponding to the vital sign measurements of the object from the machine learning model.

Claims

1. A method executed by at least one processor, the method comprising: Get the image of the object; The image of the object is preprocessed; The preprocessed image is input into a machine learning model trained according to a first frequency distribution, wherein the first frequency distribution corresponds to a first ground truth obtained from one or more sensors that perform vital sign measurements on one or more test subjects. as well as Estimates of the signals corresponding to the vital sign measurements of the object are obtained from the machine learning model.

2. The method of claim 1, further comprising: A second frequency distribution corresponding to a second true value of the vital signs measurement is obtained from one or more sensors that perform the vital signs measurement on the object; The estimated value of the signal corresponding to the vital sign measurement of the object is converted into a frequency domain signal; Determine the error between the frequency domain signal and the second frequency distribution; and The machine learning model is updated based on the determined error.

3. The method as described in claim 1, wherein, The first frequency distribution is a Gaussian distribution centered on the first true value and having a first standard deviation, wherein the first standard deviation is a function of N frames sampled per second (f frames), where N and f are positive integers.

4. The method of claim 2, wherein, The first frequency distribution is a Gaussian distribution centered on the first true value and having a second standard deviation, wherein the second standard deviation is a function of N frames sampled per second (f frames), where N and f are positive integers.

5. The method of claim 3, wherein, The operation of determining the error includes: Determine the mean square error (MSE) loss between the frequency domain signal and the second frequency distribution.

6. The method of claim 5, wherein, The operation of determining the error further includes: The signal-to-noise ratio (SNR) is determined based on the ratio of the power centered at the peak frequency of the frequency domain signal to the sum of the power between the lower and upper cutoff frequencies of the frequency domain signal.

7. The method of claim 6, wherein, The operation of determining the error further includes: The irrelevant power ratio (IPR) is determined based on the ratio of the power between the lower and upper cutoff frequencies of the frequency domain signal to the total power of the frequency domain signal.

8. The method of claim 1, wherein, The preprocessing operations for the image of the object include: Detecting the region of interest of the image object; and Adjust the size of the region of interest of the image object.

9. The method of claim 8, wherein, The region of interest is at least a portion of the object's face.

10. The method of claim 1, wherein, The vital signs measurement is one of pulse rate, blood pressure, and oxygen saturation level.

11. The method of claim 1, wherein, The machine learning model is a three-dimensional 3D convolutional neural network (CNN).

12. An apparatus comprising: Memory; Processing circuitry is incorporated into the memory, and the processing circuitry is configured to: Get the image of the object; The image of the object is preprocessed; The preprocessed image is input into a machine learning model trained according to a first frequency distribution, wherein the first frequency distribution corresponds to a first ground truth obtained from one or more sensors that perform vital sign measurements on one or more test subjects. as well as Estimates of the signals corresponding to the vital sign measurements of the object are obtained from the machine learning model.

13. The device as claimed in claim 12, wherein, The processing circuit is further configured to: A second frequency distribution corresponding to a second true value of the vital signs measurement is obtained from one or more sensors that perform the vital signs measurement on the object; The estimated value of the signal corresponding to the vital sign measurement of the object is converted into a frequency domain signal; Determine the error between the frequency domain signal and the second frequency distribution; and The machine learning model is updated based on the determined error.

14. The device as claimed in claim 12, wherein, The first frequency distribution is a Gaussian distribution centered on the first true value and having a first standard deviation, wherein the first standard deviation is a function of N frames sampled per second (f frames), where N and f are positive integers.

15. The device as claimed in claim 13, wherein, The first frequency distribution is a Gaussian distribution centered on the first true value and having a second standard deviation, wherein the second standard deviation is a function of N frames sampled per second (f frames), where N and f are positive integers.