Segmented voice processing system
A system for extracting health metrics from structured vowel sounds in telephony systems addresses AGC distortions by using time-domain and frequency-domain analyses, ensuring accurate heart rate and variability measurements with real-time feedback.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- VITAL AUDIO SYSTEMS INC
- Filing Date
- 2025-11-06
- Publication Date
- 2026-05-15
AI Technical Summary
Conventional telephony systems apply automatic gain control (AGC) and other modulations that distort speech signals, complicating the extraction of health metrics like heart rate from voice data, and lack robust mechanisms for segmenting speech and validating signal integrity.
A system that prompts users to articulate structured vowel sounds, implementing time-domain and frequency-domain analyses to extract fundamental frequency and cardiac modulations, with signal integrity checks and real-time alerting to generate high-quality health metrics.
Minimizes carrier-induced distortions, enabling accurate extraction of heart rate and variability metrics by bypassing AGC and other modulations, and providing real-time feedback for improved signal quality.
Smart Images

Figure US2025054424_15052026_PF_FP_ABST
Abstract
Description
VITAL01-PCT / 142238-0103 PATENTSEGMENTED VOICE PROCESSING SYSTEMCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 718,079, filed November 8, 2024, which is incorporated by reference in its entirety.TECHNICAL FIELD
[0002] The present disclosure generally relates to systems and methods for detecting health metrics using audio signals. More specifically, embodiments pertain to computer-implemented techniques for extracting health metrics from segmented voiced-speech data for remote-user vital evaluation and monitoring.BACKGROUND
[0003] The healthcare industry evolution has recently been catalyzed with innovative technologies. This has created a secondary avenue for healthcare delivery namely, telehealth or telemedicine using telecommunications and computing technologies for remote consultations and data collection and analysis. For instance, telephony-based communications technologies allow for remote consultations and interpersonal interactions for collecting information or vitals that help healthcare providers gather and assess key vitals quickly, easily, and accurately.
[0004] It is often beneficial to gather information about the patient through a telephony communication using speech-based analysis of the patient’s speech. However, a significant challenge in measuring vitals, such as heart rate, through speech-based methods is that speech modulation techniques and operations, such as automatic gain control (AGC), are often applied automatically by carriers, devices, or other networking elements. These existing approaches to speech modulation and handling voice signal data in conventional telephony systems (e.g., cellular systems) impact the audio signal data in a manner that complicates or prevents the capacity to determine or collect healthcare vitals using the patient’s speech, such as heart rate computation. These conventional telephony systems and prior modulation techniques are often not under the control of the user or a developer who is trying to disable such modulation techniques.Page 1 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENTSUMMARY
[0005] Embodiments described herein address the shortcomings in existing technologies mentioned above and may provide any number of additional or alternative benefits. Embodiments described herein include systems and methods for capturing voice signal data and producing digitized voice signals without sacrificing qualities or characteristics needed for computing health related information about a speaking end-user. Embodiments described herein provide a computer- implemented system and method for detecting physiological parameters from segmented audio signals. The system prompts users to articulate structured vowel sounds, enabling robust extraction of fundamental frequency and cardiac-related modulations while minimizing carrier-induced distortions. The system implements signal processing algorithms, including time-domain and frequency-domain analyses, to generate health metrics such as heart rate and heart rate variability. The system may implement various process for quality control, including signal integrity checks, real-time alerting, and adaptive reacquisition protocols to generate high-quality data and reliable outputs.
[0006] Embodiments may include a computer-implemented method for detecting physiological parameters of using audio signals, the method including: obtaining, by a computer, an input audio signal including a plurality of voiced sound segments for a speaker and one or more acoustic pauses; determining, by the computer, a fundamental frequency for one or more voiced sound segments of the input audio signal based on at least one of a time-domain periodicity of the one or more voiced sound segments or a frequency-domain harmonic-spacing of the one or more voiced sound segments; identifying, by the computer, a periodic modulation corresponding to cardiac pulses of the speaker for the fundamental frequency, the periodic modulation based upon a plurality of perturbations of one or more acoustic characteristics of the fundamental frequency; and generating, by the computer, one or more physiological parameters of the speaker for the input audio signal based on the periodic modulation of the fundamental frequency of the input audio signal.
[0007] In some aspects, the techniques described herein relate to a method, further including determining, by the computer, one or more signal integrity criteria for the input audio signal, and wherein the computer determines the one or more physiological parameters in responsePage 2 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT to the computer determining that the input audio signal satisfies the one or more signal integrity criteria.
[0008] In some aspects, the techniques described herein relate to a method, wherein determining the one or more signal integrity criteria includes detecting, by the computer, each voiced sound segment of the plurality of voiced sound segments in the input audio signal.
[0009] In some aspects, the techniques described herein relate to a method, wherein determining the one or more signal integrity criteria includes determining, by the computer, that each voiced sound segment satisfies a minimum energy threshold;
[0010] In some aspects, the techniques described herein relate to a method, wherein determining the one or more signal integrity criteria includes determining, by the computer, that each voiced sound segment corresponds with the fundamental frequency.
[0011] In some aspects, the techniques described herein relate to a method, wherein determining the one or more signal integrity criteria includes determining, by the computer, that at least one voiced sound segment corresponds with a harmonics characteristic for a vowel sound.
[0012] In some aspects, the techniques described herein relate to a method, wherein the one or more acoustic characteristics include at least of an amplitude, a frequency, an amplitude envelope, or an instantaneous frequency; and wherein the computer identifies the periodic modulation of the fundamental frequency corresponding to the cardiac pulses based upon at least one of the amplitude envelope or the instantaneous frequency of the fundamental frequency.
[0013] In some aspects, the techniques described herein relate to a method, further including generating, by the computer, a time-frequency representation of the input audio signal using one or more transform functions and the input audio signal.
[0014] In some aspects, the techniques described herein relate to a method, wherein the one or more physiological parameters includes at least one of a heart rate or heart rate variability.
[0015] In some aspects, the techniques described herein relate to a method, further including generating, by the computer, an alert for display at a user interface in response toPage 3 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT determining that a value of a physiological parameter satisfies an alert threshold corresponding to the physiological parameter.
[0016] Embodiments may include a system for detecting physiological parameters of using audio signals, the system including: a computer including at least one processor, configured to: obtain an input audio signal including a plurality of voiced sound segments for a speaker and one or more acoustic pauses; determine a fundamental frequency for one or more voiced sound segments of the input audio signal based on at least one of a time-domain periodicity of the one or more voiced sound segments or a frequency-domain harmonic-spacing of the one or more voiced sound segments; identify a periodic modulation corresponding to cardiac pulses of the speaker for the fundamental frequency, the periodic modulation based upon a plurality of perturbations of one or more acoustic characteristics of the fundamental frequency; and generate one or more physiological parameters of the speaker for the input audio signal based on the periodic modulation of the fundamental frequency of the input audio signal.
[0017] In some aspects, the techniques described herein relate to a system, wherein the computer is further configured to determine one or more signal integrity criteria for the input audio signal, and wherein the computer determines the one or more physiological parameters in response to the determining that the input audio signal satisfies the one or more signal integrity criteria.
[0018] In some aspects, the techniques described herein relate to a system, wherein when determining the one or more signal integrity criteria the computer is further configured to detect each voiced sound segment of the plurality of voiced sound segments in the input audio signal.
[0019] In some aspects, the techniques described herein relate to a system, wherein when determining the one or more signal integrity criteria the computer is further configured to determine that each voiced sound segment satisfies a minimum energy threshold;
[0020] In some aspects, the techniques described herein relate to a system, wherein when determining the one or more signal integrity criteria the computer is further configured to determine that each voiced sound segment corresponds with the fundamental frequency.
[0021] In some aspects, the techniques described herein relate to a system, wherein when determining the one or more signal integrity criteria the computer is further configured toPage 4 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT determine that at least one voiced sound segment corresponds with a harmonics characteristic for a vowel sound.
[0022] In some aspects, the techniques described herein relate to a system, wherein the one or more acoustic characteristics include at least of an amplitude, a frequency, an amplitude envelope, or an instantaneous frequency; and wherein the computer identifies the periodic modulation of the fundamental frequency corresponding to the cardiac pulses based upon at least one of the amplitude envelope or the instantaneous frequency of the fundamental frequency.
[0023] In some aspects, the techniques described herein relate to a system, the computer is further configured to generate a time-frequency representation of the input audio signal using one or more transform functions and the input audio signal.
[0024] In some aspects, the techniques described herein relate to a system, wherein the one or more physiological parameters includes at least one of a heart rate or heart rate variability.
[0025] In some aspects, the techniques described herein relate to a system, the computer is further configured to generate an alert for display at a user interface in response to the computer determining that a value of a physiological parameter satisfies an alert threshold corresponding to the physiological parameter.
[0026] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the invention as claimed.BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The present disclosure can be better understood by referring to the following figures. The components in the figures are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosure. In the figures, reference numerals designate corresponding parts throughout the different views.
[0028] FIG. 1 shows components of a system for processing voice audio signals and performing health diagnostics, according to embodiments.Page 5 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT
[0029] FIG. 2 depicts dataflow amongst components of a system for segmented voice processing, according to embodiments.
[0030] FIG. 3 shows dataflow amongst components of a system for analyzing voice signals in audio data to generate health metrics for speaker-users, according to embodiments.
[0031] FIG. 4 shows operations of a method for performing speech data analysis according to a segmented articulation process, according to embodiments.
[0032] FIGS. 5A-5B depict a pulse sequence and pulse-spectrum for a fundamental frequency for an ideal audio signal generated by the analytics server, according to embodiments.
[0033] FIGS. 6A-6B depict a pulse sequence and pulse-spectrum generated by an analytics server for a fundamental frequency for moderately variability in timing for an audio signal, according to embodiments.
[0034] FIGS. 7A-7B depict a pulse sequence and pulse-spectrum generated by an analytics server for a fundamental frequency for elevated timing variability for an audio signal, according to embodiments.
[0035] FIGS. 8A-8B depict a pulse sequence and pulse-spectrum generated by an analytics server for a fundamental frequency for substantial timing variability for an audio signal, according to embodiments.
[0036] FIG. 9 is a flowchart of a computer-implemented method for detecting physiological parameters (e g., heart rate) or other health metrics data outputs using audio signals, according to embodiments.DETAILED DESCRIPTION
[0037] Reference will now be made to the illustrative embodiments illustrated in the drawings, and specific language will be used here to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended. Alterations and further modifications of the inventive features illustrated here, and additional applications of the principles of the inventions as illustrated here, which would occur to a person skilled in the relevant art and having possession of this disclosure, are to be considered within the scope of the invention.Page 6 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT
[0038] The present disclosure addresses longstanding technological shortcomings in the field of remote physiological monitoring and telehealth assessment using audio signals. Conventional systems for extracting health metrics from speech audio, such as heart rate, heart rate variability, and related parameters, have been several technical shortcomings or challenges. For instance, existing telephony and voice communication platforms frequently apply automatic gain control (AGC), noise suppression, and other carrier-induced modulations that distort the original speech signal. These artifacts can obscure subtle physiological modulations, such as cardiac-induced amplitude or frequency perturbations, rendering conventional analysis techniques unreliable or inaccurate. Additionally, prior approaches often lack robust mechanisms for segmenting speech, validating signal integrity, and distinguishing genuine physiological features from artifacts introduced by the communication channel, user movement, or environmental noise.
[0039] As mentioned, a common challenge is AGC modulation feature applied by communication carrier networks, devices, or other elements. AGC modulates the speech signal, complicating heart rate computation. This modulation is often beyond the control of the patients, care providers, or third-party software developers that build software programs (e.g., telehealth software) that compute or capture healthcare vitals for patients.
[0040] For instance, to capture healthcare vitals, such as heart rate, a conventional system captures audio signal data containing the patient’s audio speech signal and analyzes certain aspects of ta continuous articulated vocal sound. As an example, when a patient provides a continuous “ah” sound for about 8 seconds, the continuously articulated vowel sound can trigger the AGC modulation of a telephony system, which significantly attenuates the signal’s amplitude after 2-3 seconds. This attenuation results in a distorted speech signal, making heart rate computation challenging. Other signal modulations, such as noise suppression and echo cancellation, are often applied by communication layers and carrier services further complicate the process of heart rate computation. These modulations can further distort the original speech signal.
[0041] Separately, some patients could find it difficult to hold a tone for the required duration of continuous articulation. This patient-side difficulty adds another potential challenge of obtaining a clean and consistent speech signal for determining patient vital data, such as heart rate measurement. Further, many conventional systems do not provide real-time feedback or adaptive reacquisition protocols, resulting in missed opportunities for clinical intervention or user guidancePage 7 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT when signal quality is insufficient. This could be useful when patients have difficulty holding a continuous tone. Moreover, the absence of integrated metadata handling, such as sampling rate, codec type, and session identifiers, further complicates normalization and cross-device compatibility, limiting the scalability and clinical utility of such conventional technologies.
[0042] Embodiments disclosed herein overcome these limitations by introducing a comprehensive system and method for detecting physiological parameters using segmented voiced-sound segments from input audio signals. The system prompts end-user speakers to vocally produce specific vowel sounds in structured intervals. These vowel sounds avoid triggering modulation functions, thereby minimizing the impact of AGC and other carrier artifacts. The system implements various signal processing functions, including time-domain and frequencydomain analyses, short-time Fourier transform (STFT), and discrete Fourier transform (DFT), to extract fundamental frequency components and identify periodic modulations corresponding to the speaker’s cardiac pulses. The system may implement various thresholds or signal integrity criteria, such as spectral energy level thresholds, voiced sound segment detection, and harmonic structure validation such that the system uses high-quality data for generating physiological metric.
[0043] In addition, the system may implement features for real-time alerting and feedback mechanisms, enabling immediate user guidance and reacquisition of audio samples when signal quality is insufficient or when physiological parameters fail to satisfy expected thresholds. In some cases, the system captures and references metadata associated with each audio sample for normalization and preprocessing, allowing for consistent and accurate analysis across diverse devices and environments. The system may generate a wide range of physiological parameters, such as heart rate, heart rate variability, respiratory rate, phonation quality, and proxies for lung capacity and oxygen saturation, thereby enhancing the reliability, accuracy, and clinical relevance of remote health assessments.
[0044] FIG. 1 shows components of a system 100 for processing voice audio signals and performing health diagnostics, according to embodiments. The system 100 includes an analytics system 101 for remote patient monitoring and acoustic data analysis, caller end-user devices 114a- 114d (generally referred to as caller devices 114), telephony carrier networks 130a-130b, which may include originating carrier networks 130a and terminating carrier networks 130b (generally referred to as telephony networks 130 or carrier networks 130), and healthcare software applicationPage 8 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT provider systems (sometimes referred to as a provider system 110). The system 100 may include hardware and software components hosting and executing components of cloud middleware service 133, which may be hosted and executed at devices of the telephony network 130 or at devices of a separate cloud-based computing system (not shown). The various components of the system 100 may communicate with one another via one or more networks 105, which may include computing networks for inter-device communications or telephony -based communications via the one or more telephony carrier networks 130. The system 100 may include any number of databases 104a-104b (generally referred to as databases 104), which may be components of and hosted by the analytics system 101 or the provider system 110. The analytics system 101 may include analytics servers 102 and analytics databases 104a. The provider system 110 may include provider servers 111, provider databases 104b, and agent devices 113 that may be operated by called users.
[0045] Embodiments may comprise additional or alternative components or omit certain components from what is shown in FIG. 1, yet still fall within the scope of this disclosure. For ease of description, FIG. 1 shows only one instance of various aspects the illustrative embodiment. However, other embodiments may comprise any number of the components. For instance, it will be common for there to be multiple provider systems 110, or for an analytics system 101 to have multiple analytics servers 102. Although FIG. 1 shows the illustrative system 100 having only a few of the various components, embodiments may include or otherwise implement any number of devices capable of performing the various features and tasks described herein. For example, in the illustrative system 100, an analytics server 102 is shown as a distinct computing device from an analytics database 104a; but in some embodiments the analytics database 104a may be integrated into the analytics server 102, such that these features are integrated within a single device.
[0046] The illustrative system 100 of FIG. 1 comprises various network infrastructures 101, 110, including the analytics system 101 and the provider system 110. The network infrastructures 101, 110 may be a physically and / or logically related collection of devices owned or managed by some enterprise organization, where the devices of each infrastructure 101, 110 are configured to provide the intended services of the particular infrastructure 101, 110 and responsible organization.
[0047] The caller devices 114 may be any electronic device comprising hardware and software components that callers operate to place a call to callee-destinations (e.g., providerPage 9 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT systems 110) via one or more carrier networks 130. Non-limiting examples of the caller devices 114 include landline phones 114a or mobile phones 114b. The caller devices 114 are not limited to telecommunications-oriented, telephony-based devices (e.g., telephones 114a, 114b). As an example, a caller device 114 may include an electronic device comprising a processor and / or software, such as a computer 114c or loT device 114d, configured to implement voice-over-IP (VoIP) or other telephony-based telecommunications protocols. As another example, a caller device 114 may include an electronic device comprising a processor and / or software, such as an loT device 114d (e.g., voice assistant device, “smart device”), capable of utilizing telecommunications features of a paired or otherwise internetworked caller device 114, such as mobile phone 114b. A caller device 114 may comprise hardware (e.g., microphone) and / or software (e.g., codec) for detecting and converting sound (e.g., caller’s spoken utterance, ambient noise) into electrical audio signals. The caller device 114 then transmits the audio signal, along with other forms of call data, according to one or more telephony or other communications protocols to a called, destination provider system 110 for an established telephone call.
[0048] The networks 105 may include hardware and software components for devicecommunications networks, such as TCP / IP -based or packet-based communications networks. The components of the one or more networks 105 may execute or perform programming and protocols for device communications for transmitting routing computing-network data packets for an IPbased networking or the like.
[0049] The networks 105 of the system 100 include one or more telephony networks 130 for telephony-based communications. In some cases, the caller devices 114 use the telephonybased networks 130 for communicating with the customer-facing service provider systems 110 or the analytics system 101 according to telephony and telecommunications protocols. The telephony network 130 includes hardware, and software capable of hosting, transporting, and exchanging call data, including audio data and metadata according to the particular telephony protocol(s). The telephony networks 130 may include hardware and software components for hosting, managing, and conducting telephony -based calls, between the caller devices 114 and the provider systems 110 or the analytics system 101. The telephony networks 130 can be any suitable telephonic system, including wireless and / or wired, and can utilize standard telephonic protocols. The telephony networks 130 may include, for example, various hardware and software components forPage 10 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT managing, hosting, and conducting the calls according to the telephony protocols and standards. The telephony networks 130 may include, for example, switches and exchanges of public switched telephone network (PSTN) infrastructure, and private branch exchange (PBX) infrastructures, among others. Non-limiting examples of telecommunications and / or computing networking hardware may include switches and trunks, among other additional or alternative hardware used for hosting, routing, or managing data communication, circuits, and signaling via the Internet or other device communications medium. Non-limiting examples of software and protocols for telecommunications may include SS7, SIGTRAN, SCTP, ISDN, and DNIS among other additional or alternative software and protocols used for hosting, routing, or managing telephone calls, circuits, and signaling. In some cases, the caller devices 114 execute software and protocols for performing telephony-based communications via one or more computing networks.
[0050] In some embodiments, the interceding in the call may be performed at the switchlevel or exchange-level using an existing PSTN infrastructure for capturing and forwarding call data (e.g., voice audio signals, metadata) to the analytics system 101. In some embodiments, the interceding is performed at the PBX-level, local to an entity of the provider system 110, such as a hospital or medical facility. The computing device of the telephony networks 130 may capture, convert, and forward the call data to the analytics system 101 via the one or more networks 105. In some embodiments, the call is conducted and hosted using Voice-over-Internet Protocol (VoIP) programming and protocols, such that the interceding is performed by computing devices for transmitting routing computing-network data packets for an IP network or the like. In some embodiments, the interceding is performed via an application on the caller device 114, such that the caller device 114 captures and forwards the call data to the analytics system 101 via the one or more networks 105.
[0051] Various different entities manage or organize the components of the telecommunications systems of the telephony networks 130, including carriers, networks, and exchanges, among others. For instance, the carrier networks 130 (e.g., originating carriers 130a, terminating carriers 130b) may host, operate, and administer components of the telephony networks 130 on behalf of the callers, caller devices 114, and the provider systems 110.
[0052] Computing devices of the system 100 (e.g., caller device 114, analytics server 102, computing device of the telephony networks 130) executes software programming for capturingPage 11 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT call data, including the voice signal audio, for the analytics system 101. Such computing device and / or the analytics server 102 may implement an application programming interface (API) or other software programming that intercedes into, or interfaces with, a call hosted by or conducted through the telephony networks 130. Using the API, the analytics server 102 obtains the various forms or types of call data, including the voice audio signal, for calls placed by end-users using the caller devices 114.
[0053] The cloud middleware system 133 comprises hardware and software components configured to facilitate secure and reliable transmission of segmented speech data between user devices 114, provider system 110, carrier networks 130, and the analytics system 101. The cloud middleware system 133 operates as an intermediary layer that manages communication protocols, data compression, and routing functions necessary for telephony-based or VoIP-based communications. In some embodiments, the cloud middleware system 133 includes one or more servers executing middleware services that provide session management, authentication, and data integrity checks for audio streams received from user devices 114.
[0054] The cloud middleware system 133 includes software programming to receive segmented audio signals from a carrier network 130 system and apply standardized encoding and packetization routines to prepare the audio data for downstream analysis. The middleware system 133 may implement APIs or communication frameworks that enable interoperability with multiple carrier networks 130, provider system 110, and the analytics system 101. In certain embodiments, the cloud middleware system 133 supports adaptive bitrate control and error correction techniques to maintain audio fidelity under variable network conditions.
[0055] The cloud middleware system 133 further provides a secure channel for transmitting audio data to the provider system 110 or analytics system 101. The functions of the cloud middleware system 133 may include encryption of audio packets, token-based authentication, and compliance with privacy standards for handling sensitive health-related data. The software of the middleware system 133 may also perform preliminary validation of audio data, such as confirming the presence of expected segmentation markers or metadata tags, before forwarding the audio data to the analytics system 101 or provider system 110 for further operations.Page 12 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT
[0056] The cloud middleware system 133 may also implement advanced security and compliance features for handling sensitive health-related audio data. These features may include encryption of audio packets during transit, token-based authentication for session initiation, and adherence to privacy regulations such as HIPAA. In some embodiments, the middleware system 133 performs integrity checks on segmented audio streams by validating segmentation markers and metadata tags before forwarding the audio data to the analytics system 101. The middleware system 133 may further support adaptive bitrate control and error correction techniques to maintain audio fidelity under variable network conditions.
[0057] The analytics system 101 is operated by an audio analytics service that provides various audio processing and health-related metrics services to the provider systems 110 of various customer organizations (e.g., hosting organization of health-related software applications; healthcare providers; insurance companies). A caller may place a telephone call to the provider system 110 of a particular organization using a caller device 114. When the caller device 114 originates the telephone call, the call data for the telephone call is generated by the caller device 114 and by the components of the one or more telephony networks 130, such as switches and trunks. An originating telephony carrier network 130a of the caller device 114 routes the call to a terminating carrier 130b of the destination provider system 110 according to the call metadata, such as telephony protocol messages (e.g., SIP INVITE). The terminating carrier 130b extracts and forwards the call data to the analytics system 101 and / or to the provider system 110 via the one or more networks 105. The components of the analytics system 101 (e.g., analytics server 102) execute various analytics operations using the call data, which generate and provide various health- related caller health metrics to the provider systems 110.
[0058] The analytics system 101 comprises an analytics server 102, an admin device (not shown), and an analytics database 104a. The call analytics server 102 may receive the call data from the telephony networks 130 via the networks 105. The analytics server 102 may also retrieve various types of end-user caller data from the analytics database 104a or provider database 104b, including information about the caller device 114 or the caller, among other types of information.
[0059] An analytics server 102 may be any computing device comprising various hardware components (e.g., at least one processor, non-transitory machine-readable storage) and software components, and capable of performing the various processes and tasks described herein. ThePage 13 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT analytics server 102 may be in network-communication with the analytics database 104 or other components of the system 100 to receive various types of data related to the caller or caller device 114. The analytics server 102 may receive the call data of particular calls from the carrier device or telephony devices of the telephony networks 130 via the networks 105. Although FIG. 1 shows a single analytics server 102, it should be appreciated that, in some embodiments, the analytics server 102 may include any number of computing devices. In some cases, the computing devices of the analytics server 102 may perform all or sub-parts of the processes and benefits of the analytics server 102. It should also be appreciated that, in some embodiments, the analytics server 102 may comprise any number of computing devices operating in a cloud computing or virtual machine configuration.
[0060] The analytics server 102 executes various software operations for generating health metrics or vitals for the call via voice signals in the call data of a phone call. For a particular call that originated at a caller device 114, the analytics server 102 obtains the call data from the caller device 114 or the computing device of the telephony network 130, through the one or more APIs of the analytics system 101. The health metrics generated by the analytics server 102 may be transmitted, via the one or more networks 105, to the provider system 110 for display on the agent devices 113. The transmission may occur via secure APIs or encrypted communication channels managed by the cloud middleware system 133. The GUI of the agent devices 113 may present real-time physiological parameters (e.g., heart rate, heart rate variability) and, in some cases, alert notifications generated by the analytics server 102. In some embodiments, the health metrics are stored in the provider database 104b for subsequent review and integration with patient care operations of the provider server 111 or other component of the provider system 110.
[0061] In some implementations, the analytics server 102 or provider server 111 generates an audio output containing speech prompt. The speech prompt requests the end-user patient utter a spoken sound for a set duration at the caller device 114. The caller device 114 or the telephony networks 130 captures or generates an audio file or audio signal as the call data (or portion of the call data) containing the voice signal of the end-user, which the caller device 114 or the telephony network 130 forwards to the analytics server 102 via the networks 105. The analytics server 102 obtains the audio signal calculates various health metrics for the end-user based on the voice signal in the call data. In some embodiments, the speech prompt instructs the patient to utter a sound forPage 14 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT a duration, for example, in the range of 1 second to 20 seconds, 5 seconds to 10 seconds, 6 seconds to 8 seconds, about 7 seconds, or any other suitable duration. In some cases, the speech prompt instructs the patient to utter a vowel sound for at least a portion of the voice signal.
[0062] The analytics server 102 may transmit the healthcare metrics, among other types of data, to the provider system 110 and / or the caller device 114 via the one or more networks 105. The healthcare metrics, as computed by the analytics server 102, may be presented at a user interface of the caller device 114 (e.g., end-user device of a calling speaker, computing device of a medical practitioner) and / or agent device 113 of the provider system 110.
[0063] In some implementations, the analytics server 102 or other computing device of the system 100 includes programming for automatically ceasing or rejecting the call data of an ongoing call or removing the analytics server 102 from the ongoing call. For instance, when the analytics server 102 transmits the healthcare metrics to the provider system 110 or the caller device 114, the analytics server 102 may automatically terminate any ongoing telephony communications or data communications that may be transmitted using the APIs.
[0064] The system 100 may include any number of databases 104, including the analytics database 104a of the analytics system 101 and the provider database 104b of the provider system 110, among other potential types of the databases 104. The analytics database 104a may include at least one computing device comprising hardware (e.g., at least one processor, non-transitory machine-readable storage) and software components for storing various types of data or information related to end-user or operational services. As an example, the analytics database 104a may store audio files, audio signals, and / or health metrics outputs or results, as received or generated by the analytics server 102. The analytics server 102 may reference the data in the analytics database 104a when generating the various health metrics for the end-user. As another example, the provider database 104b may store healthcare information or user profile information related to the end-users or caller devices 114. In some cases, the provider database 104b provides advantages in patient privacy and ease-of-use since health care information (e.g., vitals data) may be stored only on the analytics database 104b and not on the caller device 114. In some cases, at least one database 104 of the system 100 includes a portion of an electronic medical record (EMR) database.Page 15 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT
[0065] The provider system 110 compri ses hardware and software components configured to host healthcare-related applications and manage telephony-based interactions with end-users. The provider system 110 may include computing infrastructure for executing patient-facing and clinician-facing services, such as scheduling, telehealth consultations, and real-time health metrics display. In some embodiments, the provider system 110 operates as a secure environment for receiving call data from telephony networks 130 and transmitting health-related outputs generated by the analytics system 101. The provider system 110 may implement compliance frameworks for handling sensitive health information, including encryption protocols, authentication mechanisms, and audit logging. The provider system 110 may also provide interoperability with external healthcare systems, such as electronic health record (EHR) platforms, through standardized APIs or HL7 / FHIR interfaces.
[0066] The provider server 111 is a computing device of the provider system 110 responsible for managing inbound and outbound telephony sessions, coordinating with agent devices 113, and interfacing with the analytics system 101. The provider server 111 may execute software programming for call routing, session control, and integration with healthcare applications. In some embodiments, the provider server 111 hosts a web-based dashboard or portal that displays real-time health metrics received from the analytics system 101, enabling clinicians to monitor patient vitals during an active call. The provider server 111 may also perform authentication of caller devices 114, manage user profiles, and store operational data in the provider database 104b. The provider server 111 may include hardware components such as at least one processor, non-transitory machine-readable storage, and network interfaces for secure communication over networks 105. In certain implementations, the provider server 111 supports redundancy and load balancing to ensure high availability for telehealth services.
[0067] The agent devices 113 are computing devices operated by personnel of the provider system 110, such as healthcare practitioners, call center agents, or administrative staff. The agent devices 113 may include desktop computers, laptops, tablets, or other electronic devices capable of executing provider-facing applications and displaying health metrics generated by the analytics system 101. Each agent device 113 may present a graphical user interface (GUI) that provides realtime physiological parameters, such as heart rate and heart rate variability, along with alert notifications during an active telephony session. In some embodiments, the GUI enables agents toPage 16 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT review historical health data, confirm patient identity, and initiate follow-up actions, such as scheduling additional consultations or transmitting care instructions.
[0068] The agent devices 113 may also include software programming for secure communication with the provider server 111 and analytics system 101 via networks 105. This programming may implement encryption protocols, token-based authentication, and compliance features for handling sensitive health-related data. In certain implementations, the agent devices 113 support multi-modal interaction, allowing agents to receive audio prompts, view health metrics dashboards, and input clinical notes during a call. The agent devices 113 may further provide interoperability with external healthcare systems through APIs or integrated modules, enabling seamless updates to electronic health records (EUR) or other patient management platforms.
[0069] FIG. 2 illustrates dataflow amongst components of a system 200 for segmented voice processing. The system 200 includes hardware and software components for addressing technical challenges in speech signal modulation that are commonly imposed by communication carrier networks, such as those standards and protocols often implemented in cellular networks or other types of telephony-based communications networks. The system 200 includes hardware and software components of a segmented articulation engine 202, an analog-to-digital converter engine (ADC engine 204), digital speech data 206, a network carrier system 208, a modulation engine 210, a middleware communications engine 212, an analytics system 214, a speech integrity detection engine 216, a clean speech detection engine 217, a speech data analysis engine 218, health metrics data outputs 220, and an alert notification engine 222.
[0070] The segmented articulation engine 202 comprises hardware and software components configured to generate instructions for a caller to produce segmented voiced sounds for input audio signal data during a telephony session. The segmented articulation engine 202 may provide prompts requesting the caller to articulate vowel-based sounds (e.g., “ah,” “eh”) in short bursts separated by pauses, thereby mitigating signal modulation artifacts such as automatic gain control (AGC) commonly imposed by carrier networks. In some embodiments, the segmented articulation engine 202 operates in conjunction with the provider system 110 or analytics system 214 (e.g., analytics system 101) to deliver real-time guidance to the caller device 114. The segmented articulation engine 202 comprises instructions to the user / patient to articulate a pitch-Page 17 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT based sound (e.g. “ah,” “eh,” ... etc.) that exhibit a fundamental frequency (fD) and harmonics (fD*k, k is harmonic index) that is segmented - 2-3 seconds of “ah” interleaved by approximately 1 -second pauses. This process allows bypassing of carriers’ AGC feature that is oftentimes automatically applied to the speech signal.
[0071] The ADC engine 204 converts the analog speech signal captured by the caller device 114 into a digital representation suitable for downstream processing. For instance, the ADC engine 204 converts the segmented analog speech signal to a discrete digital sample version, and outputs the digitized audio samples as digital speech data 206. The ADC engine 204 includes hardware and software for generating one or more segments of speech signals by converting the analog speech signals into the discrete digital samples, as the digital speech data 206. The ADC engine 204 receives the continuous or segmented analog voice input signals, which include the periodic analog voice sounds, and performs various operations for quantizing the input analog signals into the digital format of the digital speech data 206. For instance, upon receiving the analog audio input, the ADC engine 204 performs sampling at a preconfigured sampling rate, converting the continuous analog audio signal into a series of discrete numerical values and outputting the digital audio signal or digital audio samples in the form of the digital speech data 206 that represent the analog audio samples. The ADC engine 204 may transform continuous analog speech signals into the discrete digital samples. The conversion operations of the ADC engine 204 enables components the system 200 to perform digital processing and analysis operations on voice signals in cellular communications. The ADC engine 204 may implement sampling routines at a predefined rate (e.g., 8 kHz) and apply quantization techniques to preserve the integrity of the segmented audio signal. The ADC engine 204 be an executable component of the caller device 114 or infrastructure of the network carrier 208.
[0072] The ADC engine 204 operates by taking the segmented voice input signals of the audio signal, which includes periodic voiced sounds such as “ah” interspersed with brief pauses, and quantizing these voice signal and / or audio signals into a digital format of the digital speech data 206. In some cases, the operations beneficially preserve the integrity of the original voice speech patterns while allowing for efficient transmission of the digitized audio signals, as the digital audio signals of the digital speech data 206, and for subsequent computational analysis. The digital audio samples of the digital speech data 206 may retain the essential frequency andPage 18 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT amplitude characteristics of the original analog voice signal, such that the subsequent components of the system 200 may accurately analyze and interpret the digital speech data 206.
[0073] After the ADC engine 204 digitizes the input analog speech signal and outputs the digital speech samples, the digital speech samples become part of the digital speech data 206. This digital speech data 206 is then transmitted through hardware and software components of a network carrier 208 for further processing and analysis at the various downstream components of the system 200. The digital speech data 206 represents the encoded audio stream generated by the ADC engine 204. This data includes segmented voiced audio segments and associated metadata, such as timestamps and segmentation markers. The digital speech data 206 is transmitted through the network carrier system 208 and serves as the primary input for subsequent analytics operations. In some implementations, the digital speech data 206 may be packetized for transport over IPbased networks or compressed using standardized codecs.
[0074] The system 200 may include hardware and software components of one or more network carriers 208 (e g., telephony carrier networks 130) that provides telephony services to the end-user’s calling device (e.g., caller device 114). The network carrier system 208 comprises hardware and software of a telephony infrastructure responsible for routing the digital speech data 206 between the caller device 114 and downstream systems. The carrier system 208 may include originating and terminating carrier components, switches, and signaling protocols for managing call sessions. In some embodiments, the carrier system 208 applies modulation features such as AGC, echo cancellation, and noise suppression, among others.
[0075] The network carrier 208 may include hardware and software components of a carrier modulation engine 210 for performing various modulation operations, such as an Automatic Gain Control (AGC) engine. The network carrier system 208 comprises telephony infrastructure responsible for routing the digital speech data 206 between the caller device 114 and downstream systems. The carrier system 208 may include originating and terminating carrier components, switches, and signaling protocols for managing call sessions. In some embodiments, the carrier system 208 applies modulation features such as AGC, echo cancellation, and noise suppression, which the segmented articulation protocol is designed to circumvent.Page 19 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT
[0076] The AGC engine of the carrier modulation engine 210 may regulate the amplitude of voice signals to maintain consistent audio levels during telephony services. The AGC engine operates by automatically adjusting the gain of the audio signal, particularly during sustained articulation of certain voiced sounds, such as the voiced-sound “ah.” The modulation operations of the AGC engine of the carrier modulation engine 210 may be activated automatically by sustained articulation of such voiced-sounds.
[0077] In some circumstances, it may be desirable to circumvent the AGC engine and modulation functions in order to more accurately capture the characteristics of the voice audio signal that are used for generating various health metrics outputs 220 (e.g., heart rate) of the caller. In such cases, the ADC engine 204, segmented articulation engine 202, or other components of the system 200 may beneficially perform the speech segmentation techniques that are capable of bypassing the AGC engine of the network carrier 208, thereby preserving the integrity of the original voice patterns in the digital speech data 206, which components of the system 200 reference to generate the health metrics outputs 220 for the caller, such as the heart rate of the caller.
[0078] The carrier modulation engine 210 may perform various additional or alternative modulation operations or other operations for handling the digital speech data 206. For instance, the ADC engine 204, modulation engine 210, or other components of the system 200 may execute data compression-and-decompression operations, data encoding-and-decoding operations, and audio sampling operations, among others. As an example, the voice audio signals in the digital speech data 206 in telephony communications may be compressed and sampled at 8kHz (fs = 8kHz).
[0079] The middleware communications engine 212 operates as an intermediary layer for secure and reliable transmission of segmented audio data to the analytics system 214. The engine 212 may implement session management, encryption, and error correction protocols to maintain audio fidelity under variable network conditions. In some embodiments, the middleware engine 212 validates segmentation markers and metadata before forwarding the audio stream for analysis. The network carrier 208 or other components of the system 200 may forward or transmit the digital speech data 206, from the carrier modulation engine 210 to components of the middleware communications system 212 (e.g. Twillio®). Generally, the middleware cloud communicationPage 20 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT service 212 may provide automated middleware software services for connecting callers and called devices or receivers with various cloud-based operations or services, similar to that of a human operator.
[0080] The middleware communications engine 212 may forward or transmit the segmented speech recording of the digital speech data 206, from the network carrier 208 to the analytics system 214. Upon receiving segmented audio streams from the network carrier 208, the middleware communications engine 212 may perform a series of preparatory operations, such as validating the integrity of the audio data, confirming the presence of expected segmentation markers, and applying standardized encoding or packetization routines to ensure compatibility with downstream analytics processes. The middleware communications engine 212 may also implement session management protocols, encryption mechanisms, and error correction techniques to maintain audio fidelity and protect sensitive health-related information during transit. Once these operations are complete, the middleware communications engine 212 forwards the segmented speech recording, now in a format suitable for analysis, to the analytics system 214, where advanced signal processing and health metric computations are performed.
[0081] The analytics system 214 includes hardware and software components for generating the health metrics of the caller based on the digital speech data 206 (such as an analytics server 102 of FIG. 1). The analytics system 214 analyze and process the speech data 206 to render the health metrics outputs 220, such as the heart rate and other information
[0082] The analytics system 214 comprises computing resources for processing segmented audio signals and generating health metrics. The analytics system 214 may include one or more servers executing algorithms for speech integrity detection, signal analysis, and physiological parameter computation. In certain implementations, the analytics system 214 interfaces with provider systems 110 to transmit computed metrics for display at end-user devices, such as caller devices 114 or agent devices 113. In the system 200, the analytics system 214 includes a speech integrity detection module 216.
[0083] The speech integrity detection engine 216 evaluates the quality of the received audio signal to confirm compliance with segmentation requirements. The speech integrity detection engine 216 may perform voiced-sound or unvoiced- sound detection or classification,Page 21 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT energy threshold checks, and segment counting to confirm whether the analytics system 214 received sufficient data for accurate analysis. If integrity criteria are not met, the 216 / / may trigger an alert or re-prompt containing a request for additional audio samples.
[0084] The software programming of the analytics system 214 analyzes the speech data 206 provided by the user (via a middleware communication node of the middleware communications components 212) and generates various health metrics outputs 220, such as the user’s heart rate and heart-rate variability values, among others. The functions of analytics system 214 may include, for example, voiced / unvoiced sound detection to determine if the speech data includes a strong fundamental frequency typically present in the vocalization of “ah”, computing the number of “ah” segments, and determining energy levels of the segments.
[0085] The analytics system 214 may execute software operations for clean speech detection 217. The clean speech detection engine 217 detects “clean speech” or otherwise determines whether the audio signal is free from artifacts, such as excessive noise or incomplete segmentation. The clean speech detection engine 217 may apply statistical and spectral analysis techniques to validate the presence of strong fundamental frequency components and harmonic structures indicative of vowel sounds. In these operations, the analytics system 214 determines whether the speech data 206 provided by the user is “clean” based on outputs of the speech integrity detection engine 216.
[0086] For speech data that meets the criteria for cleanliness and robustness, the analytics system 214 proceeds with advanced speech analysis and physiological signal processing routines, as in the functions of the speech data analysis engine 218. Non-limiting examples of these operations of the speech data analysis engine 218 include: extracting time-domain and frequencydomain features from the audio signal, identifying periodic modulations associated with cardiac activity, and performing spectral analysis to isolate harmonics indicative of heart rate. The speech data analysis engine 218 or other aspects of the analytics system 214 executes transformation functions, such as STFT or DFT, and data evaluation operations, such as autocorrelation and harmonic spacing evaluation, to detect pulse-like amplitude modulations within the voiced segments. After the speech data analysis engine 218 / determines a fundamental frequency and related periodic perturbations, the speech data analysis engine 218 computes health metrics outputsPage 22 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT220, such as the user’s heart rate and heart rate variability (HRV) by mapping the detected frequency components of the fundamental frequency to physiological metrics.
[0087] The speech data analysis engine 218 may generate additional types of health metrics data outputs 220, such as the consistency and stability of the extracted features of the acoustic data across multiple segments to enhance the reliability of the computed health metrics. The health metrics outputs 220 generated by the speech data analysis engine 218 or other aspects of the analytics system 214 may include a comprehensive set of physiological and speech-derived indicators. Non-limiting examples of the health metrics data outputs 220 include: patient vitals such as heart rate, heart rate variability (HRV), lung capacity (estimated from sustained vowel articulation and respiratory patterns), oxygen saturation (inferred from spectral characteristics and amplitude envelope), and electrocardiogram (ECG) trace proxies derived from periodicity in the speech signal. The analytics system 214 may also detect and quantify slurred speech as a marker for neurological assessment, compute blood pressure estimates using multi-modal signal fusion, and determine mean arterial pressure through advanced modeling of speech and acoustic features. Additional health metrics data outputs 220 may include respiratory rate, phonation quality, and alert flags for anomalies such as arrhythmias or abnormal variability.
[0088] For digital speech data 206 that is determined to be non-robust or not clean, such as user-provided signals affected by excessive noise, incomplete segmentation, or insufficient voiced energy, the analytics system 214 initiates feedback and reacquisition operations. In such cases, the system generates an alert via the alert notification engine 222, which may include specific guidance for improving signal quality (e.g., instructions to repeat the articulation, increase loudness, or reduce background noise). The alert notification engine 222 of the analytics system 214 transmits alert notifications to a provider system 110 or directly to a caller device 114 or agent device 113. The analytics system 214 may log the reason for reacquisition and track subsequent attempts into a database 104 to indicate instances of failed minimum signal integrity criteria. The process then repeats, prompting the user to provide a new speech sample until the captured audio meets the required standards for accurate health metric computation.
[0089] In some cases, the speech data analysis engine 218 performs advanced signal processing operations to extract physiological parameters from the segmented audio signal for the health metrics data outputs 220. These operations may include time-domain periodicity analysis,Page 23 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT frequency-domain harmonic spacing evaluation, and detection of modulation patterns associated with cardiac activity. The speech data analysis engine 218 generates health metrics data outputs 220, such as heart rate and heart rate variability for downstream use.
[0090] The health metrics data outputs 220 represent the computed physiological parameters derived from the speech data analysis engine 218. These health metrics data outputs 220 may include heart rate, heart rate variability, and related indicators of physiological parameters or autonomic function of the caller. The health metrics data outputs 220 are transmitted to provider systems 110 for integration with clinical workflows and may be displayed on agent devices 113 during an active session.
[0091] FIG. 3 shows dataflow amongst components of a system 300 for analyzing voice signals in audio data to generate health metrics for speaker-users, according to embodiments. The system 300 depicts two processes 301, 303 for generating health metrics data: a segmented articulation process 301 and an unsegmented articulation process 303, each configured to produce audio signals. The system 300 may implement a segmented articulation process 301 and / or a continuous vowel articulation process 303 for analyzing voice signals in audio signals to generate the health metrics, including physiological data, for speaker-users. The figure further illustrates the resulting effects of carrier-imposed modulation and the corresponding impact on health metric accuracy.
[0092] The segmented articulation process 301 includes user-directed prompts or instructions for a caller to vocally articulate vowel-based, voiced sounds in short bursts separated by pauses. This process is designed to mitigate automatic gain control (AGC) effects commonly applied by telephony carrier systems. As shown in FIG. 3, the segmented articulation process 301 generates a sequence of voiced segments interleaved with silence intervals as a segmented articulation audio sequence, represented by a first audio sequence 305. In the first audio sequence 305, each voiced sound segment may have a duration in the range of approximately two to three seconds, followed by a pause of approximately one second, resulting in a total articulation period of about twelve seconds.
[0093] The first audio sequence 305, representing the segmented articulation audio sequence generated by the segmented articulation process 301, illustrates an example pattern ofPage 24 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT alternating voiced segments and pauses. This segmented articulation process 301 prevents sustained amplitude conditions that typically trigger AGC, thereby preserving the original signal characteristics. The segmented articulation process 301 may output the first audio sequence 305 and a resulting audio representation, represented as first speech data 307. The segmented articulation process 301 avoids the AGC functions, and as a result, the first speech data 307 demonstrates the absence of AGC-induced attenuation, maintaining consistent amplitude across all voiced segments. This audio integrity enables accurate extraction of physiological parameters from the audio signal.
[0094] When employing the segmented articulation process 301, a non-continuous articulation of “ah,” as a series of shorter “ahs” (e.g. 2-3 seconds) interleaved by short pauses (e.g. 1 second) results. When the first audio sequence 305 is subjected to modulation operations of the communication carrier, the AGC is not triggered, preserving the full spectrum of the “ah” sound, which is then divided into multiple “ah” voiced segments for the first speech data 307. These voiced segments of the first speech data 307 are analyzed by the first speech data analysis operations 309 either collectively or individually by the first speech data analysis operations 309 to generate accurate heart rate and other related measurements.
[0095] The first speech data analysis operations 309 represent functions executed by an analytics server using the first speech data 307, as generated by the segmented articulation process 301. These first speech data operations 309 may include, for example, time-domain periodicity analysis, frequency-domain harmonic spacing evaluation, and detection of modulation patterns associated with cardiac activity. Because the segmented articulation process 301 avoids AGC distortion when producing the first speech data 307, the first analysis operations 309 generates correct health metrics, such as heart rate and heart rate variability, as physiological parameters for the speaker-user.
[0096] The unsegmented or continuous articulation process 303 includes user-directed prompts or instructions for a caller to vocally articulate a continuous vowel voiced-sound for an extended duration, such as eight seconds, to produce a continuous articulation sequence, represented by the second audio sequence 311. While the continuous articulation process 303 may appear suitable for capturing periodic modulations, the continuous articulation process 303 oftenPage 25 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT triggers AGC mechanisms within carrier networks, resulting in amplitude compression and distortion of the signal, or other modulation operations that impact the audio signal integrity.
[0097] The continuous articulation sequence in the second audio sequence 311 illustrates a sustained vowel sound without pauses. This sustained amplitude condition activates AGC within the carrier system, as depicted by the second audio data 313 produced using the AGC modulation functions. The AGC effect significantly attenuates the signal after a few seconds, reducing the amplitude of the voiced sound and masking subtle variations in the second audio data 313 required for accurate physiological analysis.
[0098] The second speech data analysis operations 315 represents operations performed on audio signals, using the second audio data 313, generated by the unsegmented articulation process 303. Due to AGC-induced distortion, the analysis may yield erroneous health metrics, including inaccurate heart rate and variability values. This limitation underscores the technical advantage of the segmented articulation process 301 over conventional continuous articulation methods.
[0099] In the continuous articulation process 303, the user provides a vowel sound (e.g., “ah” sound) that the user holds for some duration of time (e.g., approximately 8 seconds). This prior approach presents a few challenges in obtaining a sufficient amount of voice signal data for generating the health metrics. For instance, some user may find it difficult to hold the vowel sound tone for a required duration. Moreover, telephony communications software layers and carrier services often apply signal modulations operations on the call audio signal (e.g., AGC, noise suppression, echo cancellation) that can reduce the quality of the voice signal data or otherwise complicate the process of heart rate computation. The signal modulations operations often distort the original speech signal in a manner that makes heart rate computation challenging.
[0100] For example, as shown in the continuous articulation process 303, after an “ah” is articulated for the duration of time (e.g., eight seconds), an AGC engine is triggered for AGC operations, whereby the signal amplitude is attenuated after two or three seconds to generate the second audio data 313. While technically it is possible to disable certain modulation features at the carrier-side, such as the AGC engine, control of these features are typically unavailable at the user device or at a middleware node of a middleware node system (e.g., Twillio®). The second speechPage 26 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT data analysis operations 315 modulated speech signal in the second audio data 313 is then processed and analyzed by the second speech data analysis operations 315 for generating health metric data, such as physiological data, heart rate, and other information, which often results in erroneous information due to signal modulation operations, such as AGC modulation.
[0101] FIG. 4 shows operations of a method 400 for performing speech data analysis according to a segmented articulation process. The method 400 is described as performed by a server that provides healthcare data using the speech analysis on speech audio data for the user (e.g., analytics server 102). Embodiments, however, are not so limited. The method 400 depicts a dataflow amongst system components that implement whole-signal and per-segment analysis techniques to extract heart rate and related health metrics from voice signals transmitted over carrier network. The process 400 analyzes segmented speech signals as a whole and as individual segments to accurately compute health metrics (e.g., heart rate values) and other information about a speaker-user.
[0102] In operation 401, the server obtains speech audio data for a call that originated at a user device (e.g., caller device 114) of the user. In some implementations, the server receives the speech audio data from a network carrier or other telephony system for handling calls. In some implementations, the server receives the speech audio data through a middleware server of a middleware system (e.g., Twillio®). In some cases, the input audio stream may include segmented vowel sounds articulated by the caller in accordance with a segmentation protocol designed to mitigate automatic gain control (AGC) effects.
[0103] Optionally, in operation 403, the server executes a segmentation operation on the speech audio data obtained in the operation 401. The software functions of the server performs segmentation operations to generate the segmented audio data 405 by generating or inserting breaks or pauses into the speech audio data. In some cases, however, the speech audio signal is already segmented when obtained by the server, such that the server obtains the segmented audio data 405 in the operation 401. For example, the user may be prompted to provide a tone or vowel sound for a certain amount of time and then stop over multiple intervals. In this way, the server obtains the segmented audio data 405 and may not need to execute the segmentation operation (as in the current operation 403).Page 27 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT
[0104] In operation 407, the server executes analysis operations for the entire audio signal using the collection of the segmented speech signal 405 as a whole, which includes performing a transformation function, such as short-Fourier Transform (STFT), and executing a spectral analysis on the transformed segmented audio data in the transform domain (e.g., frequency domain). This transformation may be configured to provide a maximum frequency resolution by employing a window size that is the size of the entire signal. The “pauses” can optionally be zeroed to render more robust heart rate related outputs.
[0105] In operation 409, for the whole segmented speech signal, each frequency bin is analyzed to determine heart rate pulse patterns. The server analyzes the entire segmented speech signal in the frequency domain to identify heart-rate pulse patterns that persist across time. In some embodiments, the server computes the STFT over the collection or concatenated segments of the segmented audio data 405 comprising the user’s vowel segments, and, to enhance frequency resolution while preserving the segmentation effect, optionally zeroes or de-weights the inter-segment pauses before analysis. For each frequency bin (i.e., each time-evolving “bin track”), the server applies hard / soft gating to attenuate inter-harmonic regions and retain energy located at or near harmonics of the vocal fundamental (fo). The server then forms a pulse-spectrum by applying a discrete Fourier transform (DFT) to the amplitude trajectory of each bin track and tests whether pulse-like periodicity consistent with a cardiac rhythm is present. Bins failing voiced-speech criteria (e.g., insufficient harmonicity) or minimum signal-to-noise thresholds are discarded. The server may further employ peak-refinement (e.g., parabolic or spline interpolation) and windowing selections to improve frequency estimates while mitigating spectral leakage.
[0106] The server may determine a whole harmonic distance 411 for the entire audio signal using the collection of the segmented audio data 405. Using the whole-signal pulse-spectral results, the server determines inter-peak spacings across candidate spectral maxima to estimate a harmonic distance associated with the cardiac fundamental. As an example, the server generates a histogram of frequency differences between adjacent candidate peaks aggregated over multiple harmonics and bin tracks. The server evaluates even-harmonic and odd-harmonic hypotheses: if peaks follow f_k ~ k fo (even / complete harmonic series), spacings are aggregated as difference Af = fb; if peaks follow f_k ~ (2k+l) fo (odd-only series), spacings are aggregated as difference Af ~ 2 fo. Robust statistics (e.g., median-of-pairwise differences, trimmed means) may be used to suppress outliersPage 28 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT introduced by network compression or transient artifacts. The resulting value represents the whole harmonic distance 411 for the full audio record in Hertz.
[0107] The server determines an average (e.g., mean, median mode), frequent, or otherwise dominant spacing for the whole audio signal (e.g., whole max mode 413) and a whole heart rate 414 for the entire audio signal using the collection of the segmented audio data 405. The server identifies the whole max mode 413 from the whole-signal spacing histogram as a candidate (fo), mapping to beats per minute (e.g., BPM = 60 fo), and enforces plausibility constraints (e.g., allowable adult ranges such as 35-220 BPM; expected continuity with prior accepted measurements; limits on frame-to-frame acceleration). The server may compute confidence from peak prominence, harmonic-series agreement, cross-bin consensus, and the proportion of voiced energy retained after gating. If multiple modes have comparable strength, the server can apply tie-breakers favoring the mode with higher harmonic consistency, greater cross-segment agreement, or lower uncertainty. The result of this path is a whole-signal heart-rate estimate with an associated confidence score.
[0108] In operation 415, the server executes analysis operations for each particular segment of the segmented audio data 405. For each segment, the server analyzes each frequency bin to determine segment-level heart rate pulse patterns. In parallel, the server performs per-segment analysis for each segment of the segmented audio data 405. For each segment, the server computes an STFT on the segment-bounded waveform (without zero-padding external pauses), executes the segment-level gating and bin-track construction, and generates a per-segment pulse-spectrum. This per-segment path increases robustness to localized disturbances (e.g., brief user motion, packet loss) and mitigates the impact of carrier-side automatic gain control (AGC) by exploiting the designed pauses that suppress AGC triggering.
[0109] For each segment of the set of segmented audio data 405, the server determines a segment harmonic distance 417 using the segment audio signal of the particular segment. For each segment, the server determines the segment harmonic distance 417 using the segment’s pulse-spectrum. Because segments are shorter (e.g., two to three seconds), the server may increase effective frequency resolution by interpolation (e.g., quadratic peak fitting, spline interpolation, or poly-fit) and by aggregating spacing estimates across multiple validated harmonics. In some cases,Page 29 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT each segment’s spacing estimate is associated with quality indicators, such as segment SNR, number of validated bin tracks, and voiced-speech confidence.
[0110] For each segment of the set of segmented audio data 405, determines a segment max mode 419 and a segment heart rate 420 using the segment audio signal of the particular segment. For each segment, the server identifies the segment max mode from the segment-level spacing histogram and converts it to a segment BPM estimate. Segment outliers can be rejected using robust estimators (e.g., Hampel filters, RANSAC-style consensus) subject to physiologic and continuity constraints. The server stores, for each segment, the segment BPM estimate, segment BPM confidence, the selected harmonic hypothesis (even or odd), and diagnostic metrics (e g., peak sharpness, inter-harmonic agreement) for downstream fusion functions.
[0111] In operation 421, the server combines candidate outputs from the whole-signal and per-segment analyses using a constraint-based majority-vote and mode-analysis algorithm. The server filters candidates by likelihood thresholds representing plausibility (e.g., human physiologic limits, maximum allowable change from the prior accepted rate, minimum voiced-segment count) and weighted by evidence quality (e g., SNR, harmonic-series conformity, peak prominence, number of agreeing segments, and whether pauses were successfully identified / zeroed). The server may prefer or select the whole-signal estimate having a higher frequency resolution and satisfies one or more constraints; otherwise, the server may generate or select a robust aggregate (e.g., weighted median of top- - modes). If disagreement exceeds a divergence threshold (e.g., >15 BPM), the server may downgrade confidence and flag the sample for reacquisition. This fusion resolves discrepancies introduced by variable network conditions, compression artifacts, or residual AGC effects.
[0112] At operation 423, the server generates and outputs the computed physiological parameter, such as heart rate, based on aggregated results from both analysis paths. This dual-path approach enhances accuracy and reliability under variable network conditions. The server generates and outputs a physiological parameter, such as BPM, based on the fused result and, in some cases, auxiliary metadata (e.g., confidence score, quality flags indicating “insufficient voiced segments,” “low SNR,” “possible carrier AGC”). In some embodiments, the server also computes an HRV proxy from spectral-peak dispersion or inter-peak variability metrics collected during the foregoing analyses and includes this proxy as an ancillary output. The dual-path design of thePage 30 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT process 400 improves accuracy and reliability across heterogeneous devices, codecs, and carrier conditions.
[0113] FIGS. 5A-5B illustrate an idealized pulse formation and corresponding spectral analysis as performed by a computing device (e.g., analytics server 102, provider server 111), according to embodiments. FIG. 5A depicts a pulse sequence 500a generated by the analytics server 102. The pulse sequence 500a represents a periodic amplitude modulation observed on energy concentrated along harmonic tracks of a voiced vowel segment (e.g., “ah”) of a speakeruser, which may be received from user device 114. FIG. 5B depicts a pulse-spectrum 500b generated by the analytics server 102. The analytics server 102 executes a discrete Fourier transform (“DFT”) applied to the pulse sequence 500a (as depicted in FIG. 5A) to produce the pulse-spectrum 500b (depicted in FIG. 5B) having peaks or structural characteristics according to a harmonic relationship with an underlying cardiac rate. As used herein, a bin track denotes the time-evolving magnitude of a single STFT frequency bin; a spacing histogram denotes a histogram of adjacent spectral -peak frequency differences computed across candidate harmonics; comb-coherence denotes an agreement score between observed spectral peaks and an ideal harmonic comb under a selected hypothesis (complete-series or odd-dominant).
[0114] In operation, a user device 114 (e.g., a telephone handset, softphone, or mobile application) initiates a call or session with a provider server 111 over one or more communications networks 130 and / or a cloud middleware service 133. The provider server 111 transmits a segmented-articulation prompt to the user device 114, instructing a speaker-user to articulate a sustained, voiced vowel (e.g., “ah”) in 2-3 second segments separated by ~1 -second pauses for a configured duration (e.g., 10-15 seconds total). During the session, the provider server 111 can replay a brief audio demonstration and / or render on-screen cues to assist the user in producing steady phonation at a comfortable loudness.
[0115] The provider server 111 captures the incoming audio stream, timestamps packet arrival, and forwards the audio (or a recording thereof) to an analytics server 102 via one or more networks 105. In some embodiments, the provider server 111 attaches session metadata (e.g., codec type, nominal sampling rate, call leg identifiers, channel energy statistics) that the analytics server 102 uses to normalize subsequent processing. The agent device 113 mayPage 31 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT concurrently display capture state (e.g., “segments 1-5 recorded,” “insufficient voiced energy,” “re-prompt needed”) to a provider-user.
[0116] Upon receipt, the analytics server 102 performs a codec-normalization pipeline that may include any subset of: (i) p-1 aw or A-law companding decode; (ii) channel de-emphasis or DC removal; (iii) resampling to a target analysis rate (e.g., 8 kHz or 16 kHz) using band-limited interpolation; (iv) channel-wise energy normalization to a configured RMS window; and (v) optional narrowband pass filtering (e.g., 80-3400 Hz for PSTN) to align spectra across heterogeneous call paths. The analytics server 102 logs the nominal sampling rate, effective bandwidth, and decode pathway as diagnostic metadata associated with the analyzed record.
[0117] In some embodiments, the analytics server 102 executes a lightweight voice-activity detector (VAD) and a voiced-speech detector to identify candidate frames for voiced segments and to exclude non-voiced regions from subsequent harmonic-track analysis. Frames failing minimum energy and voiced periodicity criteria are masked prior to STFT computation to improve downstream gating and bin-track stability.
[0118] To generate the pulse sequence 500a, as in FIG. 5A, the analytics server 102 computes a time-frequency representation (e.g., short-time Fourier transform) of the speech signal and applies gating thresholds to attenuate regions between harmonics while retaining energy on harmonic tracks associated with the voiced source. The analytics server 102 aggregates energy across one or more validated harmonic tracks to produce a univariate amplitude trajectory that emphasizes periodic “pulse-like” rises and falls resulting from the physiological modulation of the vocal signal by the heartbeat. To improve robustness, the analytics server 102 may normalize pertrack amplitudes, reject tracks failing voiced or SNR thresholds, and optionally “zero” or “de-weighf ’ pauses introduced by the segmented-articulation protocol. The resulting time-domain pulse sequence 500a is an idealized, evenly-spaced series of pulses.
[0119] The analytics server 102 executes the DFT function applied to the values of the pulse sequence 500a to generate the frequency-domain representation of a pulse structure 500b, shown in FIG. 5B. Under ideal conditions, evenly spaced pulses yield a comb-like spectrum with a fundamental frequency at fo and harmonics at integer multiples of the fundamental frequency. Based on the effective duty cycle of the pulse sequence 500a, the spectral structure 500b mayPage 32 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT accentuate or emphasize a complete harmonic series (approximately, f_k » k fo), or an odddominant series (approximately f_k = (2k+l) fo). The inter-peak spacing encodes fo in Hertz, which the analytics server 102 converts to heart rate in beats per minute (BPM) as BPM = 60 fo.
[0120] In some embodiments, the analytics server 102 determines a harmonic distance by measuring frequency differences between adjacent candidate peaks in the pulse-spectrum 500b and compiling these differences into a spacing histogram. For a complete series hypothesis, adjacent differences approximate the fundamental frequency fo; for an odd-dominant hypothesis, adjacent differences approximate two fundamental frequencies, which the analytics server 102 correspondingly scales. The analytics server 102 identifies a max mode of the spacing histogram and adopts the associated frequency as the candidate fundamental frequency fo. The analytics server 102 may perform tie-breaking selections between competing modes based on, for example, peak prominence, harmonic-series conformity, and agreement across multiple gated bin tracks or multiple segments.
[0121] To mitigate finite-window effects and codec-induced blur, the analytics server 102 may apply windowing, zero-padding, and peak-interpolation (e.g., parabolic interpolation, spline interpolation) to refine spectral-peak locations prior to spacing analysis. In some cases, the analytics server 102 combines spacing data across several harmonics to reduce variance and applies robust statistics (e.g., trimmed means, medians) to suppress spurious intervals caused by noise, compression artifacts, or transient microphone handling.
[0122] As an example, the analytics server 102 may implement STFT windows having the range of 256-2048 samples with 50-75% overlap (at 8-16 kHz analysis rates), and Hann or Kaiser windows (0—5—8). The analytics server 102 may evaluate N=3-6 harmonics for comb-coherence scoring. Spacing histograms may use bin widths of 0.01-0.05 Hz with parabolic or spline peak interpolation.
[0123] In some embodiments, the analytics server 102 enforces plausibility and quality constraints when interpreting the pulse-spectrum 500b. Constraints may include minimum voiced-segment counts, minimum peak-to-noise ratios, and physiological bounds on heart rate. The analytics server 102 may also score each hypothesis (e.g., complete harmonic series, odd-dominant harmonic series) using series-conformity metrics (e.g., how well multiples of thePage 33 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT candidate fundamental frequency fo predict observed peak positions), and select the hypothesis with the higher conformity score. Estimates that fail prescribed thresholds are flagged for reacquisition or are down-weighted in subsequent fusion operations.
[0124] The analytics server 102 may generate, alongside the fo / BPM estimate derived from the pulse-spectrum 500b, confidence indicators that reflect peak sharpness, spacing-histogram dominance, cross-harmonic consensus, and stability across time. Diagnostic metadata can include the number of peaks used, hypothesis selected (complete or odd-dominant), interpolation method applied, and any gating or zeroing operations performed on pauses. These indicators facilitate downstream fusion with other analysis paths and support auditability.
[0125] In some embodiments, the analytics server 102 computes a composite confidence score for a heart-rate estimate using a weighted combination of (i) comb -coherence (agreement to an ideal harmonic comb at the selected hypothesis); (ii) spacing-histogram dominance (mode prominence vs. next-best mode); (iii) peak-prominence statistics across the first N harmonics; (iv) cross-domain agreement (frequency-domain fo vs. time-domain IPI); and (v) segment consensus (dispersion of per-segment BPM). Weights may be configured based on deployment context (e.g., narrowband PSTN vs. wideband VoIP).
[0126] The reporting policy may include thresholds that: (a) report BPM with a confidence value when the composite score > a reporting threshold; (b) report BPM with a quality flag (e.g., “elevated variability,” “possible AGC”) when the composite score is between a warning and reporting threshold; or (c) issue a no-estimate with a quality-of-capture notice when the score is below a minimum confidence threshold or when spacing histograms are multi-modal without a clear max mode. The provider server 111 may then prompt user device 114 to reacquire audio using guidance rules described herein.
[0127] The analysis results generated by the analytics server 102 and illustrated inFIGS. 5A-5B serves as a canonical, “clean-case” model that the analytics server 102 adapts to real -world conditions in other figures. For example, where segmented articulation yields multiple short, high-integrity segments, per-segment pulse-DFT analyses, analogous to FIGS. 5A-5B, can be executed and combined via constraint-based fusion. Conversely, when a full concatenated signal affords higher frequency resolution, a whole-signal analysis consistent with FIGS. 5A-5BPage 34 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT may be preferred. Both implementations fall within the scope of the disclosed embodiments and may contribute candidates to a final fused heart-rate output.
[0128] FIGS. 6A-6B illustrate the effect of moderate timing variability (e.g., approximately ten percent deviation from ideal pulse centers) on pulse formation and the corresponding spectral analysis. In operation, a user device 114 (e.g., a telephone handset or mobile device) captures voiced vowel segments and transmits audio signals or audio data over a communications channel to a provider server 111, which may be transmitted via a cloud middleware system 133 and / or telephony networks 130. The provider server 111 forwards the audio stream data or recording to an analytics server 102 for analysis via one or more networks 105. FIG. 6A depicts a pulse sequence 600a, derived or generated by the analytics server 102 from the audio data, in which individual pulses are jittered in time relative to ideal, evenly spaced occurrences. FIG. 6B shows a frequency-domain representation of pulse-spectrum 600b (e.g., a DFT computed by the analytics server 102 on the pulse sequence 600a) in which spectral lines associated with the cardiac fundamental and corresponding harmonics exhibit broadened peaks, reduced prominence, and increased inter-peak dispersion, as compared to the idealized case ofFIGS. 5A-5B
[0129] The analytics server 102 derives the univariate pulse sequence 600a by computing a time-frequency representation of the received speech (e.g., STFT) and aggregating energy across validated harmonic tracks, with per-track normalization and voiced / SNR screening. Pauses introduced by the segmented-articulation protocol may be zeroed or down-weighted by the analytics server 102 before aggregation to prevent pause boundaries from biasing center estimates. The resulting time series, generated by the analytics server 102, exhibits bounded temporal jitter about an expected interpulse interval (IPI), visually demonstrating non-uniform spacings constrained around a common mean IPI.
[0130] The analytics server 102 generates the pulse-spectrum 600b by executing a DFT applied to the jittered pulse sequence 600a to obtain the pulse-spectrum 600b. Relative to the ideal case, moderate variability may, for example, broaden spectral lines, reduce peak-to-floor contrast, and introduce dispersion around expected harmonic locations. While the inter-peak spacing remains approximately equal to the cardiac fundamental frequency (fo) when measured in aggregate, the localization of each peak degrades, and low-level side structures may appear due toPage 35 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT time-domain irregularity. These transformations are computed at the analytics server 102 and may be previewed to personnel via an agent device 113 operating a clinical or provider dashboard.
[0131] In some embodiments, the analytics server 102 quantifies the spectral effects of the pulse-spectrum 600b using one or more variability metrics, such as: peak width (e.g., full width at half maximum) at candidate harmonic locations; comb coherence, defined as agreement between observed peak frequencies and an ideal comb at multiples of a candidate fundamental frequency fo; spacing dispersion, computed from the variance or robust spread of adjacent inter-peak differences; and peak-prominence ratio, computed as a peak amplitude normalized by a local noise-floor estimate. The analytics server 102 combines these measures into a frequency-domain variability index that increases as timing irregularity grows. The agent device 113 may display the index and supporting diagnostics to a provider-user of a provider system 110.
[0132] The analytics server 102 includes programming for AGC or suppression detectiong that analyzes amplitude envelopes and spectral floor trajectories to identify carrier-induced gain changes. In one implementation, the server flags a candidate AGC event when (i) a sustained voiced segment exhibits a step-like envelope attenuation exceeding a configured dB threshold within a short time window, and (ii) a correlated rise in the noise floor (or suppression skirt) is observed across inter-harmonic bands. When flagged, the server down-weights affected frames in the bin-track aggregation and may exclude them from spacing histograms used for harmonic-distance estimation.
[0133] To mitigate AGC effects while preserving usable data, the analytics server 102 can apply segment-aware weighting that favors earlier, pre-AGC intervals within each segment and increases fusion reliance on unaffected segments. The server records AGC detections as quality flags (e g., “possible AGC,” onset time, affected duration) that are surfaced to a provider-user via an agent device 113 and included in the session audit record.
[0134] In some embodiments, the analytics server 102 evaluates the spectral pattern under both a complete-series hypothesis (f_k « k fo) and an odd-dominant hypothesis (f_k ~ (2k+l) fo). For each hypothesis, the analytics server 102 forms a histogram of inter-peak differences, identifies a dominant mode (max mode), and rescales differences as appropriate (e.g., halving differences under odd-dominant assumptions) to obtain a candidate fundamental frequency fo. The analyticsPage 36 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT server 102 then scores each hypothesis using series-conformity metrics (e.g., predicted-vs-observed alignment error across the first N harmonics) and selects the hypothesis with the higher score for downstream heart-rate estimation and variability reporting.
[0135] In some embodiments, the analytics server 102 cross-validates frequency-domain findings by estimating pulse centers directly from the aggregated time series using, for example, autocorrelation peaks, matched filtering, or peak-picking with refractory constraints. The analytics server 102 computes summary statistics of inter-pulse intervals (e.g., mean, median, and robust spread) and compares these statistics against frequency-domain spacing dispersion to confirm consistency. Where supported, the agent device 113 may present cross-check outcomes and confidence indicators. Large disagreement between domains may reduce confidence and trigger reacquisition logic at the provider server 111 to re-prompt the user device 114 for a new sample.
[0136] In some embodiments, the analytics server 102 maps the pulse-spectrum 600b metrics to qualitative classes (e.g., “low” vs. “elevated” variability) using configurable thresholds. For approximately ten percent jitter, the analytics server 102 typically observes modest peak broadening and spacing dispersion within acceptable physiologic limits; in such cases, the analytics server 102 may extract a reliable heart-rate estimate from the dominant spacing mode, and further determine or generate an elevated-variability flag and a reduced confidence score relative to the ideal case. The analytics server 102 may generate one or more alert messages containing status indicators for transmission to the provider system 110 and / or various devices, such as caller device 114 or agent devices 113.
[0137] The analytics server 102 evaluates segment-level integrity using criteria that may include: a minimum number of validated segments (e.g., >4); a minimum voiced-duration per segment (e.g., >1.5 second voiced frames after VAD or voiced-sound screening); a minimum segment SNR (e.g., >12 dB within a harmonic band); and a harmonic-track count (e.g., >three validated tracks). Sessions failing one or more criteria trigger a re-prompt operation.
[0138] In response, the provider server 111 may cause the agent device 113 and / or user device 114 to present capture prompts (e.g., “increase loudness,” “reduce background noise,” “maintain steady ‘ah’,” “keep phone 10-15 cm from mouth”). Optionally, the analytics server 102Page 37 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT records reasons for re-prompt and the subsequent outcome in a database 104, which may be included as part of the audit log in the database 104.
[0139] In some embodiments, the analytics server 102 computes variability metrics per segment and for the whole record. The analytics server 102 weights segment-level contributions by segment quality indicators (e.g., voiced confidence, SNR, number of validated bin tracks) and aggregates results into a session-level variability index. Segments exhibiting transient artifacts (e.g., packet loss originating along the path to the provider server 111 or brief mis-phonation at the user device 114) may be down- weighted or excluded using robust outlier rules to prevent spurious inflation of variability. Aggregated metrics and exclusions may be transmitted, displayed, and reviewed at various devices, such as caller device 114 or agent devices 113.
[0140] In some embodiments, the presence of moderate variability, as in the pulse sequence 600a and pulse-spectrum 600b (of FIGS. 6A-6B), reduces the prominence and sharpness of spectral peaks but does not preclude accurate determination of the fundamental frequency fo. The analytics server 102 executes fusion logic (e g., constraint-based majority vote across whole-signal and per-segment paths) that incorporates the variability index as a weighting factor, such that candidates supported by higher comb coherence and lower spacing dispersion are favored. The analytics server 102 outputs the fused heart-rate estimate, an associated confidence value, and diagnostic metadata indicating the measured level of variability and the hypothesis selected for spectral interpretation; these outputs may be returned through the provider server 111 and presented to a clinician or other end-user at various devices, such as the caller device 114 or the agent devices 113.
[0141] In some embodiments, the analytics server 102 generates a heart-rate variability (HRV) proxy value derived from frequency-domain dispersion metrics (e.g., normalized peak-width and spacing-dispersion in the 0.7-3.0 Hz cardiac band). The HRV proxy may be categorized into qualitative bands (e.g., low, moderate, high) for human-readable informational use and / or confidence weighting.
[0142] FIGS. 7A-7B illustrate the effect of elevated timing variability (e.g., approximately fifteen percent deviation from ideal pulse centers) on pulse formation and the corresponding spectral analysis, according to embodiments. In operation, a user device 114 (e.g., a telephonePage 38 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT handset or mobile device) captures voiced vowel segments and transmits audio over one or more communications channels to a provider server 111 (optionally via cloud middleware system 133 and / or telephony networks 130). The provider server 111 forwards the audio stream or recording to an analytics server 102 for processing. FIG. 7A depicts a pulse sequence 700a, generated by the analytics server 102 from the received audio, in which pulses exhibit more pronounced temporal jitter than the moderate-variability case. FIG. 7B shows a frequency-domain representation, as a pulse-spectrum 700b, computed by the analytics server 102 (e.g., a DFT of the sequence 700a), where spectral lines associated with the cardiac fundamental frequency and harmonics may be displayed with broadened lines, reduced prominence, and / or increased dispersion relative to the idealized case of FIGS. 5A-5B and the moderate-variability case of FIGS. 6A-6B
[0143] The analytics server 102 computes a time-frequency representation of the speech (e g., STFT), applies gating thresholds to attenuate inter-harmonic regions, and retains energy along validated harmonic tracks associated with the voiced source. The analytics server 102 aggregates the gated, per-track amplitudes — after per-track normalization and voiced / SNR screening — into a univariate amplitude trajectory. Prior to aggregation, pauses introduced by the segmented-articulation protocol may be “zeroed” or down-weighted to prevent pause boundaries from biasing center estimates. The resulting time-domain pulse sequence 700a exhibits elevated temporal jitter around the expected inter-pulse interval (IPI), producing visibly non-uniform spacings that depart from the tighter clustering of the values of the pulse sequence 600a depicted in FIG. 6A
[0144] The analytics server 102 executes a DFT on the pulse sequence 700a to generate the pulse-spectrum 700b. With elevated timing variability, spectral lines broaden, peak-to-floor contrast decreases, and side structures (e.g., skirted shoulders) may emerge near expected harmonic locations. While the aggregate inter-peak spacing in 700b continues to reflect the cardiac fundamental (fo), individual peaks are less sharply localized, and local spacing estimates may deviate from the nominal fo due to increased time-domain irregularity. These transformations are computed at the analytics server 102 and may be transmitted to and displayed at user interface of end-user computing device, such as a user device 114 or agent device 113.Page 39 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT
[0145] The analytics server 102 quantifies features of the pulse-spectrum 700b using one or more metrics, including: peak width (e.g., full width at half maximum) at candidate harmonic locations; comb-coherence, defined as agreement between observed peak frequencies and an ideal comb at multiples of a candidate fundamental frequency fo; spacing dispersion, computed from the variance or robust spread of adjacent inter-peak differences; and prominence ratio, computed as a peak amplitude normalized by a local noise-floor estimate. The analytics server 102 combines these measures into a frequency-domain variability index that increases with timing irregularity, and may transmit the index and supporting diagnostics to the provider system 110 for display at a computing device, such as a user device 114 or agent device 113.
[0146] To account for sampling-rate drift and clock skew across heterogeneous endpoints, the analytics server 102 may estimate an effective sampling rate by aligning known spectral features (e.g., power-line or comfort-noise bands when present) or by maximizing short-term autocorrelation stability, then resample the signal to the corrected rate prior to STFT. When dual-channel media are received, the analytics server 102 can select the channel exhibiting higher voiced energy and stability or beamform a monophonic signal prior to analysis.
[0147] The analytics server 102 evaluates the spectral pattern under both a complete-series hypothesis (f_k « k fo) and an odd-dominant hypothesis (f_k ~ (2k+l) fo). For each hypothesis, the analytics server 102 constructs a histogram of inter-peak differences, identifies a dominant mode (max mode), and rescales differences where appropriate (e.g., halving differences under odd-dominant assumptions) to obtain a candidate fo. The analytics server 102 then scores each hypothesis using series-conformity metrics (e.g., predicted-vs-observed alignment error across the first N harmonics) and selects the hypothesis with the higher score for downstream heart-rate estimation and variability reporting.
[0148] In parallel, the analytics server 102 estimates pulse centers directly from the aggregated time series using, for example, autocorrelation peaks, matched filtering, or peak-picking with refractory constraints. The analytics server 102 computes summary statistics of the inter-pulse intervals (e.g., mean, median, robust spread) and compares these statistics against frequency-domain spacing dispersion to confirm consistency. Cross-check outcomes and confidence indicators may be presented at a user interface of an end-user device, such as the caller device 114 or agent device 113. In response to the analytics server 102 or end-user identifying aPage 40 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT threshold difference or disagreement between values of the domains may reduce confidence and trigger the analytics server 102 reacquisition logic at the provider server 111 to re-prompt the user device 114 for a new sample.
[0149] The analytics server 102 maps pulse-spectrum 700b metrics to qualitative classes (e g., “elevated variability”) using configurable thresholds. At approximately fifteen percent jitter, the analytics server 102 typically observes clear peak broadening and increased spacing dispersion beyond the moderate case, yet still within ranges that support reliable fo extraction from the dominant spacing mode. In such cases, the analytics server 102 may generate an elevated-variability flag and a corresponding reduction in confidence relative to FIGS. 6A-6B, while continuing to report a heart-rate estimate.
[0150] The analytics server 102 computes variability metrics per segment and for the whole record, weighting segment-level contributions by quality indicators (e.g., voiced confidence, SNR, number of validated bin tracks). Segments exhibiting transient artifacts — such as packet loss along the path to the provider server 111 or brief “mis-phonation” at the user device 114, which may be down-weighted or excluded via robust outlier rules to prevent spurious inflation of variability. Aggregated metrics and exclusions may be returned through the provider server 111 for review on a computing device, such as a user device 114 or agent device 113.
[0151] The presence of elevated variability, as in the pulse sequence 700a and pulse-spectrum 700b of FIGS. 7A-7B, reduces spectral peak sharpness and prominence but does not preclude accurate determination of the fundamental frequency fo. The analytics server 102 executes fusion logic (e.g., constraint-based majority vote across whole-signal and per-segment paths) that incorporates the variability index as a weighting factor, favoring candidates with higher comb-coherence and lower spacing dispersion. The analytics server 102 outputs a fused heart-rate estimate, an associated confidence value, and diagnostic metadata indicating the measured variability level and the selected harmonic-series hypothesis; these outputs of the analytics server 102 may be delivered via the provider server 111 and presented to end-users on the user device 114 and / or the agent device 113.
[0152] FIGS. 8A-8B illustrate the effect of substantial timing variability (e.g., approximately twenty percent deviation from ideal pulse centers) on pulse formation and thePage 41 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT corresponding spectral analysis, according to embodiments. In operation, a user device 114 (e.g., a telephone handset or mobile device) captures voiced vowel segments and transmits audio signals over one or more communications channels to a provider server 111, which may be transmitted via a cloud middleware system 133 and / or telephony networks 130. The provider server 111 forwards the audio stream data or recording to an analytics server 102 for analysis via one or more networks 105. FIG. 8A depicts a pulse sequence 800a, derived or generated by the analytics server 102 from the audio data, in which inter-pulse intervals vary widely around a nominal period. FIG. 8B shows a frequency domain representation of a pulse-spectrum 800b (e.g., a DFT computed by the analytics server 102 on the pulse sequence 800a) in which the comb like structure associated with the cardiac fundamental and corresponding harmonics is markedly degraded relative to the idealized case of FIGS. 5A-5B and the elevated-variability case of FIGS. 7A-7B.
[0153] The analytics server 102 derives the univariate pulse sequence 800a by computing a time-frequency representation of the received speech (e.g., STFT) and aggregating energy across validated harmonic tracks, with per track normalization and voiced / SNR screening. Pauses introduced by the segmented articulation protocol may be zeroed or down weighted by the analytics server 102 before aggregation to prevent pause boundaries from biasing center estimates. The resulting time series, generated by the analytics server 102, exhibits substantial temporal jitter about an expected inter-pulse interval (IPI), with visible short-long and long-short runs and larger departures from uniform spacing than those observed in FIG. 7A.
[0154] The analytics server 102 generates the pulse-spectrum 800b by executing a DFT applied to the substantially jittered pulse sequence 800a. Under substantial variability, spectral lines broaden (e.g., increased full width at half maximum), peak to floor contrast decreases, and side structures (e.g., skirted shoulders or partial peak splitting) may emerge near expected harmonic locations. While the aggregate inter-peak spacing remains informative of the cardiac fundamental frequency (fo) when measured in aggregate, local peak localization degrades and local spacing estimates may deviate from the nominal fo due to increased time domain irregularity. These transformations are computed at the analytics server 102 and may be previewed to personnel via an agent device 113 operating a clinical or provider dashboard.
[0155] In some embodiments, the analytics server 102 quantifies the spectral effects of the pulse-spectrum 800b using one or more variability metrics, such as: peak width (e.g., full width atPage 42 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT half maximum) at candidate harmonic locations; comb coherence, defined as agreement between observed peak frequencies and an ideal comb at multiples of a candidate fundamental frequency fo; spacing dispersion, computed from the variance or robust spread of adjacent inter peak differences; and peak prominence ratio, computed as a peak amplitude normalized by a local noise floor estimate. The analytics server 102 combines these measures into a frequency domain variability index that increases as timing irregularity grows. The agent device 113 may display the index and supporting diagnostics to a provider user of a provider system 110.
[0156] In some embodiments, the analytics server 102 evaluates the spectral pattern under both a complete series hypothesis (f_k » k fo) and an odd dominant hypothesis (f_k ~ (2k+l) • fo). For each hypothesis, the analytics server 102 forms a histogram of inter peak differences, identifies a dominant mode (max mode), and rescales differences as appropriate (e.g., halving differences under odd dominant assumptions) to obtain a candidate fundamental frequency fo. The analytics server 102 then scores each hypothesis using series conformity metrics (e.g., predicted vs observed alignment error across the first N harmonics) and selects the hypothesis with the higher score for downstream heart rate estimation and variability reporting.
[0157] In some embodiments, the analytics server 102 cross validates frequency domain findings by estimating pulse centers directly from the aggregated time series using, for example, autocorrelation peaks, matched filtering, or peak picking with refractory constraints. The analytics server 102 computes summary statistics of inter pulse intervals (e.g., mean, median, and robust spread) and compares these statistics against frequency domain spacing dispersion to confirm consistency. Where supported, the agent device 113 may present cross check outcomes and confidence indicators. Large disagreement between domains may reduce confidence and trigger reacquisition logic at the provider server 111 to re prompt the user device 114 for a new sample.
[0158] In some embodiments, the analytics server 102 maps pulse-spectrum 800b metrics to qualitative classes (e.g., “high variability”) using configurable thresholds. For approximately twenty percent jitter, the analytics server 102 typically observes pronounced peak broadening and increased spacing dispersion beyond the elevated case, yet a reliable fo may still be extracted from a dominant spacing mode when present. In such cases, the analytics server 102 may generate a high variability flag and a corresponding reduction in confidence relative to FIGS. 7A-7B, while continuing to report a heart rate estimate. If dominance is not achieved (e.g., a multi modal spacingPage 43 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT histogram with no clear max mode), the analytics server 102 may decline to report BPM and instead generate a quality of capture notice.
[0159] In some embodiments, the analytics server 102 computes variability metrics per segment and for the whole record. The analytics server 102 weights segment level contributions by segment quality indicators (e.g., voiced confidence, SNR, number of validated bin tracks) and aggregates results into a session level variability index. Segments exhibiting transient artifacts (e.g., packet loss originating along the path to the provider server 111 or brief mis phonation at the user device 114) may be down weighted or excluded using robust outlier rules to prevent spurious inflation of variability. Aggregated metrics and exclusions may be transmitted, displayed, and reviewed at devices such as the user device 114 or the agent device 113.
[0160] In some embodiments, the presence of substantial variability, as in the pulse sequence 800a and pulse-spectrum 800b of FIGS. 8A-8B, reduces the prominence and sharpness of spectral peaks and may complicate local spacing estimation, but does not necessarily preclude accurate determination of the fundamental frequency fo. The analytics server 102 executes fusion logic (e.g., constraint-based majority vote across whole signal and per segment paths) that incorporates the variability index as a weighting factor, such that candidates supported by higher comb coherence and lower spacing dispersion are favored. The analytics server 102 outputs the fused heart rate estimate, an associated confidence value, and diagnostic metadata indicating the measured level of variability and the hypothesis selected for spectral interpretation; these outputs may be returned through the provider server 111 and presented to a clinician or other end user at various devices, such as the user device 114 or the agent device 113.
[0161] In some embodiments, when confidence remains below a reporting threshold after fusion, the analytics server 102 and provider server 111 may coordinate a guidance loop. The provider server 111 generates an instruction to end-user device, such as a caller device 114 or an agent device 113 (e.g., provider dashboard), to prompt the speaker-user to provide a re-recording with steadier phonation, increased loudness, or reduced ambient noise. The agent device 113 can trigger the user device 114 to present an instructional prompt or play a demonstration sample before reacquisition. Guidance outcomes and capture quality may be logged to support longitudinal review and auditability associated with FIGS. 8A-8B.Page 44 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT
[0162] FIG. 9 is a flowchart of a computer-implemented method 900 for detecting physiological parameters (e.g., heart rate) or other health metrics data outputs using audio signals, according to embodiments. The operations of the method 900 are performed by a computer, though any computing device or collection of computing devices may be used (e.g., analytics server 102 of analytics system 101). Embodiments may include additional or alternative operations, or omit operations, as those of the method 900 and still fall within the scope of this disclosure.
[0163] At operation 910, a computer (e.g., analytics server 102) obtains an input audio signal comprising a plurality of voiced sound segments for a speaker and one or more acoustic pauses. The input audio signal may be captured during a telephony session, remote consultation, or via a dedicated application on a user device 114. The system prompts the speaker to produce specific vowel sounds (such as “ah”) in segmented intervals of voiced sound segments, typically ranging from 2-3 seconds per segment, interleaved with brief pauses of approximately 1 second. This segmentation operation may beneficially minimize AGC and other carrier-induced modulations, thereby preserving the integrity of the physiological information encoded in the speech signal. The computer may also receive metadata associated with the audio signal, such as sampling rate, codec type, and session identifiers, which are used to normalize and preprocess the signal prior to analysis. In some cases, the computer may also receive metadata associated with the audio signal, such as sampling rate, codec type, and session identifiers, which are used for normalization and preprocessing.
[0164] The computer may determines whether the input audio signal satisfies one or more signal integrity criteria, such as minimum energy thresholds, sufficient number of voiced segments, and presence of harmonic structure. If the criteria are not met, the computer may generate an alert and prompt the user to provide a new sample.
[0165] At operation 920, the computer determines a fundamental frequency for one or more voiced sound segments of the input audio signal. This determination may be based on timedomain periodicity, such as identifying regular intervals between peaks in an amplitude envelope, or frequency-domain harmonic-spacing, which involves determining analyzing the spectral structure or characteristics of the input audio signal using, for example, STFT or DFT. The computer applies gating and filtering functions to isolate energy concentrated along harmonic tracks associated with the voiced sound source, and may use voice-activity detection (V D)Page 45 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT software functions to exclude non-voiced or unvoiced sound regions. For each segment, the system computes the fundamental frequency (fo) and validates the presence of the fundamental frequency using criteria, such as minimum energy thresholds, harmonic consistency, and signal-to-noise ratio (SNR). The computer may aggregate or algorithmically combine (e.g., average) the resulting segment fundamental frequency estimates across segments to enhance robustness and accuracy.
[0166] At operation 930, the computer identifies a periodic modulation corresponding to cardiac pulses of the speaker for the fundamental frequency. The periodic modulation may be based upon a plurality of perturbations of one or more acoustic characteristics of the fundamental frequency. The computer identifies the periodic modulation corresponding to cardiac pulses of the speaker for the fundamental frequency based upon detecting a series of perturbations, such as amplitude modulations or subtle frequency shifts, in the fundamental frequency or in harmonics of the fundamental frequency.
[0167] The computer identifies the periodic modulation by analyzing the time-varying characteristics or trajectory of the voiced sound segments, focusing on the amplitude envelope and spectral content of the audio signal. The computer may execute various signal processing techniques, such STFT or DFT, or other spectral analysis methods, to identify or uncover periodic fluctuations that are temporally aligned with expected cardiac rhythms. These modulations may manifest as slight, recurring changes in loudness (amplitude) or as small, cyclical shifts in the pitch (frequency) of the voiced sound, which may be influenced by the mechanical and hemodynamic effects of the cardiac cycle.
[0168] The computer distinguishes between physiological modulations and artifacts by leveraging both time-domain and frequency-domain analyses. In addition to amplitude and frequency perturbations, the computer may examine harmonics of the fundamental frequency, as cardiac-driven modulations can induce coherent shifts across multiple harmonic components. The computer may execute filtering and gating functions to suppress noise, exclude unvoiced or artifact-laden regions, and isolate a true physiological signal. The computer may validate the presence, periodicity, and regularity of these perturbations against expected heart rate ranges and dynamics, such that only genuine cardiac modulations are considered for generating physiological parameters (as in operation 940, below). The computer may aggregate findings across multiplePage 46 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT segments and cross-validate results using independent detection functions, such as autocorrelation or matched filtering, to enhance robustness and accuracy.
[0169] The computer analyzes a time-varying trajectory of the amplitude envelope and executes a spectral analysis function to detect pulse-like periodicity that aligns with expected cardiac rhythms. The computer may distinguish physiological modulations from artifacts introduced by the communication channel or user movement. The computer may also crossvalidate findings using both time-domain and frequency-domain approaches, ensuring that detected modulations are consistent with known patterns of cardiac activity.
[0170] At operation 940, the computer generates one or more physiological parameters of the speaker for the input audio signal based on the periodic modulation of the fundamental frequency of the input audio signal. These physiological parameters may include heart rate (beats per minute, BPM), heart rate variability (HRV), and additional metrics (e.g., respiratory rate, phonation quality, proxy values indicating lung capacity and oxygen saturation). The computer applies plausibility constraints and confidence scoring to confirm that reported values fall within physiologically expected or reasonable threshold ranges. If the computed physiological parameters meet quality and integrity criteria thresholds, the computer outputs the physiological parameters for integration with downstream clinical workflows, user interfaces or dashboards, or electronic health records databases. If the signal quality is insufficient, the computer may trigger a feedback loop, alert function, or re-prompting operations, prompting the user to provide a new sample or adjust recording conditions to improve data quality. In some implementations, if the generated physiological parameters meet or exceed alert thresholds (e.g., abnormal heart rate, insufficient signal quality), the computer generates an alert notification for display at a user interface, and may initiate a feedback loop to reacquire audio data.
[0171] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the describedPage 47 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0172] Embodiments implemented in computer software may be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment or machine-executable instructions may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, attributes, or memory contents. Information, arguments, attributes, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0173] The actual software code or specialized control hardware used to implement these systems and methods is not limiting of the invention. Thus, the operation and behavior of the systems and methods were described without reference to the specific software code being understood that software and control hardware can be designed to implement the systems and methods based on the description herein.
[0174] When implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in a processor-executable software module which may reside on a computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable media includes both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. A non-transitory processor-readable storage media may be any available media that may be accessed by a computer. By way of example, and not limitation, such non-transitory processor- readable media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that may be used to store desired program code in the form of instructions or data structures and that may be accessed by a computer or processor. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-Ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers.Page 48 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENTCombinations of the above should also be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of codes and / or instructions on a non-transitory processor-readable medium and / or computer- readable medium, which may be incorporated into a computer program product.
[0175] The preceding description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the spirit or scope of the invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.
[0176] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.Page 49 of 544899-1684-8759
Claims
VITAL01-PCT / 142238-0103 PATENTCLAIMSWhat is claimed is:
1. A computer-implemented method for detecting physiological parameters of using audio signals, the method comprising: obtaining, by a computer, an input audio signal comprising a plurality of voiced sound segments for a speaker and one or more acoustic pauses; determining, by the computer, a fundamental frequency for one or more voiced sound segments of the input audio signal based on at least one of a time-domain periodicity of the one or more voiced sound segments or a frequency-domain harmonic-spacing of the one or more voiced sound segments; identifying, by the computer, a periodic modulation corresponding to cardiac pulses of the speaker for the fundamental frequency, the periodic modulation based upon a plurality of perturbations of one or more acoustic characteristics of the fundamental frequency; and generating, by the computer, one or more physiological parameters of the speaker for the input audio signal based on the periodic modulation of the fundamental frequency of the input audio signal.
2. The method of claim 1, further comprising determining, by the computer, one or more signal integrity criteria for the input audio signal, and wherein the computer determines the one or more physiological parameters in response to the computer determining that the input audio signal satisfies the one or more signal integrity criteria.
3. The method of claim 2, wherein determining the one or more signal integrity criteria includes detecting, by the computer, each voiced sound segment of the plurality of voiced sound segments in the input audio signal.
4. The method of claim 2, wherein determining the one or more signal integrity criteria includes determining, by the computer, that each voiced sound segment satisfies a minimum energy threshold;Page 50 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT5. The method of claim 2, wherein determining the one or more signal integrity criteria includes determining, by the computer, that each voiced sound segment corresponds with the fundamental frequency.
6. The method of claim 2, wherein determining the one or more signal integrity criteria includes determining, by the computer, that at least one voiced sound segment corresponds with a harmonics characteristic for a vowel sound.
7. The method of claim 1, wherein the one or more acoustic characteristics include at least of an amplitude, a frequency, an amplitude envelope, or an instantaneous frequency; and wherein the computer identifies the periodic modulation of the fundamental frequency corresponding to the cardiac pulses based upon at least one of the amplitude envelope or the instantaneous frequency of the fundamental frequency.
8. The method of claim 1, further comprising generating, by the computer, a time-frequency representation of the input audio signal using one or more transform functions and the input audio signal.
9. The method of claim 1, wherein the one or more physiological parameters includes at least one of a heart rate or heart rate variability.
10. The method of claim 1, further comprising generating, by the computer, an alert for display at a user interface in response to determining that a value of a physiological parameter satisfies an alert threshold corresponding to the physiological parameter.
11. A system for detecting physiological parameters of using audio signals, the system comprising: a computer comprising at least one processor, configured to: obtain an input audio signal comprising a plurality of voiced sound segments for a speaker and one or more acoustic pauses; determine a fundamental frequency for one or more voiced sound segments of the input audio signal based on at least one of a time-domain periodicity of the one or more voiced sound segments or a frequency-domain harmonic-spacing of the one or more voiced sound segments;Page 51 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT identify a periodic modulation corresponding to cardiac pulses of the speaker for the fundamental frequency, the periodic modulation based upon a plurality of perturbations of one or more acoustic characteristics of the fundamental frequency; and generate one or more physiological parameters of the speaker for the input audio signal based on the periodic modulation of the fundamental frequency of the input audio signal.
12. The system of claim 11, wherein the computer is further configured to determine one or more signal integrity criteria for the input audio signal, and wherein the computer determines the one or more physiological parameters in response to the determining that the input audio signal satisfies the one or more signal integrity criteria.
13. The system of claim 12, wherein when determining the one or more signal integrity criteria the computer is further configured to detect each voiced sound segment of the plurality of voiced sound segments in the input audio signal.
14. The system of claim 12, wherein when determining the one or more signal integrity criteria the computer is further configured to determine that each voiced sound segment satisfies a minimum energy threshold;15. The system of claim 12, wherein when determining the one or more signal integrity criteria the computer is further configured to determine that each voiced sound segment corresponds with the fundamental frequency.
16. The system of claim 12, wherein when determining the one or more signal integrity criteria the computer is further configured to determine that at least one voiced sound segment corresponds with a harmonics characteristic for a vowel sound.
17. The system of claim 11, wherein the one or more acoustic characteristics include at least of an amplitude, a frequency, an amplitude envelope, or an instantaneous frequency; and wherein the computer identifies the periodic modulation of the fundamental frequency corresponding to the cardiac pulses based upon at least one of the amplitude envelope or the instantaneous frequency of the fundamental frequency.Page 52 of 544899-1684-8759VITAL01-PCT / 142238-0103 PATENT18. The system of claim 11, the computer is further configured to generate a time-frequency representation of the input audio signal using one or more transform functions and the input audio signal.
19. The system of claim 11, wherein the one or more physiological parameters includes at least one of a heart rate or heart rate variability.
20. The system of claim 11, the computer is further configured to generate an alert for display at a user interface in response to the computer determining that a value of a physiological parameter satisfies an alert threshold corresponding to the physiological parameter.Page 53 of 544899-1684-8759