Deriving insights on health by analyzing audio data generated by a digital stethoscope

By using an electronic stethoscope system and machine learning algorithms, the problem of frequency attenuation when amplifying sounds inside a living body by traditional stethoscopes has been solved. This has enabled efficient amplification and accurate identification of internal sounds, providing near real-time respiratory rate calculation and visualization analysis, and improving diagnostic accuracy.

CN115884709BActive Publication Date: 2026-03-27HEROIC FAITH MEDICAL SCI CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-14
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Traditional acoustic stethoscopes suffer from frequency attenuation and changes in ear sensitivity when amplifying sounds within a living body, making it difficult to accurately diagnose illnesses, especially since low-frequency sounds are difficult to hear.

Method used

An electronic stethoscope system is used to convert sound waves into electrical signals through a microphone and amplify them. At the same time, machine learning algorithms and deep learning models are used to analyze audio data, identify and filter out external noise, and enhance the recognition and diagnosis of internal sounds.

Benefits of technology

It improves the amplification of in vivo sounds, enhances the accuracy of identifying internal sounds such as inhalation and exhalation, and provides near real-time respiratory rate calculation and visualization analysis to help medical professionals make more accurate diagnoses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115884709B_ABST
    Figure CN115884709B_ABST
Patent Text Reader

Abstract

Presented herein are computer programs and associated computer-implemented techniques for deriving insights into a patient's health by analyzing audio data generated by an electronic stethoscope system. A diagnostic platform can be responsible for examining this audio data generated by an electronic stethoscope system in order to gain insights into a patient's health. The diagnostic platform can employ heuristics, algorithms, or models that rely on machine learning or artificial intelligence to perform auscultation in a manner that is demonstrably superior to traditional methods that rely on visual analysis by a healthcare professional.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various implementation schemes involve computer programs and related computer implementation techniques for deriving insights into a patient's health by analyzing audio data generated by an electronic stethoscope system. Background Technology

[0002] Historically, acoustic stethoscopes have been used to listen to internal sounds originating from living organisms. This process (called "auscultation") is typically performed to examine biological systems, from which the performance of those systems can be inferred. A typical acoustic stethoscope comprises a single chest piece with resonators designed to fit snugly against the body, and the pair of hollow tubes connected to a stethoscope. When the resonators capture sound waves, these waves are guided through the hollow tubes to the stethoscope.

[0003] However, acoustic stethoscopes have several drawbacks. For example, they attenuate sound proportionally to the frequency of the sound source. Therefore, the sound transmitted to the stethoscope is often very weak, making accurate diagnosis difficult. In fact, due to variations in ear sensitivity, some sounds (e.g., sounds below 50 Hz) may be completely inaudible.

[0004] Some companies have begun developing electronic stethoscopes (also known as "chest sound transmitters") to address the shortcomings of acoustic stethoscopes. Electronic stethoscopes improve upon acoustic stethoscopes by electronically amplifying sound. For example, they can process faint sounds originating from within a living body by amplifying them. To achieve this, the electronic stethoscope converts sound waves detected by a microphone located in the chest piece into electrical signals, which are then amplified for optimal hearing. Attached Figure Description

[0005] Figure 1A Top perspective view including the input unit of the electronic stethoscope system.

[0006] Figures 1B to 1C include Figure 1A The bottom perspective view of the input unit shown.

[0007] Figure 2 A cross-sectional side view including the input unit of the electronic stethoscope system.

[0008] Figure 3 This illustrates how one or more input units can be connected to a hub unit to form an electronic stethoscope system.

[0009] Figure 4 This is a high-level block diagram illustrating exemplary components of the input unit and hub unit of an electronic stethoscope system.

[0010] Figure 5The network environment, including the diagnostic platform, is shown.

[0011] Figure 6 An example of a computing device capable of implementing a diagnostic platform is shown, which is designed to generate outputs that help detect, diagnose, and monitor changes in a patient's health.

[0012] Figure 7 This includes an example of a workflow diagram that shows how the audio data obtained by the diagnostic platform can first be processed by a back-end service, and then the analysis of the audio data can be presented on the interface for review.

[0013] Figure 8 Includes a high-level diagram of the computational pipeline that can be used by the diagnostic platform during the detection phase.

[0014] Figure 9 The architecture of several baseline models that can be adopted by the diagnostic platform is shown.

[0015] Figure 10 This demonstrates how the diagnostic platform can classify individual segments of audio data.

[0016] Figure 11A Includes high-level illustrations of the algorithmic methods used to estimate respiratory rate.

[0017] Figure 11B The diagram illustrates how a diagnostic platform can calculate respiratory rate using a sliding window within a predetermined length of record that is updated at a predetermined frequency.

[0018] Figure 12 An alternative method of auscultation is shown, in which the diagnostic platform uses autocorrelation to identify inspiration and expiration.

[0019] Figure 13A This demonstrates how spectrograms generated by a diagnostic platform can be presented in near real-time to facilitate diagnostic determination by healthcare professionals.

[0020] Figure 13B This demonstrates how digital elements (also known as "graphical elements") can be overlaid on a spectrogram to provide additional insights into a patient's health.

[0021] Figure 14 A flowchart is depicted for the process of detecting respiratory events by analyzing audio data and then calculating the respiratory rate based on these respiratory events.

[0022] Figure 15 A flowchart is depicted for the process of calculating respiratory rate based on analysis of audio data containing sounds emitted by a patient's lungs.

[0023] Figure 16This demonstrates how healthcare professionals can use the electronic stethoscope system and diagnostic platform described in this article to listen to auscultation sounds and detect respiratory events in near real-time.

[0024] Figure 17 This is a block diagram illustrating an example of a processing system in which at least some of the operations described herein can be implemented.

[0025] Schemes are illustrated in the accompanying drawings by way of example and not limitation. While the drawings depict various embodiments for illustrative purposes, those skilled in the art will recognize that alternative embodiments may be employed without departing from the principles of the present technology. Thus, although specific embodiments are shown in the drawings, the present technology is readily adaptable to various modifications. Detailed Implementation

[0026] An electronic stethoscope system can be designed to simultaneously monitor sounds originating from both inside and outside the living body being examined, as discussed further below. As discussed further below, the electronic stethoscope system may include one or more input units connected to a hub unit. Each input unit may have a conical resonator designed to direct sound waves toward at least one microphone configured to generate audio data indicating internal sounds originating from within the living body. These microphones may be referred to as “stethoscope microphones”. Additionally, each input unit may include at least one microphone configured to generate audio data indicating external sounds originating from outside the living body. These microphones may be referred to as “ambient microphones” or “surrounding environment microphones”.

[0027] For illustrative purposes, an “ambient microphone” can be described as capable of generating audio data that indicates “ambient sound.” However, these “ambient sounds” generally include a combination of external sounds generated by three different sources: (1) sounds originating from the surrounding environment; (2) sounds leaking through the input unit; and (3) sounds penetrating the living body being examined. Examples of external sounds include sounds directly originating from the input unit (e.g., scratches from fingers or chest) and low-frequency ambient noise penetrating the input unit.

[0028] Recording internal and external sounds separately has several advantages. Notably, internal sounds can be amplified electronically, while external sounds can be suppressed, attenuated, or filtered electronically. Therefore, electronic stethoscope systems can process faint sounds originating from a living organism being examined by manipulating audio data that indicates both internal and external sounds. However, manipulation can lead to undesirable digital artifacts, making internal sounds more difficult to interpret. For example, these artifacts can make it harder to identify patterns in the audio data generated by the stethoscope microphone that indicate inhalation or exhalation values.

[0029] This article describes computer programs and associated computer implementation techniques used to derive insights into patient health by analyzing audio data generated by electronic stethoscope systems. A diagnostic platform (also known as a "diagnostic program" or "diagnostic application") is responsible for examining the audio data generated by the electronic stethoscope system to gain insights into patient health. As discussed further below, diagnostic platforms may employ heuristics, algorithms, or models that rely on machine learning (ML) or artificial intelligence (A1) to perform auscultation in a manner significantly superior to traditional methods that rely on visual analysis by healthcare professionals.

[0030] Suppose, for example, that a diagnostic platform's task is to infer a patient's respiratory rate based on audio data generated by an electronic stethoscope system connected to the patient. The term "respiratory rate" refers to the number of breaths per minute. For a resting adult, a respiratory rate between 12 and 25 breaths per minute is considered normal. A respiratory rate less than 12 breaths per minute or greater than 25 breaths per minute at rest is considered abnormal. Because respiratory rate is a key indicator of health, it is sometimes referred to as a "vital sign." Therefore, knowing the respiratory rate can be crucial for providing appropriate medical care to a patient.

[0031] As mentioned above, an electronic stethoscope system can generate audio data indicating internal sounds and audio data indicating external sounds. The former may be referred to as "first audio data" or "internal audio data," and the latter may be referred to as "second audio data" or "external audio data." In some embodiments, the diagnostic platform utilizes the second audio data to improve the first audio data. Thus, the diagnostic platform can examine the second audio data to determine how the first audio data (if any) should be manipulated to mitigate the influence of external sounds. If, for example, the diagnostic platform detects external sounds by analyzing the second audio data, the diagnostic platform can obtain the first audio data and subsequently apply a filter to the first audio data to remove external sounds without distorting the internal sounds of interest (e.g., those corresponding to inhalation, exhalation, etc.). The diagnostic platform can generate the filter based on its analysis of the second audio data, or the diagnostic platform can identify the filter based on its analysis of the second audio data (e.g., from multiple filters).

[0032] The diagnostic platform can then apply a computer-implemented model (or simply "model") to the first audio data. This model can be designed and trained to recognize different phases of breathing, namely inhalation and exhalation. This model can be a deep learning model based on one or more artificial neural networks (or simply "neural networks"). Neural networks are a framework of ML algorithms that work together to process complex inputs. Inspired by the biological neural networks that make up the human and animal brains, neural networks can "learn" to perform tasks by considering examples without being programmed with task-specific rules. For example, a neural network can learn to recognize inhalation and exhalation by examining a series of audio data labeled as "inhalation," "exhalation," or "no inhalation or exhalation." These series of audio data can be called "training data." This training method allows the neural network to automatically learn from the series of audio data the presence of features indicating inhalation and exhalation. For example, the neural network can begin to understand patterns in the audio data indicating inhalation and patterns in the audio data indicating exhalation.

[0033] While the model's outputs can be useful, they can be difficult to interpret. For example, a diagnostic platform might generate a visual representation (also called a "visualization") of the initial audio data for review by healthcare professionals responsible for monitoring, examining, or diagnosing patients. The platform can visually highlight inhalation or exhalation to improve the interpretability of this visualization. For instance, the platform might overlay numerical elements (also called "graphical elements") onto the visualization to indicate a patient's breath. As discussed further below, the numerical elements can extend from the start of inhalation, as determined by the model, to the end of exhalation, as determined by the model. By interpreting the model's outputs in a more understandable way, the diagnostic platform can build trust with healthcare professionals who rely on these outputs.

[0034] Furthermore, diagnostic platforms can leverage the outputs of this model to gain insights into patient health. For example, suppose the model is applied to audio data associated with a patient to identify inhalations and exhalations as discussed above. In this scenario, the diagnostic platform can use these outputs to generate metrics. One example of such a metric is the respiratory rate. As mentioned above, the term "respiratory rate" refers to the number of breaths per minute. However, counting the actual number of breaths in the last minute is simply impractical. If a patient suffers from hypoxia (even with a limited timeframe), significant irreversible damage can occur. Therefore, it is crucial for healthcare professionals to have a consistent insight into respiratory rates that are updated near real-time. To achieve this, diagnostic platforms can continuously calculate the respiratory rate using a sliding window algorithm. At a higher level, this algorithm defines a window containing a portion of the audio data and then continuously slides the window to cover different portions of the audio data. As discussed further below, the algorithm can adjust the window so that its boundaries always contain a predetermined number of inhalations or exhalations. For example, the algorithm can be programmed such that the window must include three inhalations within its boundaries, where the first inhalation represents the "beginning" of the window and the third inhalation represents the "end" of the window. The diagnostic platform can then calculate the respiratory rate based on the intervals between the inhalations included in the window.

[0035] For illustrative purposes, implementation schemes may be described in the context of instructions executable by a computing device. However, aspects of this technology may be implemented via hardware, firmware, software, or any combination thereof. For example, a model may be applied to audio data generated by an electronic stethoscope system to identify outputs representing a patient's inspiration and expiration. An algorithm may then be applied to the outputs produced by the model to generate a measure indicating the patient's respiratory rate.

[0036] Implementation schemes may be described with reference to specific computing devices, networks, and healthcare professionals and organizations. However, those skilled in the art will recognize that the features of these implementation schemes can be similarly applied to other computing devices, networks, and healthcare professionals and organizations. For example, while implementation schemes may be described in the context of deep neural networks, the models employed by a diagnostic platform may be based on alternative deep learning architectures, such as deep belief networks, recurrent neural networks, and convolutional neural networks.

[0037] the term

[0038] The following are brief definitions of the terms, abbreviations and phrases used throughout this application.

[0039] The terms “connection,” “coupled,” and any variations thereof are intended to include any direct or indirect connection or coupling between two or more elements. This connection or coupling can be physical, logical, or a combination thereof. For example, objects can be electrically or communicatively connected to each other even without sharing a physical connection.

[0040] The term "module" can be used broadly to refer to a component implemented via software, firmware, hardware, or any combination thereof. Generally, a module is a functional component that generates one or more outputs based on one or more inputs. A computer program may include one or more modules. Therefore, a computer program may include multiple modules responsible for performing different tasks or a single module responsible for performing all tasks.

[0041] Overview of Electronic Stethoscope Systems

[0042] Figure 1A The image shows a top perspective view of the input unit 100 of an electronic stethoscope system. For convenience, the input unit 100 may be referred to as a "stethoscope patch," even though the input unit may only include a subgroup of components required for auscultation. The input unit 100 may also be referred to as a "chest piece," as it is typically attached to the chest of the body. However, those skilled in the art will recognize that the input unit 100 may also be attached to other parts of the body (e.g., the neck, abdomen, or back).

[0043] As further described below, the input unit 100 can collect sound waves representing the bioactivity within the body being examined, convert these sound waves into electrical signals, and then digitize these electrical signals (e.g., for easier transmission, to ensure higher fidelity, etc.). The input unit 100 may include a structure 102 made of a rigid material. Typically, the structure 102 is made of metal, such as stainless steel, aluminum, titanium, or a suitable metal alloy. To prepare the structure 102, molten metal is typically die-cast and then machined or extruded into a suitable form.

[0044] In some embodiments, the input unit 100 includes a housing that suppresses the exposure of the structure 102 to the surrounding environment. For example, the housing may prevent contamination, improve cleanliness, etc. Generally, the housing encloses substantially the entire structure 102, except for the conical resonator disposed along its bottom side. The following will refer to... Figure 1B The conical resonator is described in more detail in section C. The housing may be made of silicone rubber, polypropylene, polyethylene, or any other suitable material. Furthermore, in some embodiments, the housing contains additives that limit microbial growth, ultraviolet (UV) degradation, etc.

[0045] Figures 1B to 1CThe image includes a bottom perspective view of an input unit 100, which comprises a structure 102 having a distal portion 104 and a proximal portion 106. To initiate a stethoscope procedure, an individual (e.g., a healthcare professional, such as a doctor or nurse) can fix the proximal portion 106 of the input unit 100 against the surface of the body being examined. The proximal portion 106 of the input unit 100 may include a wider opening 108 of a conical resonator 110. The conical resonator 110 may be designed to direct sound waves collected through the wider opening 108 toward a narrower opening 112 that leads to a stethoscope microphone. Typically, the wider opening 108 is approximately 30 mm to 50 mm, 35 mm to 45 mm, or 38 mm to 40 mm. However, since the input unit 100 described herein may have automatic gain control, a smaller conical resonator may be used. For example, in some embodiments, the wider opening 108 is less than 30 mm, 20 mm, or 10 mm. Therefore, the input unit described herein may be able to support a wide variety of conical resonators with different sizes and designed for different applications.

[0046] With regard to the terms “distal” and “proximal”, unless otherwise specified, these terms refer to the relative position of the input unit 100 with respect to the body. For example, when referring to the input unit 100 suitable for being fixed to the body, “distal” may refer to a first position close to where a cable suitable for transmitting digital signals can be connected to the input unit 100, and “proximal” may refer to a second position close to where the input unit 100 contacts the body.

[0047] Figure 2 This image shows a cross-sectional side view of the input unit 200 of an electronic stethoscope system. Typically, the input unit 200 includes a structure 202 having an internal cavity defined therein. The structure 202 of the input unit 200 may have a conical resonator 204 designed to direct sound waves toward a microphone residing within the internal cavity. In some embodiments, a diaphragm 212 (also referred to as a "diaphragm") extends across a wider opening (also referred to as an "external opening") across the conical resonator 204. The diaphragm 212 can be used to hear sharp sounds, such as the sharp sounds often produced by the lungs. The diaphragm 212 may be formed from a thin plastic disc made of an epoxy glass fiber compound or glass fiber.

[0048] To improve the clarity of the sound waves collected by the conical resonator 204, the input unit 200 can be designed to simultaneously monitor sounds originating from different locations. For example, the input unit 200 can be designed to simultaneously monitor sounds originating from within the body being examined and sounds originating from the surrounding environment. Therefore, the input unit 200 may include at least one microphone 206 (referred to as a "stethoscope microphone") configured to generate audio data indicating internal sounds and at least one microphone 208 (referred to as an "ambient microphone") configured to generate audio data indicating ambient sounds. Each stethoscope microphone and ambient microphone may include a transducer capable of converting sound waves into electrical signals. The electrical signals generated by the stethoscope microphones and ambient microphones 206, 208 can then be digitized before being transmitted to the hub unit. Digitization allows the hub unit to easily clock or synchronize signals received from multiple input units. Digitization also ensures that the signals received by the hub unit from the input units have a higher fidelity than would otherwise be possible.

[0049] These microphones can be omnidirectional microphones designed to pick up sound from all directions or directional microphones designed to pick up sound from a specific direction. For example, input unit 200 may include a stethoscope microphone 206 oriented to pick up sound originating from a space adjacent to the external opening of the conical resonator 204. In such embodiments, ambient microphone 208 may be an omnidirectional or directional microphone. Alternatively, a group of ambient microphones 208 may be equidistantly spaced within the structure 202 of input unit 200 to form a phased array capable of capturing highly directional ambient sound to reduce noise and interference. Thus, stethoscope microphone 206 may be arranged to focus on the path of incoming internal sound (also referred to as the "stethoscope path"), while ambient microphone 208 may be arranged to focus on the path of incoming ambient sound (also referred to as the "ambient path").

[0050] Conventionally, electronic stethoscopes subject the electrical signals indicating sound waves to digital signal processing (DSP) algorithms, which filter out unwanted artifacts. However, such actions suppress almost all sounds within certain frequency ranges (e.g., 100Hz to 800Hz), thus greatly distorting the internal sounds of interest (e.g., those corresponding to inhalation, exhalation, or heartbeat). Here, however, the processor may employ an active noise cancellation algorithm that separately examines the audio data generated by stethoscope microphone 206 and the audio data generated by ambient microphone 208. More specifically, the processor may analyze the audio data generated by ambient microphone 208 to determine how the audio data generated by stethoscope microphone 206 should be modified (if any). For example, the processor may identify certain digital features that should be modified (e.g., since they correspond to internal sounds), attenuated (e.g., since they correspond to ambient sounds), or removed entirely (e.g., since they represent noise). This technique can be used to improve the clarity, detail, and quality of the sound recorded by input unit 200. For example, the application of noise cancellation algorithms can be an integral part of the denoising process employed by an electronic stethoscope system that includes at least one input unit 200.

[0051] For privacy purposes, the stethoscope 206 and ambient microphone 208 may not be allowed to record when the conical resonator 204 is away from the body guide. Therefore, in some embodiments, the stethoscope 206 and / or ambient microphone 208 do not begin recording until the input unit 200 is attached to the body. In such embodiments, the input unit 200 may include one or more accessory sensors 210a-c responsible for determining whether the structure 202 has been correctly mounted to the surface of the body.

[0052] Input unit 200 may include any subgroup of accessory sensors shown herein. For example, in some embodiments, input unit 200 may include only accessory sensors 210a-b positioned near a wider opening of the conical resonator 204. As another example, in some embodiments, input unit 200 may include only accessory sensors 210c positioned near a narrower opening (also referred to as an “internal opening”) of the conical resonator 204. Furthermore, input unit 200 may include different types of accessory sensors. For example, accessory sensor 210c may be an optical proximity sensor designed to emit light (e.g., infrared light) through the conical resonator 204 and then determine the distance between input unit 200 and the body surface based on the light reflected back into the conical resonator 204. As another example, accessory sensors 210a-c may be audio sensors designed to determine whether structure 202 is securely sealed against the body surface based on the presence of ambient noise (also referred to as “ambient noise”) by means of an algorithm programmed to determine the drop-off of a high-frequency signal.

[0053] For example, attachment sensors 210a-b can be pressure sensors designed to determine whether structure 202 is firmly sealed against the body surface based on the amount of applied pressure. Some embodiments of input unit 200 include each of these different types of attachment sensors. By taking into account the outputs of these attachment sensors 210a-c plus the aforementioned active noise cancellation algorithm, the processor may be able to dynamically determine the adhesion state. That is, the processor may be able to determine whether input unit 200 has formed a seal against the body based on the outputs of these attachment sensors 210a-c.

[0054] Figure 3 The diagram illustrates how one or more input units 302a-n can be connected to a hub unit 304 to form an electronic stethoscope system 300. In some embodiments, multiple input units are connected to the hub unit 304. For example, the electronic stethoscope system 300 may include four, six, or eight input units. Generally, the electronic stethoscope system 300 will include at least six input units. An electronic stethoscope system with multiple input units may be referred to as a "multi-channel stethoscope". In other embodiments, only one input unit is connected to the hub unit 304. For example, a single input unit can move across the body in a manner that simulates an array of multiple input units. An electronic stethoscope system with one input unit may be referred to as a "single-channel stethoscope".

[0055] like Figure 3As shown, each input unit 302a-n can be connected to the hub unit 304 via a corresponding cable 306a-n. Generally, the transmission path formed between each input unit 302a-n and the hub unit 304 via the corresponding cable 306a-n is designed to be substantially interference-free. For example, electronic signals can be digitized by the input unit 302a-n before being transmitted to the hub unit 304, and signal fidelity can be ensured by preventing the generation / contamination of electromagnetic noise. Examples of cables include ribbon cables, coaxial cables, Universal Serial Bus (USB) cables, High Definition Multimedia Interface (HDMI) cables, RJ45 Ethernet cables, and any other cables suitable for transmitting digital signals. Each cable includes a first end connected (e.g., via a physical port) to the hub unit 304 and a second end connected (e.g., via a physical port) to the corresponding input unit. Therefore, each input unit 302a-n may include a single physical port, and the hub unit 304 may include multiple physical ports. Alternatively, all input units 302a-n can be connected to hub unit 304 using a single cable. In such embodiments, the cable may include a first end capable of mating with hub unit 304 and a series of second ends, each capable of mating with a single input unit. Based on the number of second ends, such a cable may be referred to as, for example, a "two-to-one cable," a "four-to-one cable," or a "six-to-one cable."

[0056] When all input units 302a-n connected to hub unit 304 are in auscultation mode, the electronic stethoscope system 300 can employ an adaptive gain control algorithm programmed to compare internal sounds with ambient sounds. The adaptive gain control algorithm analyzes the target auscultation sound (e.g., normal breath sounds, wheezing, crackles, etc.) to determine if a sufficient sound level has been achieved. For example, the adaptive gain control algorithm can determine whether the sound level exceeds a predetermined threshold. The adaptive gain control algorithm can be designed to achieve gain control up to 100 times (e.g., in two different phases). The gain level can be adaptively adjusted based on the number of input units in input unit array 308 and the sound level recorded by the auscultation microphone in each input unit. In some embodiments, the adaptive gain control algorithm is programmed to be deployed as part of a feedback loop. Therefore, the adaptive gain control algorithm can apply gain to the audio recorded by the input unit, determine whether the audio exceeds a pre-programmed intensity threshold, and dynamically determine whether additional gain is necessary based on this determination.

[0057] Because the electronic stethoscope system 300 can deploy an adaptive gain control algorithm during post-processing, the input unit array 308 can collect information related to a wide variety of sounds caused by the heart, lungs, etc. Since the input units 302a-n in the input unit array 308 can be placed in different anatomical locations along the body surface (or on completely different parts of the body), the electronic stethoscope system 300 can simultaneously monitor different biometric characteristics (e.g., respiratory rate, heart rate, or the intensity of wheezing, crackles, etc.).

[0058] Figure 4 This is a high-level block diagram illustrating exemplary components of an input unit 400 and a hub unit 450 of an electronic stethoscope system. Embodiments of the input unit 400 and hub unit 450 may include... Figure 4 This includes any subgroup of components shown and additional components not shown herein. For example, input unit 400 may include a biometric sensor capable of monitoring biometric characteristics of the body, such as perspiration (e.g., based on skin moisture), temperature, etc. Additionally or alternatively, the biometric sensor may be designed to monitor breathing patterns (also known as "respiratory patterns"), record cardiac electrical activity, etc. As another example, input unit 400 may include an inertial measurement unit (IMU) capable of generating data from which posture, orientation, or position can be derived. An IMU is an electronic component designed to measure the force, angular rate, tilt angle, and / or magnetic field of an object. Generally, an IMU includes an accelerometer, gyroscope, magnetometer, or any combination thereof.

[0059] The input unit 400 may include one or more processors 404, a wireless transceiver 406, one or more microphones 408, one or more accessory sensors 410, a memory 412, and / or a power supply unit 414 electrically coupled to a power interface 416. These components may reside within a housing 402 (also referred to as a "structure").

[0060] As described above, microphone 408 converts acoustic waves into electrical signals. Microphone 408 may include a stethoscope microphone configured to generate audio data indicating internal sounds, an ambient microphone configured to generate audio data indicating ambient sounds, or any combination thereof. The audio data representing the values ​​of the electrical signals may be stored, at least temporarily, in memory 412. In some embodiments, processor 404 processes the audio data before transmitting it downstream to hub unit 450. For example, processor 404 may apply algorithms designed for digital signal processing, denoising, gain control, noise cancellation, artifact removal, feature recognition, etc. In other embodiments, processor 404 performs minimal processing before transmitting it downstream to hub unit 450. For example, processor 404 may simply append metadata identifying the input unit 400 to the audio data or check the metadata already added to the audio data by microphone 408.

[0061] In some implementations, input unit 400 and hub unit 450 transmit data to each other via a cable connected between corresponding data interfaces 418, 470. For example, audio data generated by microphone 408 can be forwarded to data interface 418 of input unit 400 for transmission to data interface 470 of hub unit 450. Alternatively, data interface 470 may be part of wireless transceiver 456. Wireless transceiver 406 may be configured to automatically establish a wireless connection with wireless transceiver 456 of hub unit 450. Wireless transceivers 406, 456 can communicate with each other via bidirectional communication protocols such as Near Field Communication (NFC), Wireless USB, etc. They communicate using cellular data protocols (such as LTE, 3G, 4G, or 5G) or proprietary point-to-point protocols.

[0062] Input unit 400 may include a power supply component 414 capable of providing power to other components residing within housing 402 when needed. Similarly, hub unit 450 may include a power supply component 466 capable of providing power to other components residing within housing 452. Examples of power supply components include rechargeable lithium-ion (Li-Ion) batteries, rechargeable nickel-metal hydride (NiMH) batteries, rechargeable nickel-cadmium (NiCad) batteries, etc. In some embodiments, input unit 400 does not include a dedicated power supply component and therefore must receive power from hub unit 450. Cables designed to facilitate power transfer (e.g., via physical connections of electrical contacts) may be connected between power interface 416 of input unit 400 and power interface 468 of hub unit 450.

[0063] For illustrative purposes only, the power channel (i.e., the channel between power interface 416 and power interface 468) and the data channel (i.e., the channel between data interface 418 and data interface 470) have been shown as separate channels. Those skilled in the art will recognize that these channels can be included in the same cable. Therefore, a single cable capable of carrying both data and power can be coupled between input unit 400 and hub unit 450.

[0064] Hub unit 450 may include one or more processors 454, wireless transceivers 456, displays 458, codecs 460, one or more light-emitting diode (LED) indicators 462, memory 464, and power supply units 466. These components may reside within housing 452 (also referred to as "structure"). As described above, embodiments of hub unit 450 may include any subgroup of these components as well as additional components not shown herein.

[0065] like Figure 4 As shown, embodiments of hub unit 450 may include a display 458 for presenting information such as the respiratory status or heart rate of the individual being examined, network connection status, power connection status, connection status of input unit 400, etc. The display 458 may be controlled via a tactile input mechanism (e.g., a button accessible along the surface of housing 452), an audio input mechanism (e.g., a microphone), etc. Alternatively, some embodiments of hub unit 450 may include an LED indicator 462 for operation guidance instead of display 458. In such embodiments, the LED indicator 462 may convey information similar to that presented on display 458. Again, some embodiments of hub unit 450 may include both display 458 and LED indicator 462.

[0066] Upon receiving audio data representing an electrical signal generated by the microphone 408 of the input unit 400, the hub unit 450 may provide the audio data to the codec 460, which is responsible for decoding the incoming data. The codec 460 may, for example, decode the audio data (e.g., by reversing the encoding applied by the input unit 400) to prepare it for editing, processing, etc. The codec 460 may be designed to process audio data generated by the stethoscope microphone in the input unit 400 and audio data generated by the ambient microphone in the input unit 400 sequentially or simultaneously.

[0067] The processor 454 can then process the audio data. Much like the processor 404 of the input unit 400, the processor 454 of the hub unit 450 can apply algorithms designed for digital signal processing, denoising, gain control, noise cancellation, artifact removal, feature recognition, etc. Some of these algorithms may be unnecessary if they have already been applied by the processor 404 of the input unit 400. For example, in some embodiments, the processor 454 of the hub unit 450 applies algorithms to discover diagnostically relevant features in the audio data, while in other embodiments, this action may be unnecessary if the processor 404 of the input unit 400 has already discovered diagnostically relevant features. Alternatively, the hub unit 450 can forward the audio data to a destination (e.g., a diagnostic platform running on a computing device or a decentralized system) for analysis, as discussed further below. Generally, diagnostically relevant features will correspond to patterns in the audio data whose values ​​match predetermined pattern definition parameters. For example, in some implementations, the processor 454 of the hub unit 450 applies algorithms to reduce noise in the audio data, thereby improving the signal-to-noise ratio (SNR), while in other implementations, these algorithms are applied by the processor 404 of the input unit 400.

[0068] In addition to power interface 468, hub unit 450 may also include a power port. The power port (also referred to as a "power jack") allows hub unit 450 to be physically connected to a power source (e.g., a power outlet). The power port may be able to mate with different connector types (e.g., C13, C15, C19). Additionally or alternatively, hub unit 450 may include a power receiver having an integrated circuit (also referred to as a "chip") capable of wirelessly receiving power from an external source. Similarly, for example, if input unit 400 and hub unit 450 are not physically connected to each other via a cable, input unit 400 may include a power receiver having a chip capable of wirelessly receiving power from an external source. The power receiver may be configured to receive power transmitted according to the Qi standard developed by the Wireless Power Consortium or some other wireless power standard.

[0069] In some implementations, the housing 452 of the hub unit 450 includes an audio port. An audio port (also called an "audio jack") is a socket that allows signals (such as audio) to be transmitted to a suitable plug for an accessory (such as headphones). An audio port typically includes one, two, three, or four contacts that facilitate the transmission of audio signals when a suitable plug is inserted into the audio port. For example, most headphones include a plug designed for a 3.5 mm audio port. Additionally or alternatively, the wireless transceiver 456 of the hub unit 450 may be able to transmit audio signals directly to wireless headphones (e.g., via NFC, wireless USB, Bluetooth, etc.).

[0070] As described above, the processor 404 of the input unit 400 and / or the processor 454 of the hub unit 450 can apply various algorithms to support different functions. Examples of such functions include attenuation of lost data packets in audio data, noise-related volume control, dynamic range compression, automatic gain control, equalization, noise suppression, and acoustic echo cancellation. Each function may correspond to a separate module residing in memory (e.g., memory 412 of the input unit 400 or memory 464 of the hub unit 450). Therefore, the input unit 400 and / or the hub unit 450 may include an attenuation module, a volume control module, a compression module, a gain control module, an equalization module, a noise suppression module, an echo cancellation module, or any combination thereof.

[0071] It should be noted that in some implementations, input unit 400 is configured to transmit audio data generated by microphone 408 directly to a destination other than hub unit 450. For example, input unit 400 may forward audio data to wireless transceiver 406 for transmission to a diagnostic platform responsible for analyzing the audio data. Audio data may be transmitted to the diagnostic platform as an alternative to or supplement to hub unit 450. If audio data is also forwarded to the diagnostic platform in addition to hub unit 450, input unit 400 may generate copies of the audio data and then forward these individual copies (e.g., forwarding to wireless transceiver 406 for transmission to the diagnostic platform, forwarding to data interface 418 for transmission to hub unit 450). As discussed further below, the diagnostic platform typically resides on a computing device communicatively connected to input unit 400, but aspects of the diagnostic platform may reside on input unit 400 or hub unit 450.

[0072] Additional information regarding electronic stethoscope systems can be found in U.S. Patent No. 10,555,717, the entire contents of which are incorporated herein by reference.

[0073] Overview of the diagnostic platform

[0074] Figure 5A network environment 500 including a diagnostic platform 502 is shown. Individuals can interact with the diagnostic platform 502 via an interface 504. For example, patients may be able to access the interface to provide information related to audio data, disease, treatment, and feedback. Similarly, healthcare professionals may be able to access the interface to review audio data and its analysis to determine appropriate diagnoses, monitor health, etc. At a higher level, the interface 504 can be intended to serve as an information dashboard for patients or healthcare professionals.

[0075] like Figure 5 As shown, the diagnostic platform 502 may reside in a network environment 500. Therefore, the diagnostic platform 502 may connect to one or more networks 506a-b. Networks 506a-b may include personal area networks (PANs), local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), cellular networks, the Internet, etc. Additionally or alternatively, the diagnostic platform 502 may be coupled to one or more computing devices via short-range wireless connectivity technologies such as Bluetooth, NFC, Wi-Fi Direct (also known as "Wi-Fi P2P").

[0076] Interface 504 can be accessed via a web browser, desktop application, mobile application, or over-the-top (OTT) application. For example, a healthcare professional may be able to access the interface to enter patient-related information. Such information may include name, date of birth, diagnosis, symptoms, or medications. Alternatively, this information may be automatically populated into the interface by a diagnostic platform 502 (e.g., based on data stored in a network-accessible server system 508), but the healthcare professional may be allowed to tailor the information as needed. As discussed further below, a healthcare professional may also be able to access the interface to present audio data and its analysis for review. Using this information, the healthcare professional may be able to easily determine the patient's condition (e.g., breathing or not breathing), provide a diagnosis (e.g., based on the presence of wheezing, crackles, etc.), etc. Therefore, interface 504 can be viewed on computing devices such as mobile workstations (also known as "medical carts"), personal computers, tablets, mobile phones, wearable electronic devices, and virtual reality or augmented reality systems.

[0077] In some implementations, at least some components of the diagnostic platform 502 are hosted locally. That is, a portion of the diagnostic platform 502 may reside on a computing device used to access one of the interfaces 504. For example, the diagnostic platform 502 may be embodied as a mobile application running on a mobile phone associated with a healthcare professional. However, it should be noted that the mobile application may communicatively connect to a network-accessible server system 508, on which other components of the diagnostic platform 502 are hosted.

[0078] In other implementations, the diagnostic platform 502 is entirely performed by a cloud computing service, such as Amazon Web Services. Google Cloud Platform TM or Microsoft Operation. In such embodiments, the diagnostic platform 502 may reside on a network-accessible server system 508, which includes one or more computer servers. These computer servers may include models, algorithms (e.g., for processing audio data, calculating respiratory rates, etc.), patient information (e.g., profiles, credentials, and health-related information such as age, date of birth, geographic location, disease classification, disease status, healthcare provider, etc.), and other assets. Those skilled in the art will recognize that this information may also be distributed within the network-accessible server system 508 and one or more computing devices or across decentralized network infrastructure such as blockchain.

[0079] Figure 6 An example of a computing device 600 capable of implementing a diagnostic platform 610 is shown, designed to generate outputs that aid in the detection, diagnosis, and monitoring of changes in a patient's health. As discussed further below, the diagnostic platform 610 can apply models to patient-associated audio data to identify the occurrence of respiratory events (also referred to as "breathing events" or "breathing"), and then apply algorithms to these respiratory events to gain insights into the patient's health. The terms "respiratory event" and "breathing event" can be used to refer to either inhalation or exhalation. For example, the algorithm can output a metric representing the respiratory rate, as discussed further below. Thus, the diagnostic platform 610 can not only discover patterns of diagnostically relevant values ​​in audio data but can also generate visualizations of these patterns in a manner that aids in understanding the patient's current health. For illustrative purposes, the implementation is described in the context of audio data generated by an electronic stethoscope system. However, those skilled in the art will recognize that audio data can be obtained from another source.

[0080] Typically, computing device 600 is associated with healthcare professionals or healthcare facilities. For example, computing device 600 may be a mobile workstation located in a hospital operating room, or it may be a mobile phone or tablet computer accessible to healthcare professionals when providing services to patients. Alternatively, computing device 600 may be a hub unit of an electronic stethoscope system.

[0081] The computing device 600 may include a processor 602, a memory 604, a display mechanism 606, and a communication module 608. Each of these components will be discussed in more detail below. Those skilled in the art will recognize that different combinations of these components may exist depending on the nature of the computing device 600.

[0082] Processor 602 may have general-purpose characteristics similar to a general-purpose processor, or processor 602 may be an application-specific integrated circuit (ASIC) that provides control functions to computing device 600. For example... Figure 6 As shown, the processor 602 can be directly or indirectly connected to all components of the computing device 600 for communication purposes.

[0083] Memory 604 can be constructed from any suitable type of storage medium, such as static random access memory (SRAM), dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or registers. In addition to storing instructions executable by processor 602, memory 604 can also store data generated by processor 602 (e.g., when executing modules of diagnostic platform 610). It should be noted that memory 604 is merely an abstract representation of the storage environment. Memory 604 can be constructed from actual memory chips or modules.

[0084] Display mechanism 606 can be any component operable to visually convey information. For example, display mechanism 606 can be a panel including LEDs, organic LEDs, liquid crystal elements, or electrophoretic elements. In an embodiment where computing device 600 represents a hub unit of an electronic stethoscope system, display mechanism 606 can be a display panel (e.g., Figure 4 The display 458) or LED indicator (e.g., Figure 4 (LED indicator 462). In some embodiments, the display mechanism 606 is touch-sensitive. Therefore, an individual may be able to provide input to the diagnostic platform 610 by interacting with the display mechanism 606. In embodiments where the display mechanism 606 is not touch-sensitive, an individual may be able to interact with the diagnostic platform 610 using control devices (not shown) such as a keyboard, physical elements (e.g., mechanical buttons or knobs), or pointing devices (e.g., a computer mouse).

[0085] The communication module 608 can manage communication between components of the computing device 600, or it can manage communication with other computing devices (e.g., Figure 5 The communication module 608 can be a wireless communication circuit designed to establish communication channels with other computing devices. Examples of wireless communication circuits include antenna modules configured for cellular networks (also known as "mobile networks") and chips configured for NFC, wireless USB, Bluetooth, etc.

[0086] For convenience, the diagnostic platform 610 may be referred to as a computer program residing in memory 604 and executed by processor 602. However, the diagnostic platform 610 may consist of software, firmware, or hardware implemented in or accessible to computing device 600. According to the embodiments described herein, the diagnostic platform 610 may include a training module 612, a processing module 614, a diagnostic module 616, an analysis module 618, and a graphical user interface (GUI) module 620.

[0087] Training module 612 is responsible for training the model to be used by diagnostic platform 610. Training can be performed in a supervised, semi-supervised, or unsupervised manner. For example, suppose training module 612 receives input instructing the training model to recognize breathing events in audio data. In this scenario, training module 612 can obtain an untrained model and then train the model using audio data labeled, for example, "breathing event" or "no breathing event." Therefore, labeled audio data can be provided to the model as training data, allowing the model to learn how to recognize breathing events. Typically, the model will learn to recognize patterns in the values ​​indicating breathing events in the audio data.

[0088] Processing module 614 can process the audio data obtained by diagnostic platform 610 into a format suitable for other modules. For example, processing module 614 can apply rules, heuristics, or algorithms to the audio data in preparation for analysis by diagnostic module 616. Similarly, processing module 614 can apply rules, heuristics, or algorithms to the output generated by diagnostic module 616 in preparation for analysis by analysis module 618. Therefore, processing module 614 is responsible for ensuring that appropriate data is accessible to other modules of diagnostic platform 610. Furthermore, processing module 614 is responsible for ensuring that the output generated by other modules of diagnostic platform 610 is suitable for storage (e.g., in memory 604) or transmission (e.g., via communication module 608).

[0089] The diagnostic module 616 may be responsible for identifying appropriate models to apply to the audio data acquired by the diagnostic platform 610. For example, the diagnostic module 616 may identify appropriate models based on the audio data, the patient, or the attributes of the electronic stethoscope system. These attributes may be specified in the metadata accompanying the audio data. For example, an appropriate model may be identified based on the anatomical region for which internal sounds have been recorded. Alternatively, the diagnostic module 616 may identify appropriate models based on metrics that will be generated by the analysis module 618. For example, if the task of the analysis module 618 is to calculate the respiratory rate, the diagnostic module 616 may identify modules capable of detecting respiratory events. The desired metrics may be specified by healthcare professionals or determined by the diagnostic platform 610.

[0090] Generally speaking, the model applied to the audio data by the diagnostic module 616 is one of several models stored in the memory 604. These models can be associated with different respiratory events, minor illnesses, etc. For example, a first model can be designed and subsequently trained to recognize inhalation and exhalation, and a second model can be designed and subsequently trained to recognize wheezing or crackling sounds.

[0091] At a higher level, each model can represent a collection of algorithms that, when applied to audio data, produce outputs that convey information that can provide insights into a patient's health. For example, if the model applied by the diagnostic module 616 identifies inhalation and exhalation, the output can be used to determine whether the patient is breathing normally or abnormally. Similarly, if the model applied by the diagnostic module 616 identifies wheezing or crackling sounds, the output can be used to determine whether the patient has a given disease.

[0092] In some scenarios, the output generated by the model applied by the diagnostic module 616 may not be particularly useful on its own. The analysis module 618 can be responsible for considering the context of these outputs in a more comprehensive sense. For example, the analysis module 618 can generate one or more measures representing a patient's health based on the output generated by the model applied by the diagnostic module 616. These measures can provide insights into a patient's health without requiring a full analysis or understanding of the output generated by the model. Suppose, for example, that the model applied by the diagnostic module 616 identifies respiratory events based on the analysis of audio data. While knowing about respiratory events may be useful, healthcare professionals may be more interested in measures such as respiratory rate. The analysis module 618 can calculate the respiratory rate based on the respiratory events identified by the diagnostic module 616.

[0093] The GUI module 620 is responsible for determining how information is presented for review on the display device 606. Various types of information can be presented depending on the nature of the display device 606. For example, information derived, inferred, or otherwise obtained by the diagnostic module 616 and analysis module 618 can be presented on the interface for display to healthcare professionals. Similarly, visual feedback can be presented on the interface to indicate when a patient has experienced changes in their health.

[0094] Figure 7 This includes an example workflow diagram illustrating how audio data obtained from the diagnostic platform can first be processed by a backend service, and then the analysis of the audio data can be presented on the interface for review. Figure 7As shown, data can be acquired and processed by the "back end" of the diagnostic platform, while analysis of that data can be presented by the "front end" of the diagnostic platform. Individuals (such as healthcare professionals and patients) may also be able to interact with the diagnostic platform through their front end. For example, commands can be issued through the interface shown on the display. Furthermore, the diagnostic platform may be able to generate notifications, as discussed further below. For example, if the diagnostic platform determines that a patient's respiratory rate has fallen below a defined threshold, the platform can generate a notification that serves as an alert. This notification may be presented visually via a display and / or audibly via a speaker.

[0095] Methods for determining respiratory rate

[0096] Conventional methods for calculating respiratory rate have several drawbacks.

[0097] Some methods rely on observing inspiration and expiration over relatively long time intervals. These "observation intervals" can easily last 60 seconds or longer. Because the observation intervals are so long, these methods cannot account for short-term events such as apnea. Therefore, healthcare professionals may be largely (if not completely) unaware of temporary cessation of breathing because the impact on the respiratory rate over the entire observation interval will be minimal. Misunderstanding the respiratory rate can cause significant harm, as it may require immediate intervention for respiratory arrest.

[0098] Other methods rely on monitoring chest movement or end-tidal carbon dioxide (CO2) in exhaled air. However, these methods are impractical or unsuitable in many situations. For example, chest movement may not be an accurate indicator of respiratory events, especially when the patient is under general anesthesia or experiencing a medical event (e.g., epilepsy), and the composition of exhaled air may be unknown if the patient is not wearing a mask.

[0099] This article introduces a method to address these shortcomings by calculating respiratory rate through analysis of audio data. As discussed above, this method relies on detecting inspiration and expiration in a consistent and accurate manner. During auscultation, these respiratory events are important for diagnostic determination by healthcare professionals.

[0100] At a higher level, this method involves two phases: a first phase of detecting inhalation and exhalation, followed by a second phase of using these inhalations and exhalations to calculate the respiratory rate. The first phase can be referred to as the "detection phase," and the second phase as the "calculation phase."

[0101] A. Respiratory event detection

[0102] Figure 8This includes a high-level diagram of the computational pipeline that can be used by the diagnostic platform during the detection phase. One advantage of this computational pipeline (also known as the "computational framework") is its modular design. Each "unit" can be tested individually and then tuned for optimal overall performance. Furthermore, the output of some units can be used for multiple purposes. For example, spectrograms generated during preprocessing can be provided as input to the model and / or published to an interface for real-time review.

[0103] For simplicity, this framework is divided into three parts: preprocessing, analysis, and postprocessing. Preprocessing may include not only processing the audio data but also employing feature engineering techniques. Analysis may involve employing a model trained to recognize breathing events. As mentioned above, this model may include a neural network designed to produce a sequence of detections (e.g., breathing events) as output, rather than individual detections (e.g., classifications or diagnoses). Each detection may represent a separate prediction made independently by the model. Finally, postprocessing may include examining and / or splitting the detections produced by the model. Typically, this is done by a processing module (e.g., ...). Figure 6 The processing module 614) performs preprocessing, which is performed by the diagnostic module (e.g., Figure 6 The diagnostic module 616) performs the analysis, and the analysis module (e.g., Figure 6 The analysis module 618) performs post-processing.

[0104] Further details relating to each section will be provided below in the context of the examples. Those skilled in the art will recognize that the figures provided below are intended to be illustrative only. An important aspect of the framework is its flexibility, and therefore other figures may be applicable or appropriate in other scenarios.

[0105] I. Preprocessing

[0106] In this example, recorded audio data representing internal sounds emitted by the lungs is processed at a sampling frequency equivalent to 4 kHz. A high-pass filter with a 10th-order cutoff frequency of 80 Hz is then applied to the audio data to remove electrical interference (approximately 60 Hz) and internal sounds emitted by the heart (approximately 1 Hz to 2 Hz) or another internal organ. Therefore, the diagnostic platform can apply a high-pass filter with a sufficient cutoff frequency to filter sounds emitted by another internal organ of less interest. The filtered audio data is then processed using a Short-Time Fourier Transform (STFT). In this example, the STFT has a Hamming window with a window size of 256 and an overlap ratio of 0.25. Thus, approximately 15 seconds of signal can be transformed into a corresponding spectrogram of size 938 × 129. To utilize the spectral information of the internal sounds of interest, the diagnostic platform extracts (i) the spectrogram, (ii) the Mel-frequency cepstral coefficients (MFCC), and (iii) the energy summation. In this example, the spectrogram is a 129-bin logarithmic amplitude spectrogram. For MFCC, the diagnostic platform extracts 20 static coefficients, 20 differential coefficients, and 20 acceleration coefficients. For this, the platform uses 40 Mel bands within the frequency range of 0Hz to 4,000Hz. The width used to calculate the differential and acceleration coefficients is 9 frames. This produces 60 bin vectors per frame. Simultaneously, the platform calculates the energy summation for three different frequency bands (i.e., 0Hz to 250Hz, 251Hz to 500Hz, and 501Hz to 1,000Hz), producing three values ​​per frame.

[0107] After extracting these features, the diagnostic platform concatenates them to form a 938×193 feature matrix. Then, the platform applies minimum-maximum normalization to each feature, ensuring that the normalized features fall within the range of values ​​between 0 and 1.

[0108] II. Analysis

[0109] As mentioned above, the features extracted during preprocessing are fed as input to several models trained to recognize respiratory events. To determine the best model, six models are tested using the same extracted features. These models include Uni-RNN, Uni-LSTM, Uni-GRU, Bi-RNN, Bi-LSTM, and Bi-GRU. These models are collectively referred to as the "baseline models".

[0110] The architecture of these baseline models is shown in Figure 9The baseline model is designed to detect respiratory events based on the analysis of audio data. The first and second layers are recursive layers, which can be RNNs, LSTMs, or GRUs. These recursive layers process temporal information from the features. Since breathing is typically periodic, these recursive layers learn the nature of the respiratory cycle from labeled examples (also known as "training data"). To detect the start and end times of respiratory events in the audio data, the diagnostic platform uses a temporally distributed fully connected layer as the output layer. This approach produces a model that outputs a sequence of detections (e.g., inspiration or no movement) rather than a single detection. A sigmoid function is used as the activation function in the temporally distributed fully connected layer.

[0111] Each output produced by the baseline model is a detection vector of size 938×1. If this value is above a threshold, each element in the vector is set to 1 to indicate the presence of inhalation or exhalation at the corresponding time point; otherwise, the value is set to 0. In this example, a single-task learning method is used for the baseline model, but a multi-task learning method can also be used.

[0112] For the baseline model, Adaptive Moments Estimation (ADAM) is used as the optimizer. ADAM is a method that uses stochastic optimization to compute an adaptive learning rate for the parameters, a crucial process in deep learning and machine learning. The initial learning rate is set to 0.0001 with a step decay of 0.2× when the validation loss does not decrease within 10 epochs. The learning process stops when there is no improvement within 50 consecutive epochs.

[0113] III. Post-processing

[0114] The detection vectors generated by the baseline model can then be further processed for various purposes. For example, a diagnostic platform can transform each predicted vector from a frame into a time-varying format for real-time monitoring by healthcare professionals. Furthermore, it can be used as follows: Figure 9 This applies domain knowledge as shown. Since breathing is known to be performed by a living organism, the duration of respiratory events is typically within a certain range. When a detection vector indicates the presence of consecutive respiratory events (e.g., inspiration) with small intervals between them, the diagnostic platform can examine the continuity of these respiratory events and then decide whether to merge them. More specifically, the diagnostic platform can calculate the frequency difference (|p|) between the energy peaks of the j-th respiratory event and the i-th respiratory event. j -p iThe intervals between these respiratory events are less than T seconds. If the difference is less than a given threshold P, these respiratory events are merged into a single respiratory event. In this example, T is set to 0.5 seconds and P is set to 25 Hz. If a respiratory event is shorter than 0.05 seconds, the diagnostic platform can simply delete the respiratory event entirely. For example, the diagnostic platform can adjust the label applied to the corresponding segment to indicate that no respiratory event has occurred.

[0115] IV. Task Definition and Assessment

[0116] In this example, the diagnostic platform performs two different tasks.

[0117] The first task is to classify the audio data into segments. To achieve this, the recording of each breath event is initially transformed into a spectrogram. The temporal resolution of the spectrogram depends on the window size and overlap ratio of the STFT, as discussed above. For convenience, these parameters are fixed such that each spectrogram is a matrix of size 938 × 128. Therefore, each recording is divided into 938 segments, and based on... Figure 10 The actual labels shown are used to automatically tag each segment.

[0118] After the records undergo preprocessing and analysis, the diagnostic platform will be able to access the corresponding output with a sequence of detections of size 938×1. This output can be referred to as the "inference result". By comparing the sequence detections with the true segments, the diagnostic platform can define true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). The sensitivity and specificity of the model used to classify these segments can then be calculated.

[0119] The second task is to detect breathing events in the audio data. After obtaining the sequence detection performed by the model, the diagnostic platform can assemble connection segments with the same label into corresponding breathing events. For example, the diagnostic platform can specify that a connection segment corresponds to inspiration, expiration, or no action. Furthermore, the diagnostic platform can determine the start and end times of each assembled breathing event. In this example, the Jacobian index (J1) is used to determine whether a breathing event predicted by the model correctly matches a real event. If the value of J1 is greater than 0.5, the diagnostic platform designates the assembled breathing event as a TP event. If the value of J1 is greater than 0 but falls below 0.5, the diagnostic platform designates the assembled breathing event as an FN event. If the value of J1 is 0, the diagnostic platform designates the assembled breathing event as an FP event.

[0120] To evaluate performance, the diagnostic platform examined the accuracy, sensitivity, positive predictions, and F1 score of each baseline model. All bidirectional models outperformed the unidirectional models. This result is primarily attributed to the increasing complexity (and number of trainable parameters) of bidirectional models. Overall, Bi-GRU shows the most promise among these baseline models, but there may be other baseline models that are more suitable than Bi-GRU in certain situations.

[0121] B. Respiratory Rate Calculation

[0122] Reliable estimation of respiratory rate plays a crucial role in the early prediction of various diseases. Abnormal respiratory rates can indicate different diseases as well as other pathological or psychological factors. The search for a clinically acceptable and accurate technique for continuously establishing respiratory rate has proven to be futile. Several methods have been developed in an attempt to fill this clinical gap discussed above, but none have gained sufficient trust from healthcare professionals or become standard treatment.

[0123] Figure 11A Includes high-level illustrations of the algorithmic methods used to estimate respiratory rate. (e.g.) Figure 11A As shown, this algorithm relies on respiratory events predicted by a model employed by the diagnostic platform. Accurate identification of respiratory events is crucial for estimating the respiratory rate; therefore, the detection and calculation phases must be performed with consistent high accuracy. Furthermore, the detection and calculation phases can be performed continuously, enabling near real-time estimation of the respiratory rate (e.g., every few seconds rather than every 30 to 60 seconds).

[0124] At a higher level, this algorithmic approach relies on tracking the occurrence of respiratory events and then continuously calculating the respiratory rate using a sliding window. Initially, the algorithm executed by the diagnostic platform defines the starting point of the window, which contains a portion of the acquired audio data used for analysis. The diagnostic platform can then monitor respiratory events inferred by the model as discussed above to define the ending point of the window. The starting and ending points of the window can be defined such that its boundaries contain a predetermined number of respiratory events. Here, for example, the window includes three respiratory events, but in other embodiments, it may include more or fewer than three respiratory events. In an embodiment where the window is programmed to include three respiratory events, the starting point of the window corresponds to the beginning of the first respiratory event 1102 included in the window, and the ending point of the window corresponds to the beginning of the fourth respiratory event 1108. Although the fourth respiratory event 1108 defines the ending point of the window, the fourth respiratory event 1108 is not initially included in the window.

[0125] like Figure 11AAs shown, the diagnostic platform can calculate the respiratory rate based on the intervals between inspiratory breaths included in this window. Here, the window includes a first respiratory event 1102, a second respiratory event 1104, and a third respiratory event 1106. As mentioned above, a fourth respiratory event 1108 can be used to define the end point of this window. The first time period, indicated by "I", is defined by the time interval between the start of the first respiratory event 1102 and the start of the second respiratory event 1104. The second time period, indicated by "II", is defined by the time interval between the start of the second respiratory event 1104 and the start of the third respiratory event 1106. The third time period, indicated by "III", is defined by the time interval between the start of the third respiratory event 1106 and the start of the fourth respiratory event 1108.

[0126] Each time segment corresponds to multiple frames (and therefore time intervals). Suppose, for example, the audio data obtained by a diagnostic platform is a record with a total duration of 15 seconds, and during processing, this record is divided into 938 segments, each with a duration of approximately 0.016 seconds. The term "segment" is used interchangeably with the term "frame". Here, the first, second, and third time segments are approximately 5 seconds long, but the first respiratory event, the second respiratory event, and the third respiratory events 1102, 1104, and 1106 are of different lengths in terms of segments and seconds. However, those skilled in the art will recognize that the time segments between respiratory events need not be identical.

[0127] like Figure 11A As shown, the respiratory rate can be calculated per frame based on these time periods. For example, consider the fifteenth second, represented by "A". At this time, the third time period has not yet been defined because the third respiratory event 1106 has not yet started. The respiratory rate can be calculated by dividing 120 by the sum of the first and second time periods in seconds. Starting from the sixteenth second, represented by "B", the diagnostic platform learns of the third respiratory event 1106. Therefore, starting from the sixteenth second, the respiratory rate can be calculated by dividing 120 by the sum of the second and third time periods in seconds.

[0128] Typically, when an additional respiratory event is identified, the diagnostic platform calculates the respiratory rate continuously. In other words, when a new time period is identified, the respiratory rate can be calculated in a "rolling" manner. Whenever a respiratory event is identified (and thus a new time period is defined), the diagnostic platform uses the most recent time period to calculate the respiratory rate. For example, suppose the first and second time periods correspond to 5 seconds and 7 seconds, respectively. In this scenario, the respiratory rate would be 10 breaths per minute (i.e., For example, suppose the first and second time periods correspond to 4 seconds and 4 seconds respectively. In this scenario, the breathing rate would be 15 breaths per minute (i.e., After another time period (i.e., the third time period) has been identified, the second and third time periods can be used instead of the first and second time periods to calculate the respiratory rate.

[0129] As mentioned above, since breathing is known to be performed by a living organism, the duration of respiratory events is typically within a certain range. Therefore, the algorithm used by a diagnostic platform to calculate the respiratory rate can be programmed to increment the denominator if no new respiratory event is inferred within a certain number of frames or seconds. Figure 11A In this algorithm, the denominator is programmed to increment in response to the determination that no new respiratory event is detected within 6 seconds. Therefore, at 22 seconds (represented by "C"), the diagnostic platform can calculate the respiratory rate by dividing 120 by the sum of the second time interval plus 6. Simultaneously, at 23 seconds (represented by "D"), the diagnostic platform can calculate the respiratory rate by dividing 120 by the sum of the second time interval plus 7. For example... Figure 11A As shown, the diagnostic platform can continue to increase the denominator every second until a new respiratory event is detected.

[0130] This method ensures near real-time updates to the respiratory rate to reflect changes in the patient's health. If a patient stops breathing completely, the respiratory rate calculated by the diagnostic platform may not immediately have a value of zero, but the respiratory rate will show a downward trend almost instantaneously, which can serve as a warning to healthcare professionals.

[0131] To address scenarios where the respiratory rate calculated by the diagnostic platform is abnormally low or high, the algorithm can be programmed with a first threshold (also known as the "lower threshold") and a second threshold (also known as the "upper threshold"). If the respiratory rate drops below the lower threshold, the algorithm can indicate that a warning should be issued by the diagnostic platform. Furthermore, if the diagnostic platform determines that the respiratory rate has dropped below the lower threshold, it can behave as if the respiratory rate is actually zero. This approach ensures that the diagnostic platform can easily respond to scenarios where no respiratory events are detected over extended time intervals without requiring the respiratory rate to actually reach zero. Similarly, if the respiratory rate exceeds the upper threshold, the algorithm can indicate that a warning should be issued by the diagnostic platform. The lower and upper thresholds can be adjusted manually or automatically based on the patient, the services provided to the patient, etc. For example, the lower and upper thresholds could be 4 and 60 for pediatric patients, and 4 and 35 for adult patients, respectively.

[0132] For simplicity, (for example, to) Figure 13AThe respiratory rate published for review (as shown in interface B) can be an integer value, but it can be calculated to one or two decimal places (i.e., the tenths or hundredths). This helps avoid confusion between healthcare professionals and patients, especially in emergency situations. As mentioned above, the respiratory rate published for review can visually change if it falls below a lower threshold or exceeds an upper threshold. For example, if the respiratory rate falls below the lower threshold, the diagnostic platform can publish a placeholder element (e.g., "-" or "--") intended to indicate that the respiratory rate is undetectable. If the respiratory rate exceeds the upper threshold, the diagnostic platform can publish another placeholder element (e.g., "35+" or "45+") intended to indicate that the respiratory rate is abnormally high. In many cases, healthcare professionals are more interested in understanding when the respiratory rate is abnormally high than in what the exact respiratory rate is.

[0133] Generally, when recording occurs, the diagnostic platform processes the audio data stream acquired in near real-time. In other words, when a new respiratory event is detected as discussed above, the diagnostic platform can continuously calculate the respiratory rate. However, in some implementations, the diagnostic platform does not calculate the respiratory rate for every frame to conserve processing resources. Instead, the diagnostic platform can calculate the respiratory rate periodically, but typically still frequently (e.g., every few seconds) to ensure that healthcare professionals are aware of any changes in the patient's health.

[0134] Figure 11B This illustrates how a diagnostic platform can calculate respiratory rate using a sliding window within a record of a predetermined length (e.g., 12, 15, or 20 seconds) that is updated at a predetermined frequency (e.g., every 2, 3, or 5 seconds). Figure 11B In this context, the bounding box represents the boundary of a sliding window within which audio data is processed. After a predetermined amount of time has elapsed, the boundary of the sliding window moves. Here, for example, the sliding window covers 15 seconds and moves by 3 seconds, thus retaining 12 seconds of audio data previously contained within the sliding window. This approach allows for frequent calculation of the respiratory rate (e.g., whenever the boundary of the sliding window moves) without excessively consuming available processing resources.

[0135] Another method for calculating respiratory rate relies on autocorrelation rather than ML or ALI to perform auscultation. At a higher level, the term "autocorrelation" refers to the process of relating a signal to a delayed copy of that signal, which varies with delay. Autocorrelation analysis is a common mathematical tool for identifying recurring patterns and is therefore frequently used in digital signal processing to derive or infer information related to recurring events.

[0136] In an implementation where the diagnostic platform uses autocorrelation for detection, the platform uses STFT to convert the recorded audio data into a spectrogram as discussed above. The autocorrelation coefficient is then calculated and plotted against this record to determine the interval between inhalation and exhalation, such as... Figure 12 As shown, the diagnostic platform first normalizes the autocorrelation coefficients to reduce noise. Then, a high-pass filter can be applied to the autocorrelation coefficients to eliminate those that drop below a threshold (e.g., 15 Hz). Furthermore, the diagnostic platform can perform detrending modifications to further refine the autocorrelation coefficients.

[0137] After processing the autocorrelation coefficient, the diagnostic platform can calculate the respiratory rate by dividing 60 by the respiratory interval (R1) between the first and second peaks. The respiratory rate index determines which peaks to select, with high respiratory rate indices being preferred. Generally, a respiratory rate index below approximately 0.6 indicates an unstable respiratory rate and can therefore be discarded by the diagnostic platform.

[0138] Figures 13A to 13B Includes examples of interfaces that can be generated by the diagnostic platform. Figure 13A This demonstrates how spectrograms generated by a diagnostic platform can be presented in near real-time to facilitate diagnostic determination by healthcare professionals. Meanwhile, Figure 13B This illustrates how digital elements (also known as "graphical elements") can be overlaid on a spectrogram to provide additional insights into a patient's health. Here, for example, vertical bands are overlaid on portions of the spectrogram that the diagnostic platform has identified corresponding to respiratory events. It should be noted that in some embodiments, these vertical bands cover the entire respiratory cycle (i.e., inspiration and expiration), while in other embodiments, these vertical bands cover only a portion of the respiratory cycle (e.g., inspiration only).

[0139] The respiratory rate can be published (e.g., in the upper right corner) for review by healthcare professionals. If the diagnostic platform determines the respiratory rate is abnormal, the published value can be visually altered in some way. For example, the published value could be presented in a different color, at a different (e.g., larger) size, or periodically removed and subsequently republished to "flash" and attract the attention of healthcare professionals. The diagnostic platform can determine an abnormal respiratory rate if the value drops below a lower threshold (e.g., 4, 5, or 8) or exceeds an upper threshold (e.g., 35, 45, or 60), or if no respiratory event is detected within a predetermined amount of time (e.g., 10, 15, or 20 seconds).

[0140] like Figure 13BAs shown, the interface may include various icons associated with different functions supported by the diagnostic platform. When selected, the record icon 1302 initiates recording of the content displayed on the interface. For example, the recording may include a spectrogram along with any digital elements and corresponding audio data. When selected, the freeze icon 1304 freezes the interface as currently shown. Therefore, when the freeze icon 1304 is selected, no additional content may be published to the interface. When the freeze icon 1304 is selected again, the content may be presented on the interface again. When selected, the control icon 1306 controls various aspects of the interface. For example, the control icon 1306 may allow healthcare professionals to lock the interface (e.g., to semi-permanently or temporarily prevent further changes), change the layout of the interface, change the color scheme of the interface, change the input mechanism of the interface (e.g., touch to voice), etc. Supplementary information may also be presented on the interface. Figure 13B For example, the interface includes information related to the status of the active noise cancellation (ANC), playback, and connected electronic stethoscope system. Therefore, healthcare professionals may be able to easily observe whether the computing device residing on the diagnostic platform (and showing this interface) is communicatively connected to the electronic stethoscope system responsible for generating audio data.

[0141] Methods for calculating respiratory rate

[0142] Figure 14 A flowchart depicts a process 1400 for detecting respiratory events by analyzing audio data and subsequently calculating the respiratory rate based on these events. Initially, the diagnostic platform acquires recorded audio data representing sounds produced by the patient's lungs (step 1401). In some embodiments, the diagnostic platform receives the audio data directly from an electronic stethoscope system. In other embodiments, the diagnostic platform acquires the audio data from a storage medium. For example, a patient may be allowed to upload self-recorded audio data to the storage medium. Similarly, a healthcare professional may be allowed to upload self-recorded audio data to the storage medium.

[0143] The diagnostic platform can then process the audio data to prepare it for analysis by the trained model (step 1402). For example, the diagnostic platform can apply a high-pass filter to the audio data and then generate (i) a spectrogram, (ii) a series of MFCCs, and (iii) a series of values ​​representing the energy summed across different frequency bands of the spectrogram based on the filtered audio data. The diagnostic platform can concatenate these features into a feature matrix. More specifically, the diagnostic platform can concatenate the spectrogram, the series of MFCCs, and the series of values ​​into a feature matrix, which can be provided as input to the trained model. In some implementations, the diagnostic platform performs minimum-maximum normalization on the feature matrix, such that each entry has a value between 0 and 1.

[0144] The diagnostic platform can then apply the trained model to the audio data (or the analysis of the audio data) to produce a vector comprising entries arranged in chronological order (step 1403). More specifically, the diagnostic platform can apply the trained model to the feature matrix mentioned above. Each entry in the vector can represent a detection produced by the trained model, indicating whether a corresponding segment of the audio data represents a breathing event. Each entry in the vector can correspond to a different segment of the audio data, but all entries in the vector can correspond to segments of audio data of equal duration. Suppose, for example, the audio data obtained by the diagnostic platform represents a recording with a total duration of 15 seconds. As part of preprocessing, the recording can be divided into 938 segments of equal duration. For each of these segments, the trained model can produce an output indicating a detection of whether a corresponding portion of the audio data represents a breathing event.

[0145] The diagnostic platform can then perform post-processing of the entries in the vector as mentioned above (step 1404). For example, the diagnostic platform can examine the vector to identify a pair of entries corresponding to the same type of respiratory event and spaced less than a predetermined number of times apart, and then merge the pair of entries to indicate that the pair of entries represents a single respiratory event. This operation can be performed to ensure that respiratory events are not missed because there are segments in the model indicating no respiratory event. Additionally or alternatively, the diagnostic platform can examine the vector to identify a series of consecutive entries corresponding to a respiratory event of the same type that is less than a predetermined length, and then adjust the label associated with each entry in the series of consecutive entries to remove the respiratory event represented by the series of consecutive entries. This operation can be performed to ensure that respiratory events less than a predetermined length (e.g., 0.25, 0.50, or 1.00 seconds) are not identified.

[0146] The diagnostic platform can then identify (i) a first breathing event, (ii) a second breathing event, and (iii) a third breathing event by examining the vector (step 1405). The second breathing event may follow the first breathing event, and the third breathing event may follow the second breathing event. Each breathing event may correspond to at least two consecutive entries in the vector, where the at least two consecutive entries indicate that the corresponding segment of audio data represents the breathing event. The number of consecutive entries may correspond to the minimum length of a breathing event performed by the diagnostic platform.

[0147] The diagnostic platform can then determine (a) a first time period between the first and second respiratory events and (b) a second time period between the second and third respiratory events (step 1406). As discussed above, the first time period may extend from the beginning of the first respiratory event to the beginning of the second respiratory event, and the second time period may extend from the beginning of the second respiratory event to the beginning of the third respiratory event. The diagnostic platform can then calculate the respiratory rate based on the first and second time periods (step 1407). For example, the diagnostic platform can divide 120 by the sum of the first and second time periods to establish the respiratory rate.

[0148] Figure 15 A flowchart is depicted for a process 1500 for calculating the respiratory rate based on the analysis of audio data containing sounds emitted by a patient's lungs. Initially, the diagnostic platform obtains a vector comprising entries arranged in chronological order (step 1501). At a higher level, each entry in this vector can indicate whether a detection related to a corresponding segment of the audio data represents a respiratory event. As discussed above, this vector can be generated as output by a model applied by the diagnostic platform to the audio data or the analysis of the audio data.

[0149] The diagnostic platform can then identify (i) the first respiratory event, (ii) the second respiratory event, and (iii) the third respiratory event by examining the entries in the vector (step 1502). Typically, this is done by examining the vector to identify consecutive entries associated with the same type of respiratory event. For example, the diagnostic platform can parse the vector to identify events with the same detection (e.g., "inspiration" or "no action"), such as... Figure 9 (As shown) a series of consecutive entries. Therefore, each respiratory event can correspond to a series of consecutive entries that (i) exceed a predetermined length and (ii) indicate that the corresponding segment of the audio data represents the respiratory event. As discussed above, the diagnostic platform can perform post-processing to ensure that false positives and false negatives do not affect its ability to identify respiratory events.

[0150] The diagnostic platform can then determine (a) a first time period between the first and second respiratory events and (b) a second time period between the second and third respiratory events (step 1503). Figure 15 Step 1503 can be similar to Figure 14 Step 1406. Then, the diagnostic platform can calculate the respiratory rate based on the first time period and the second time period (step 1504). For example, the diagnostic platform can divide 120 by the sum of the first time period and the second time period to establish the respiratory rate.

[0151] In some implementations, the diagnostic platform is configured to trigger the display of an interface including a spectrogram corresponding to the audio data (step 1505). An example of such an interface is shown below. Figure 13A To B. In this type of implementation, the diagnostic platform may publish the respiratory rate to the interface to facilitate diagnostic determination by healthcare professionals (step 1506). Furthermore, the diagnostic platform may compare the respiratory rate to thresholds discussed above (step 1507) and then take appropriate action based on the result of the comparison (step 1508). For example, if the diagnostic platform determines that the respiratory rate has fallen below a lower threshold, the platform may generate a notification indicating this situation. Similarly, if the diagnostic platform determines that the respiratory rate has exceeded an upper threshold, it may generate a notification indicating this situation. Typically, this notification is presented on the interface on which the respiratory rate is published. For example, the respiratory rate may be visually altered in response to determining that it has fallen below a lower threshold or exceeded an upper threshold.

[0152] Unless contrary to physical possibility, it is conceivable that the above steps can be performed in various sequences and combinations.

[0153] For example, a repeatable procedure 1500 allows the respiratory rate to be continuously recalculated upon the detection of an additional respiratory event. However, as mentioned above, the diagnostic platform can be programmed to increment the denominator (and thus decrease the respiratory rate) when no respiratory event is detected.

[0154] For example, respiratory rate can be calculated using more than two time periods. Generally, calculating respiratory rate using a single time period is undesirable because the value can fluctuate too much within a relatively short timeframe, offering little benefit to healthcare professionals. However, diagnostic platforms can calculate respiratory rate using three or more time periods. Suppose, for example, that the diagnostic platform's task is to calculate respiratory rate using three time periods defined by four respiratory events. In this scenario, the diagnostic platform can primarily follow the guidelines above. Figures 14 to 15 The steps are performed as described. However, the diagnostic platform then divides 180 by the sum of the first, second, and third time periods. While the process remains largely the same for calculating the respiratory rate, the diagnostic platform multiplies 60 by "N", where "N" is the number of time periods, and divides by the sum of those "N" time periods.

[0155] In some implementations, additional steps may also be included. For example, diagnostic-related insights (such as the presence of respiratory abnormalities) may be published to an interface (e.g., Figure 13B The interface (or data structure) can be used to represent a profile associated with a patient. Therefore, diagnostic insights can be stored in a data structure encoded with audio data associated with the corresponding patient.

[0156] Exemplary use cases of electronic stethoscope systems and diagnostic platforms

[0157] The demand for moderate to severe anesthesia has gradually surpassed that for traditional intubation surgeries, and anesthesia has become a mainstream treatment in healthcare facilities. However, the risk of respiratory arrest or airway obstruction remains. Direct and continuous monitoring of the patient's respiratory status is lacking, and the auscultation-driven approach discussed above represents a solution to this problem.

[0158] According to the World Health Organization (WHO) Surgical Safety Guidelines, moderate non-intubated anesthesia requires consistent monitoring by the attending healthcare professional. This has traditionally been done in several ways, including confirming ventilation through auscultation, monitoring end-tidal CO2, or verbal communication with the patient. However, these traditional methods are impractical (if not impossible) in many cases.

[0159] The auscultation-driven methods discussed above can be used in situations where these traditional methods are not suitable. For example, auscultation-driven methods can be used in situations such as orthopedic surgery, gastrointestinal endoscopy, and dental procedures, where sedation is often used to reduce pain. As another example, auscultation-driven methods can be used in situations where there is a high risk of disease transmission (e.g., the treatment of patients with respiratory illnesses). One of the most important applications of the methods described in this article is the ability to quickly detect airway obstruction through auscultation. Electronic stethoscope systems and diagnostic platforms together allow healthcare professionals to continuously monitor patients without having to manually assess their breath sounds.

[0160] A. Plastic surgery

[0161] Certain procedures (such as rhinoplasty) will prevent end-tidal CO2 from being used to monitor respiratory status because these patients will not be able to wear the necessary mask. Peripheral pulse oximeters have an inherent delay and can only generate a notification when oxygen desaturation falls below a certain level.

[0162] B. Gastrointestinal endoscopy

[0163] Traditionally, peripheral pulse oximeters and end-tidal CO2 monitors have been used to monitor patients' respiratory status during gastrointestinal endoscopy. Both methods have inherent delays and only provide notification when oxygen desaturation or end-tidal volume changes by at least a certain amount.

[0164] C. Dental Procedures

[0165] During dental procedures, it is difficult for healthcare professionals to continuously monitor vital signs. Not only is a high level of concentration required, but it is not uncommon for only one or two healthcare professionals to be present during a dental procedure. Furthermore, patients are prone to choking or gagging on their own saliva, which is especially risky if these patients are under anesthesia.

[0166] D. Epidemic Prevention

[0167] To prevent the spread of disease, patients are typically isolated. When isolation is necessary, traditional auscultation by healthcare professionals can be dangerous. Furthermore, healthcare professionals may be unable to perform auscultation while wearing personal protective equipment (PPE) such as gowns and face shields. In short, the combination of isolation and PPE can make it difficult to accurately assess a patient's respiratory status.

[0168] Electronic stethoscope systems and diagnostic platforms can be used not only to provide a device for visualizing respiratory events, but also to initiate the playback of auscultation sounds and generate instantaneous notifications. Healthcare professionals can hear the auscultation sounds and see... Figure 16 The respiratory events shown can be used to determine when the airway has been partially obstructed or when the respiratory rate has become unstable.

[0169] Processing system

[0170] Figure 17 This is a block diagram illustrating an example of a processing system 1700 in which at least some of the operations described herein may be implemented. For example, components of the processing system 1700 may be hosted on a computing device performing a diagnostic platform. Examples of computing devices include electronic stethoscope systems, mobile phones, tablet computers, personal computers, and computer servers.

[0171] Processing system 1700 may include processor 1702, main memory 1706, non-volatile memory 1710, network adapter 1712 (e.g., network interface), video display 1718, input / output device 1720, control device 1722 (e.g., keyboard, pointing device, or mechanical input such as buttons), drive unit 1724 (which includes storage medium 1726), and signal generation device 1730 communicatively connected to bus 1716. Bus 1716 is shown as an abstraction representing one or more physical buses and / or point-to-point connections connected by appropriate bridges, adapters, or controllers. Therefore, bus 1716 may include system bus, peripheral component interconnect (PCI) bus, PCI-Express bus, HyperTransport bus, Industry Standard Architecture (ISA) bus, Small Computer System Interface (SCSI) bus, USB, internal integrated circuit (1 2 C) Bus or bus conforming to IEEE Standard 1394.

[0172] Processing system 1700 may share a computer processor architecture similar to that of the following devices: computer servers, routers, desktop computers, tablet computers, mobile phones, video game consoles, wearable electronic devices (e.g., watches or fitness trackers), network-connected (“smart”) devices (e.g., televisions or home assistant devices), augmented reality or virtual reality systems (e.g., head-mounted displays), or another electronic device capable of executing a set of instructions (sequentially or otherwise) that specifies actions to be taken by processing system 1700.

[0173] Although main memory 1706, non-volatile memory 1710, and storage medium 1726 are shown as a single medium, the terms "storage medium" and "machine-readable medium" should be considered to include a single medium or multiple media storing one or more instruction sets 1728. The terms "storage medium" and "machine-readable medium" should also be considered to include any medium capable of storing, encoding, or carrying instruction sets executed by processing system 1700.

[0174] Generally, routines executed to implement embodiments of this disclosure may be implemented as part of an operating system or a particular application, component, program, object, module, or sequence of instructions (collectively, a "computer program"). A computer program typically contains one or more instructions (e.g., instructions 1704, 1708, 1728) set at various times in various memories and storage devices in a computing device. When read and executed by processor 1702, these instructions cause processing system 1700 to perform operations to execute various aspects of this disclosure.

[0175] Although embodiments have been described in the context of a fully functional computing device, those skilled in the art will understand that various embodiments can be distributed as program products in a variety of forms. This disclosure applies regardless of the specific type of machine or computer-readable medium used to actually cause such distribution. Further examples of machine and computer-readable media include recordable media such as volatile memory, non-volatile memory 1710, removable disks, hard disk drives (HDDs), optical discs (e.g., compact disc read-only memory (CD-ROM) and digital versatile optical discs (DVDs)), cloud-based storage, and transport media such as digital and analog communication links.

[0176] Network adapter 1712 enables processing system 1700 to coordinate data in network 1714 with an external entity via any communication protocol supported by processing system 1700 and the entity outside processing system 1700. Network adapter 1712 may include a network adapter card, wireless network interface card, switch, protocol converter, gateway, bridge, hub, receiver, repeater, or transceiver, the transceiver including an integrated circuit (e.g., enabling communication via Bluetooth or Wi-Fi).

[0177] Remark

[0178] For illustrative and descriptive purposes, the foregoing description of various embodiments of the present technology has been provided. This is not intended to be exhaustive or to limit the claimed subject matter to the precise forms disclosed.

[0179] Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to best illustrate the principles of the art and its practical application, thereby enabling others skilled in the art to understand the claimed subject matter, the various embodiments, and the modifications suitable for the particular intended use.

Claims

1. A method for calculating respiratory rate, the method comprising: Obtain recorded audio data representing the sounds produced by the patient's lungs; The audio data is processed to prepare it for analysis by the trained model; The process includes: A high-pass filter is applied to the audio data. Based on the audio data, generate (i) a spectrogram, (ii) a series of Mel-frequency cepstral coefficients (MFCCs), and (iii) a series of values ​​representing the energy summed across different frequency bands of the spectrogram. The spectrum, the series of MFCCs, and the series of values ​​are concatenated into a feature matrix; The trained model is applied to the audio data to generate a vector comprising entries arranged in chronological order. Each entry in the vector indicates whether the corresponding segment of the audio data represents a breathing event; The vector is examined to identify (i) a first breathing event, (ii) a second breathing event following the first breathing event, and (iii) a third breathing event following the second breathing event; Determine (a) a first time period between the first respiratory event and the second respiratory event, and (b) a second time period between the second respiratory event and the third respiratory event; and The respiratory rate is calculated based on the first time period and the second time period.

2. The method of claim 1, wherein each entry in the vector corresponds to a different segment of the audio data, and wherein all entries in the vector correspond to segments of the audio data of equal duration.

3. The method of claim 1, wherein each breathing event corresponds to at least two consecutive entries in the vector, the at least two consecutive entries indicating that the corresponding segment of the audio data represents a breathing event.

4. The method according to claim 1, wherein the processing further comprises: The feature matrix is ​​subjected to minimum-maximum normalization so that each entry has a value between 0 and 1.

5. The method of claim 1, wherein the cutoff frequency of the high-pass filter causes sounds emitted by another internal organ of the patient to be filtered out from the audio data.

6. The method according to claim 1, further comprising: The vector is examined to identify a pair of entries that correspond to the same type of respiratory event and are separated by a predetermined number of entries; as well as The pair of entries are merged to indicate that the pair of entries represents a single breathing event.

7. The method of claim 6, wherein the pair of entries indicates that the corresponding segment of the audio data represents inhalation.

8. The method of claim 6, wherein the pair of entries indicates that the corresponding segment of the audio data represents exhalation.

9. The method according to claim 1, further comprising: The vector is examined to identify a series of consecutive entries corresponding to the same type of respiratory events of less than a predetermined length; as well as Adjust the label associated with each entry in the series of consecutive entries in order to delete the breathing events represented by the series of consecutive entries.

10. The method of claim 1, wherein the first time period extends from the beginning of the first respiratory event to the beginning of the second respiratory event, wherein the second time period extends from the beginning of the second respiratory event to the beginning of the third respiratory event, and wherein the calculation includes: Divide 120 by the sum of the first time period and the second time period to determine the respiratory rate.

11. A non-transitory medium having instructions stored thereon, the instructions, when executed by a processor of a computing device, causing the computing device to perform operations including: Obtain a vector comprising entries arranged in chronological order, which includes: Obtain recorded audio data representing the sounds produced by the patient's lungs; The audio data is processed to prepare it for analysis by the trained model; The processing described therein includes: applying a high-pass filter to the audio data. Based on the audio data, generate (i) a spectrogram, (ii) a series of Mel-frequency cepstral coefficients (MFCCs), and (iii) a series of values ​​representing the energy summed across different frequency bands of the spectrogram. The spectrum, the series of MFCCs, and the series of values ​​are concatenated into a feature matrix; The trained model is applied to the audio data to generate a vector comprising entries arranged in chronological order. Each entry in the vector indicates a prediction of whether the corresponding segment of the audio data represents a respiratory event; The entry in the vector is examined to identify (i) a first breathing event, (ii) a second breathing event following the first breathing event, and (iii) a third breathing event following the second breathing event; Determine (a) a first time period between the first respiratory event and the second respiratory event, and (b) a second time period between the second respiratory event and the third respiratory event; and The respiratory rate is calculated based on the first time period and the second time period.

12. The nontransient medium of claim 11, wherein each respiratory event is an inhalation by the patient.

13. The non-transient medium of claim 11, wherein each of the first breathing event, the second breathing event, and the third breathing event corresponds to a series of consecutive entries in the vector, the series of consecutive entries (i) exceeding a predetermined length and (ii) indicating that the corresponding segment of the audio data represents a breathing event.

14. The non-transient medium according to claim 11, wherein the calculation includes: Divide 120 by the sum of the first time period and the second time period to determine the respiratory rate.

15. The non-transient medium according to claim 11, wherein the operation further comprises: This triggers an interface that displays a spectrogram corresponding to the audio data; The respiratory rate is also published to the interface to facilitate diagnostic determination by healthcare professionals.

16. The non-transient medium according to claim 11, wherein the operation further comprises: The respiratory rate is compared with a lower threshold. and A notification is generated in response to determining that the respiratory rate has dropped below the lower threshold.

17. The non-transient medium according to claim 11, wherein the operation further comprises: The respiratory rate is compared with an upper limit threshold. and A notification is generated in response to determining that the respiratory rate exceeds the upper limit threshold.

Citation Information

Patent Citations

  • Network-connected electronic stethoscope systems

    US10555717B2

  • Vital sign measurement apparatus and body motion detection apparatus

    JP2012105762A

  • Method and apparatus for training and evaluating artificial neural networks used to determine lung pathology

    US20190088367A1