Derivation of health insights through the analysis of audio data generated by a digital stethoscope

The electronic stethoscope system addresses the limitations of acoustic stethoscopes by using conical resonators and microphones to amplify internal sounds and a diagnostic platform to analyze audio data, resulting in improved diagnostic accuracy and real-time monitoring of respiratory events.

JP7694962B2Active Publication Date: 2025-06-18HEROIC FAITH MEDICAL SCI CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022570115
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-05-15
Filing Date
2021-05-14
Publication Date
2025-06-18
Estimated Expiration
2041-05-14

AI Technical Summary

Technical Problem

Acoustic stethoscopes suffer from limitations such as sound attenuation with frequency, leading to weak sound transmission and difficulty in accurate diagnosis, especially with sounds below 50 Hertz not being heard due to differences in ear sensitivity.

Method used

An electronic stethoscope system that includes input units with conical resonators and microphones to capture and amplify internal and external sounds, utilizing a diagnostic platform that employs machine learning and artificial intelligence to analyze audio data, filter out external noise, and accurately identify respiratory events.

Benefits of technology

The electronic stethoscope system effectively amplifies internal sounds while reducing external noise, enabling accurate detection and monitoring of respiratory events, such as respiratory rate, in real-time, thereby improving diagnostic accuracy and patient care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694962000001
    Figure 0007694962000001
  • Figure 0007694962000002
    Figure 0007694962000002
  • Figure 0007694962000003
    Figure 0007694962000003
Patent Text Reader

Abstract

This application presents a computer program and related computer-implemented technology for deriving patient health insights through analysis of audio data generated by an electronic stethoscope system. A diagnostic platform may be responsible for examining the audio data generated by the electronic stethoscope system to derive patient health insights. The diagnostic platform may employ heuristics, algorithms, or models that rely on machine learning or artificial intelligence to perform auscultation in a manner significantly superior to traditional approaches that rely on visual analysis by medical professionals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various embodiments relate to a computer program for deriving insights regarding a patient's health through analysis of audio data generated by an electronic stethoscope system, and related computer-implemented techniques.

Background Art

[0002] Historically, acoustic stethoscopes have been used to listen to internal sounds generated within a living body. This process is referred to as "auscultation" and is typically performed for the purpose of examining a biological system, from which performance can be inferred from the internal sounds of those biological systems. Typically, an acoustic stethoscope includes a single chest piece with a resonator designed to be placed against the body, and a pair of hollow tubes connected to earpieces. When sound waves are captured by the resonator, the sound waves are directed towards the earpieces through the pair of hollow tubes.

[0003] However, acoustic stethoscopes suffer from several drawbacks. For example, acoustic stethoscopes are designed to attenuate sound in proportion to the frequency of the sound source. As a result, the sound transmitted to the earpiece tends to be very weak, which may make it difficult to accurately diagnose the condition. In fact, due to differences in ear sensitivity, some sounds (e.g., sounds below 50 Hertz) may not be heard at all.

[0004] Some companies have begun developing electronic stethoscopes (also referred to as "stethophones") to address the drawbacks of acoustic stethoscopes. Electronic stethoscopes improve upon acoustic stethoscopes by electronically amplifying sound. For example, an electronic stethoscope can address faint sounds generated within a living body by amplifying these sounds. To achieve this, an electronic stethoscope converts sound waves detected by a microphone in the chest piece into an electrical signal and amplifies this electrical signal for optimal listening.

Brief Description of the Drawings

[0005]

Figure 1A

Figure 1B

Figure 1C

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 12

Figure 13A

Figure 13B

Figure 14

Figure 15

Figure 16

Figure 17

[0006] Embodiments are shown in these drawings by way of example and not limitation. These drawings represent various embodiments for purposes of illustration, and those skilled in the art will recognize that alternative embodiments can be adopted without departing from the principles of the technology. Thus, while specific embodiments are shown in the drawings, the technology is amenable to various modifications. SUMMARY OF THE INVENTION

[0007] As further discussed below, an electronic stethoscope system can be designed to simultaneously monitor sounds generated from within and outside the body during an examination. As further discussed below, the electronic stethoscope system may include one or more input units connected to a hub unit. Each input unit may have a conical resonator designed to direct sound waves towards at least one microphone configured to generate audio data indicative of internal sounds generated within the body. These microphones may be referred to as "auscultation microphones". Further, each input unit may include at least one microphone configured to generate audio data indicative of external sounds generated outside the body. These microphones may be referred to as "ambient microphones" or "environmental microphones".

[0008] For purposes of explanation, "ambient microphones" may be described as being capable of generating audio data indicative of "ambient sounds". However, these "ambient sounds" generally include a combination of external sounds generated from three different sound sources: (1) sounds generated from the ambient environment, (2) sounds leaking from the input portion, and (3) sounds penetrating the body being examined. Examples of external sounds include sounds directly generated from the input unit (e.g., sounds of rubbing with a finger or chest), and low-frequency environmental noise penetrating the input unit.

[0009] There are several advantages to recording internal and external sounds separately. In particular, the internal sounds can be electronically amplified while the external sounds can be electronically dampened, attenuated, or filtered. Thus, the electronic stethoscope system can address faint sounds generated from within the body being examined by manipulating the audio data indicative of the internal and external sounds. However, unwanted digital artifacts may be generated by the manipulation, making the interpretation of the internal sounds difficult. For example, these digital artifacts may make it more difficult to identify the pattern of values of the audio data indicative of inhalation or exhalation generated by the auscultation microphones.

[0010] In this application, a computer program for deriving insights regarding a patient's health through the analysis of audio data generated by an electronic stethoscope system, and related computer-implemented techniques, are presented. A diagnostic platform (also referred to as a "diagnostic program" or "diagnostic application") can play a role in examining audio data generated by an electronic stethoscope system to obtain insights regarding a patient's health. As further discussed below, the diagnostic platform may employ heuristics, algorithms, or models that rely on machine learning (ML) or artificial intelligence (AI), thereby enabling auscultation to be performed in a manner that is significantly superior to conventional approaches that rely on visual analysis by medical experts.

Best Mode for Carrying Out the Invention

[0011] For example, assume that a diagnostic platform is tasked with estimating a patient's respiratory rate based on audio data generated by an electronic stethoscope system connected to the patient. The term "respiratory rate" refers to the number of breaths taken per minute. In the case of an adult at rest, a respiratory rate of 12 to 25 breaths per minute is considered normal. If the respiratory rate at rest is less than 12 breaths per minute or exceeds 25 breaths per minute, that respiratory rate is considered abnormal. Since the respiratory rate is an important indicator of health, it may be referred to as a "vital sign". Therefore, knowledge of the respiratory rate can be important for providing appropriate care to the patient.

[0012] As described above, the electronic stethoscope system can generate voice data indicating internal sounds and voice data indicating external sounds. The former may be referred to as "first voice data" or "internal voice data", and the latter may be referred to as "second voice data" or "external voice data". In some embodiments, the diagnostic platform utilizes the second voice data to improve the first voice data. Thus, the diagnostic platform can examine the second voice data and, if it does operate, determine how to operate the first voice data to reduce the influence of external sounds. For example, if the diagnostic platform discovers external sounds through the analysis of the second voice data, the diagnostic platform may obtain the first voice data and apply a filter to the first voice data in order to remove the external sounds without distorting the target internal sounds (such as corresponding to inhalation, exhalation, etc.). The filter may be generated by the diagnostic platform based on the analysis of the second voice data or may be identified by the diagnostic platform (e.g., from among a plurality of filters) based on the analysis of the second voice data.

[0013] Next, the diagnostic platform may apply a computer-implemented model (or simply a "model") to the first audio data. The model may be designed and trained to identify distinct phases in breathing, namely inhalation and exhalation. The model may be a deep learning model based on one or more artificial neural networks (or simply "neural networks"). A neural network is a framework of multiple ML algorithms that cooperate to process complex inputs. A neural network is inspired by the biological neural networks that make up the human or animal brain and can "learn" to perform a task by examining examples without being programmed by task-specific rules. For example, a neural network may learn to distinguish between inhalation and exhalation by examining a series of audio data labeled as "inhalation", "exhalation", or "no inhalation or exhalation". This series of audio data may be referred to as "training data". With this approach to training, the neural network can automatically learn from the series of audio data the features that indicate inhalation and exhalation by the presence of those features. For example, the neural network may come to understand the pattern of values of the audio data indicating exhalation, as well as the pattern of values of the audio data indicating inhalation.

[0014] The output generated by a model can be useful, but on the other hand, it may be difficult to interpret. For example, a diagnostic platform may generate a visual representation (also referred to as a "visualization") of first voice data for review by a medical professional responsible for monitoring, examining, or diagnosing a patient. The diagnostic platform may visually highlight inhalation or exhalation to improve the clarity of the visualization. For example, the diagnostic platform may overlay a digital element (also referred to as a "graphic element") on the visualization to indicate that the patient has inhaled. As further discussed below, the digital element can extend from the start of inhalation determined by the model to the end of exhalation determined by the model. By explaining the output generated by the model in a more understandable way, the diagnostic platform can build trust with medical professionals who rely on those outputs.

[0015] Furthermore, the diagnostic platform may utilize the output generated by the model to gain insights into the patient's health. For example, assume that the model is applied to voice data associated with a patient to identify inhalation and exhalation as described above. In such a situation, the diagnostic platform may generate a metric using those outputs. An example of such a metric is the respiratory rate. As noted above, the term "respiratory rate" refers to the number of breaths taken per minute. However, it is not practical to count the actual respiratory rate in the immediately preceding minute. Even over a limited period of time, significant irreversible damage can occur if the patient suffers from oxygen deprivation. Therefore, it is important for medical professionals to obtain consistent insights into the respiratory rate that is updated almost in real-time. To achieve this, the diagnostic platform may continuously calculate the respiratory rate using a sliding window algorithm. Broadly speaking, the algorithm defines a window that includes a portion of the voice data and then continuously slides the window to cover different portions of the voice data. As further discussed below, the algorithm may adjust the window so that a predetermined number of inhalations or exhalations are always included within its boundaries. For example, the algorithm may be programmed so that three inhalations are included within the boundaries of the window. At this time, the first inhalation represents the "start" of the window, and the third inhalation represents the "end" of the window. Subsequently, the diagnostic platform can calculate the respiratory rate based on the intervals between the inhalations included in the window.

[0016] For illustrative purposes, embodiments may be described in the context of instructions executable by a computing device. However, aspects of the present technology can be implemented via hardware, firmware, software, or any combination thereof. As an example, the model may be applied to voice data generated by an electronic stethoscope system to identify outputs representing the patient's inhalation and exhalation. And an algorithm may be applied to the output generated by the model to generate a metric indicating the patient's respiratory rate.

[0017] Embodiments may be described with reference to particular computing devices, networks, and medical professionals and facilities. However, one of ordinary skill in the art will recognize that the features of these embodiments are equally applicable to other computing devices, networks, and medical professionals and facilities. For example, embodiments may be described in the context of deep neural networks, but the models employed by the diagnostic platform may be based on other deep learning architectures such as deep belief networks, recurrent neural networks, and convolutional neural networks.

[0018] Term A brief definition of the terms, abbreviations, and phrases used throughout this application is provided below.

[0019] The terms "connected," "coupled," and any variations thereof are intended to include any connection or coupling, direct or indirect, between two or more elements. The connection or coupling can be physical, logical, or a combination thereof. For example, objects may be electrically or communicatively connected to each other even if they do not share a physical connection.

[0020] The term "module" may be used to broadly denote a component implemented via software, firmware, hardware, or any combination thereof. Generally, a module is a functional component that generates one or more outputs based on one or more inputs. A computer program may include one or more modules. Thus, a computer program may include multiple modules responsible for completing various different tasks, or may include a single module responsible for completing all tasks.

[0021] Overview of the Electronic Stethoscope System FIG. 1A includes a top perspective view of an input unit 100 for an electronic stethoscope system. For convenience, the input unit 100 may be referred to as a “stethoscope patch,” although it may include only a subset of the components necessary for auscultation. Since the input unit 100 is often attached to the chest of the body, it may also be referred to as a “chest piece.” However, those skilled in the art will recognize that the input unit 100 can be similarly attached to other parts of the body (e.g., the neck, abdomen, or back).

[0022] As further discussed below, the input unit 100 can collect sound waves representing biological activity within the body being examined, convert the sound waves into electrical signals, and then digitize the electrical signals (e.g., for purposes such as easier transmission, ensuring higher fidelity, etc.). The input unit 100 may include a structure 102 composed of a rigid material. Typically, the structure 102 is composed of a metal such as stainless steel, aluminum, titanium, or a suitable metal alloy. To fabricate the structure 102, typically, molten metal is die-cast and then machined or extruded into the appropriate shape.

[0023] In some embodiments, the input unit 100 includes a casing that prevents the structure 102 from being exposed to the surrounding environment. For example, the casing can prevent contamination and improve cleanability. Generally, the casing encloses substantially all of the structure 102 except for a conical resonator disposed along its bottom surface. The conical resonator will be described in more detail below with respect to FIGS. 1B - C. The casing can be composed of silicone rubber, polypropylene, polyethylene, or any other suitable material. Further, in some embodiments, the casing includes additives that suppress the growth of microorganisms, ultraviolet (UV) degradation, etc. due to the presence of the additives.

[0024] Figures 1B - C include a bottom perspective view of the input unit 100, and the input unit 100 includes a structure 102 having a distal portion 104 and a proximal portion 106. To initiate a auscultation procedure, a person (e.g., a medical professional such as a doctor or a nurse) can fix the proximal portion 106 of the input unit 100 to the surface of the body being examined. The proximal portion 106 of the input unit 100 may include the wide opening 108 of the conical resonator 110. The conical resonator 110 may be designed to direct sound waves collected through the wide opening 108 towards the narrow opening 112, and these sound waves can be led to the auscultation microphone. Conventionally, the wide opening 108 is about 30 - 50 millimeters (mm), 35 - 45 mm, or 38 - 40 mm. However, since the input unit 100 described herein can have an automatic gain control function, a smaller conical resonator may be used. For example, in some embodiments, the wide opening 108 is less than 30 mm, 20 mm, or 10 mm. Thus, the input unit described herein may be capable of supporting a wide variety of conical resonators, such as conical resonators of various different sizes and conical resonators designed for various different applications.

[0025] Regarding the terms "distal" and "proximal", unless otherwise specified, these terms indicate the relative position of the input unit 100 with respect to the body. For example, when referring to the input unit 100 suitable for fixation to the body, "distal" can indicate a first position close to where a cable suitable for transmitting a digital signal can be connected to the input unit 100, and "proximal" can indicate a second position close to where the input unit 100 contacts the body.

[0026] FIG. 2 includes a cross-sectional side view of an input unit 200 for an electronic stethoscope system. Often, the input unit 200 includes a structure 202 having an internal cavity defined within its structure 202. The structure 202 of the input unit 200 may have a conical resonator 204 designed to direct sound waves towards a microphone present within the internal cavity. In some embodiments, a diaphragm 212 (also referred to as a “vibrating membrane”) extends across the wide opening (also referred to as the “outer opening”) of the conical resonator 204. The diaphragm 212 can often be used to listen for high-pitched sounds such as those generated by the lungs. The diaphragm 212 can be formed from a thin plastic disk composed of an epoxy fiberglass compound or glass fibers.

[0027] To improve the clarity of the sound waves collected by the conical resonator 204, the input unit 200 can be designed to simultaneously monitor sounds generated from various different positions. For example, the input unit 200 can be designed to simultaneously monitor sounds generated from within the body being examined and sounds generated from the surrounding environment. Thus, the input unit 200 can include at least one microphone 206 (referred to as a “stethoscope microphone”) configured to generate audio data indicative of internal sounds and at least one microphone 208 (referred to as a “surround microphone”) configured to generate audio data indicative of ambient sounds. Each stethoscope and surround microphone can include a transducer capable of converting sound waves into electrical signals. Thereafter, the electrical signals generated by the stethoscope and surround microphones 206, 208 can be digitized before being transmitted to a hub unit. By digitization, the hub unit can easily clock or synchronize signals received from multiple input units. Digitization can also ensure that signals received by the hub unit from the input unit have a higher fidelity than may be possible by other means.

[0028] These microphones may be omnidirectional microphones designed to pick up sound from all directions, or may be directional microphones designed to pick up sound coming from a specific direction. For example, the input unit 200 may include one or more stethoscope microphones 206 directed to pick up sound generated from the space adjacent to the outer opening of the conical resonator 204. In such an embodiment, the one or more ambient microphones 208 may be omnidirectional microphones or may be directional microphones. As another example, a set of ambient microphones 208 can be evenly spaced within the structure 202 of the input unit 200 to form a phased array capable of capturing ambient sound that is highly directional to reduce noise and interference. Thereby, the one or more stethoscope microphones 206 can be arranged to focus on the path of the incoming internal sound (also referred to as the "stethoscope path"), and the one or more ambient microphones 208 can be arranged to focus on the path of the incoming ambient sound (also referred to as the "ambient path").

[0029] Conventionally, an electronic stethoscope has applied an electrical signal indicating a sound wave to a digital signal processing (DSP) algorithm that serves to filter out unwanted artifacts. However, such an action may suppress almost all sounds within a specific frequency range (e.g., 100 - 800 Hz), thereby significantly distorting the targeted internal sounds (e.g., sounds corresponding to inhalation, exhalation, or heartbeat). However, herein, the processor can use an active noise cancellation algorithm that separately examines the audio data generated by the auscultation microphone 206 and the audio data generated by the ambient microphone 208. More specifically, the processor can analyze the audio data generated by the ambient microphone 208 and thereby determine how the audio data generated by the auscultation microphone 206 should be modified. For example, the processor can detect that a particular digital feature should be amplified (e.g., because the feature corresponds to an internal sound), attenuated (e.g., because the feature corresponds to ambient sound), or completely removed (e.g., because the feature represents noise). Such techniques can be used to improve the clarity, detail, and quality of the sound recorded by the input unit 200. For example, the application of a noise cancellation algorithm can be an integral part of the noise removal process employed by an electronic stethoscope system including at least one input unit 200.

[0030] For privacy purposes, while the conical resonator 204 is oriented away from the body, neither the (one or more) auscultation microphones 206 nor the (one or more) ambient microphones 208 may be permitted to record. Thus, in some embodiments, the (one or more) auscultation microphones 206 and / or the (one or more) ambient microphones 208 do not start recording until the input unit 200 is attached to the body. In such embodiments, the input unit 200 may include one or more attachment sensors 210a - c that serve to determine whether the structure 202 is properly fixed to the body surface.

[0031] The input unit 200 can include any subset of the attachment sensors shown herein. For example, in some embodiments, the input unit 200 includes only the attachment sensors 210a - b disposed near the wide opening of the conical resonator 204. As another example, in some embodiments, the input unit 200 includes only the attachment sensor 210c disposed near the narrow opening (also referred to as the "inner opening") of the conical resonator 204. Further, the input unit 200 can include various different types of attachment sensors. For example, the attachment sensor 210c can be an optical proximity sensor designed to emit light (e.g., infrared light) through the conical resonator 204 and determine the distance between the input unit 200 and the structure based on the light that is reflected and returned into the conical resonator 204. As another example, the attachment sensors 210a - c can be acoustic sensors designed to determine whether the structure 202 is securely sealed against the body surface based on the presence of ambient noise (also referred to as "environmental noise") with the aid of an algorithm programmed to determine the attenuation of a high - frequency signal. As another example, the attachment sensors 210a - b can be pressure sensors designed to determine whether the structure 202 is securely sealed against the body surface based on the amount of pressure applied. Some embodiments of the input unit 200 include each of these various different types of attachment sensors. By considering the outputs of these attachment sensors 210a - c in combination with the aforementioned active noise cancellation algorithm, the processor may be able to dynamically determine the attachment state. That is, the processor may be able to determine whether the input unit 200 is forming a seal against the body based on the outputs of these attachment sensors 210a - c.

[0032] FIG. 3 shows how one or more input units 302a - n are connected to the hub unit 304 to form the electronic stethoscope system 300. In some embodiments, multiple input units are connected to the hub unit 304. For example, the electronic stethoscope system 300 may include four input units, six input units, or eight input units. Generally, the electronic stethoscope system 300 will include at least six input units. An electronic stethoscope system with multiple input units may be referred to as a "multi-channel stethoscope." In other embodiments, only one input unit is connected to the hub unit 304. For example, a single input unit may be moved across the body in such a way as to simulate an array of multiple input units. An electronic stethoscope system with one input unit may be referred to as a "single-channel stethoscope."

[0033] As shown in FIG. 3, each input unit 302a - n can be connected to the hub unit 304 via a corresponding cable 306a - n. Generally, the transmission paths formed via the corresponding cables 306a - n between each input unit 302a - n and the hub unit 304 are designed to be substantially interference - free. For example, an electronic signal may be digitized by the input units 302a - n before transmission to the hub unit 304, and the signal fidelity may be ensured by preventing the generation / contamination of electromagnetic noise. Examples of cables include ribbon cables, coaxial cables, Universal Serial Bus (USB) cables, High - Definition Multimedia Interface (HDMI) cables, RJ45 Ethernet cables, and other cables suitable for the transmission of digital signals. Each cable includes a first end connected to the hub unit 304 (e.g., via a physical port) and a second end connected to the corresponding input unit (e.g., via a physical port). Thus, each input unit 302a - n may include a single physical port, and the hub unit 304 may include multiple physical ports. Alternatively, a single cable may be used to connect all of the input units 302a - n to the hub unit 304. In such an embodiment, the cable can include a first end that can interface with the hub unit 304 and a series of second ends each of which can interface with a single input unit. Such a cable may be referred to as a "1 - to - 2 cable", "1 - to - 4 cable", or "1 - to - 6 cable", for example, based on the number of second ends.

[0034] When all of the input units 302a - n connected to the hub unit 304 are in the auscultation mode, the electronic stethoscope system 300 can employ an adaptive gain control algorithm programmed to compare internal sounds with ambient sounds. The adaptive gain control algorithm can analyze the auscultatory sounds of interest (e.g., normal breath sounds, wheezing, crepitation, crackling, etc.) to determine whether an appropriate sound level has been achieved. For example, the adaptive gain control algorithm can determine whether the sound level exceeds a predetermined threshold. The adaptive gain control algorithm can be designed to achieve gain control up to 100 times, for example, in two different stages. The gain level can be adaptively adjusted based on the number of input units within the input unit array 308 and the level of the sound recorded by the auscultation microphones within each input unit. In some embodiments, the adaptive gain control algorithm is programmed to be implemented as part of a feedback loop. Thus, the adaptive gain control algorithm can apply a gain to the sound recorded by the input unit, determine whether the sound exceeds a pre - programmed intensity threshold, and dynamically determine whether additional gain is required based on that determination.

[0035] Since the electronic stethoscope system 300 can implement the adaptive gain control algorithm during post - processing procedures, the input unit array 308 can be enabled to collect information regarding a wide range of sounds generated by the heart, lungs, etc. The input units 302a - n within the input unit array 308 can be arranged at various different anatomical positions along the body surface (or on entirely different bodies), so that various different biometric characteristics (e.g., respiratory rate, heart rate, or the degree of wheezing, crepitation, crackling, etc.) can be simultaneously monitored by the electronic stethoscope system 300.

[0036] FIG. 4 is a high-level block diagram showing exemplary components of an input unit 400 and a hub unit 450 of an electronic stethoscope system. Embodiments of the input unit 400 and the hub unit 450 can include any subset of the components shown in FIG. 4, as well as additional components not shown here. For example, the input unit 400 may include a biometric sensor capable of monitoring biometric characteristics of the body such as sweating, temperature, etc. (e.g., based on skin humidity). Additionally or alternatively, the biometric sensor may be designed to monitor a breathing pattern (also referred to as a "respiratory pattern"), record the electrical activity of the heart, etc. As another example, the input unit 400 may include an inertial measurement unit (IMU) capable of generating data from which gestures, orientations, or positions can be derived. An IMU is an electronic component designed to measure the force, angular velocity, inclination, and / or magnetic field of an object. Generally, an IMU includes an accelerometer, a gyroscope, a magnetometer, or any combination thereof.

[0037] The input unit 400 can include one or more processors 404, a wireless transceiver 406, one or more microphones 408, one or more attachment sensors 410, a memory 412, and / or a power component 414 electrically coupled to a power interface 416. These components can be present within a housing 402 (also referred to as a "structure").

[0038] As described above, the microphone 408 can convert acoustic sound waves into electrical signals. The microphone 408 can include a stethoscope microphone configured to generate voice data indicative of internal sounds, an ambient microphone configured to generate voice data indicative of ambient sounds, or any combination thereof. The voice data representing the value of the electrical signal can be stored, at least temporarily, in the memory 412. In some embodiments, the (one or more) processors 404 process the voice data before downstream transmission towards the hub unit 450. For example, the (one or more) processors 404 can apply algorithms designed for digital signal processing, noise removal, gain control, noise cancellation, artifact removal, feature identification, etc. In other embodiments, minimal processing is pre-executed by the (one or more) processors 404 before downstream transmission towards the hub unit 450. For example, the (one or more) processors 404 may simply add metadata identifying the identity of the input unit 400 to the voice data, or may inspect the metadata already added to the voice data by the microphone 408.

[0039] In some embodiments, the input unit 400 and the hub unit 450 transmit data between each other via a cable connected between the corresponding data interfaces 418, 470. For example, the voice data generated by the (one or more) microphones 408 can be transferred to the data interface 418 of the input unit 400 for transmission to the data interface 470 of the hub unit 450. Alternatively, the data interface 470 may be part of the wireless transceiver 456. The wireless transceiver 406 can be configured to automatically establish a wireless connection with the wireless transceiver 456 of the hub unit 450. The wireless transceivers 406, 456 can communicate with each other via a two-way communication protocol such as near field communication (NFC), wireless USB, Bluetooth®, Wi-Fi®, cellular data protocols (e.g., LTE, 3G, 4G or 5G), or a proprietary point-to-point protocol.

[0040] The input unit 400 may include a power component 414 that can supply power to other components present within the housing 402 as needed. Similarly, the hub unit 450 can include a power component 466 that can supply power to other components present within the housing 452. Examples of power components include rechargeable lithium-ion (Li-Ion) batteries, rechargeable nickel-metal hydride (NiMH) batteries, rechargeable nickel-cadmium (NiCad) batteries, and the like. The input unit 400 does not include a dedicated power component and thus needs to receive power from the hub unit 450. A cable designed to facilitate the transmission of power (e.g., via physical connection of electrical contacts) can be connected between the power interface 416 of the input unit 400 and the power interface 468 of the hub unit 450.

[0041] The power channel (i.e., the channel between the power interface 416 and the power interface 468) and the data channel (i.e., the channel between the data interface 418 and the data interface 470) are shown as separate channels, but this is for illustrative purposes only. Those skilled in the art will recognize that these channels can be included within the same cable. Thus, a single cable capable of carrying data and power may be coupled between the input unit 400 and the hub unit 450.

[0042] The hub unit 450 can include one or more processors 454, a wireless transceiver 456, a display 458, a codec 460, one or more light-emitting diode (LED) indicators 462, a memory 464, and a power component 466. These components can be present within the housing 452 (also referred to as the "structure"). As described above, embodiments of the hub unit 450 can include any subset of these components or additional components not shown herein.

[0043] As shown in FIG. 4, an embodiment of the hub unit 450 may include a display 458 for presenting information such as the breathing state or heart rate of the individual being examined, the network connection state, the power connection state, the connection state for the input unit 400, and the like. The display 458 may be controlled via a tactile input mechanism (e.g., a button accessible along the surface of the housing 452), an audio input mechanism (e.g., a microphone), and the like. As another example, some embodiments of the hub unit 450 include one or more LED indicators 462 for operating guidance instead of the display 458. In such embodiments, the one or more LED indicators 462 may convey the same information as that presented by the display 458. As another example, some embodiments of the hub unit 450 include the display 458 and the one or more LED indicators 462.

[0044] Upon receiving audio data representing an electrical signal generated by the microphone 408 of the input unit 400, the hub unit 450 can provide the audio data to a codec 460 responsible for decoding the incoming data. The codec 460 can decode the audio data (e.g., by reversing the encoding applied by the input unit 400) for purposes such as editing, processing, and the like. It includes the one or more stethoscope microphones within the input unit 400 and the audio data generated by the one or more ambient microphones within the input unit 400.

[0045] Thereafter, one or more processors 454 can process the voice data. Similar to the one or more processors 404 of the input unit 400, the one or more processors 454 of the hub unit 450 can apply algorithms designed for digital signal processing, noise removal, gain control, noise cancellation, artifact removal, feature identification, etc. Some of these algorithms may not be necessary if they have already been applied by the one or more processors 404 of the input unit 400. For example, in some embodiments, the one or more processors 454 of the hub unit 450 apply one or more algorithms to discover features related to diagnosis in the voice data, while in other embodiments, such actions may not be necessary if the one or more processors 404 of the input unit 400 have already discovered features related to diagnosis. Alternatively, the hub unit 450 can transfer the voice data to a destination (e.g., a diagnostic platform operating on a computing device or a distributed system) for analysis, as further discussed below. Generally, features related to diagnosis will correspond to patterns of values in the voice data that match predetermined pattern definition parameters. As another example, in some embodiments, the processor 454 of the hub unit 450 applies an algorithm to reduce the noise of the voice data to improve the signal-to-noise (SNR) ratio, while in other embodiments, these algorithms are applied by the processor 404 of the input unit 400.

[0046] In addition to the power interface 468, the hub unit 450 may include a power port. The power port (also referred to as a "power jack") enables the hub unit 450 to be physically connected to a power source (e.g., an electrical outlet). The power port may be connectable to various different connector types (e.g., C13, C15, C19, etc.). Additionally or alternatively, the hub unit 450 may include a power receiver having an integrated circuit (also referred to as a "chip") capable of receiving power wirelessly from an external source. Similarly, the input unit 400 may include a power receiver having a chip capable of receiving power wirelessly from an external source, e.g., when the input unit 400 and the hub unit 450 are not physically connected to each other via a cable. The power receiver may be configured to receive power transmitted according to the Qi standard developed by the Wireless Power Consortium, or some other wireless power standard.

[0047] In some embodiments, the housing 452 of the hub unit 450 includes an audio port. The audio port (also referred to as an "audio jack") is a receptacle that can be used to transmit signals such as audio to a suitable plug of an accessory such as headphones. The audio port typically includes 1, 2, 3, or 4 contacts that facilitate the transmission of an audio signal when a suitable plug is inserted into the audio port. For example, most headphones include a plug designed for a 3.5 millimeter (mm) audio port. Additionally or alternatively, the wireless transceiver 456 of the hub unit 450 may be capable of directly transmitting an audio signal to wireless headphones (e.g., via NFC, Wireless USB, Bluetooth, etc.).

[0048] As described above, the processor 404 of the input unit 400 and / or the processor 454 of the hub unit 450 can apply various algorithms to support various different functions. Examples of such functions include attenuation of lost data packets in audio data, volume control depending on noise, dynamic range compression, automatic gain control, equalization, noise suppression, acoustic echo cancellation, and the like. Each function may correspond to a separate module resident in a memory (e.g., the memory 412 of the input unit 400 or the memory 464 of the hub unit 450). Thus, the input unit 400 and / or the hub unit 450 may include an attenuation module, a volume control module, a compression module, a gain control module, an equalization module, a noise suppression module, an echo cancellation module, or any combination thereof.

[0049] Note that in some embodiments, the input unit 400 is configured to directly transmit the audio data generated by the microphone 408 to destinations other than the hub unit 450. For example, the input unit 400 may transfer the audio data to the wireless transceiver 406 for transmission to a diagnostic platform responsible for analyzing the audio data. The audio data may be transmitted to the diagnostic platform instead of, or in addition to, the hub unit 450. When the audio data is transferred to the diagnostic platform in addition to the hub unit 450, the input unit 400 may generate duplicate copies of the audio data and then transfer those individual copies thereto (e.g., to the wireless transceiver 406 for transmission to the diagnostic platform and to the data interface 418 for transmission to the hub unit 450). As further discussed below, the diagnostic platform typically resides on a computing device communicatively connected to the input unit 400, but the form of the diagnostic platform may be present on the input unit 400 or the hub unit 450.

[0050] Additional information regarding the electronic stethoscope system can be found in U.S. Patent No. 10,555,717, which is hereby incorporated by reference in its entirety.

[0051] Overview of the Diagnostic Platform FIG. 5 shows a network environment 500 including a diagnostic platform 502. An individual can interact with the diagnostic platform 502 via an interface 504. For example, a patient may be able to access an interface where information regarding voice data, illness, treatment, and feedback can be provided. As another example, a medical professional may be able to access an interface where voice data and the analysis of voice data can be reviewed for purposes such as making appropriate diagnostic decisions, health monitoring, etc. Generally speaking, the interface 504 can be intended to function as a valuable dashboard for either the patient or the medical professional.

[0052] As shown in FIG. 5, the diagnostic platform 502 can exist within the network environment 500. Accordingly, the diagnostic platform 502 can be connected to one or more networks 506a - b. The networks 506a - b can include a Personal Area Network (PAN), Local Area Network (LAN), Wide Area Network (WAN), Metropolitan Area Network (MAN), cellular network, the Internet, etc. Additionally or alternatively, the diagnostic platform 502 can be communicatively coupled to one or more computing devices via short - range wireless connection technologies such as Bluetooth, NFC, Wi - Fi Direct (also referred to as "Wi - Fi P2P").

[0053] The interface 504 may be accessible via a web browser, a desktop application, a mobile application, or an over-the-top (OTT) application. For example, a medical professional may have access to an interface where information about a patient can be entered. Such information may include a name, date of birth, diagnosis, symptoms, or medications. Alternatively, the information may be automatically entered into the interface by the diagnostic platform 502 (e.g., based on data stored within a network-accessible server system 508), while the medical professional may be permitted to adjust the information as needed. As further discussed below, a medical professional may have access to an interface where voice data and an analysis of the voice data can be presented for review. This information may enable the medical professional to readily establish a patient's condition (e.g., whether the patient is breathing or not breathing), make a diagnosis (e.g., based on the presence of wheezing, crepitation, crackling, etc.), and the like. Accordingly, the interface 504 may be viewable on a computing device such as a mobile workstation (also referred to as a "medical cart"), a personal computer, a tablet computer, a mobile phone, a wearable electronic device, and a virtual or augmented reality system.

[0054] In some embodiments, at least some components of the diagnostic platform 502 are hosted locally. That is, a portion of the diagnostic platform 502 may be present on a computing device used to access one of the interfaces 504. For example, the diagnostic platform 502 may be implemented as a mobile application that runs on a mobile phone associated with a medical professional. However, note that the mobile application may be communicatively connected to a network-accessible server system 508 where other components of the diagnostic platform 502 are hosted.

[0055] In other embodiments, the diagnostic platform 502 is executed entirely by, for example, cloud computing services operated by Amazon Web Service (registered trademark), Google Cloud Platform (registered trademark), or Microsoft Azure (registered trademark). In such embodiments, the diagnostic platform 502 may reside on a network-accessible server system 508 comprising one or more computer servers. These computer servers may include models, algorithms (e.g., for processing voice data, calculating respiratory rate, etc.), patient information (e.g., profiles, credential information, and health-related information such as age, date of birth, geographical location, disease classification, medical condition, healthcare provider, etc.), and other assets. Those skilled in the art will understand that such information may also be distributed between the network-accessible server system 508 and one or more computing devices, and may also be distributed across a distributed network infrastructure such as a blockchain.

[0056] FIG. 6 shows an example of a computing device 600 that can implement a diagnostic platform 610 designed to generate outputs useful for detecting, diagnosing, and monitoring changes in a patient's health state. As further discussed below, the diagnostic platform 610 can apply a model to voice data associated with a patient to identify the occurrence of respiratory events (also referred to as "breathing events" or "breaths"), and subsequently apply an algorithm to those respiratory events to understand the patient's health state. The terms "respiratory event" and "breathing event" may be used to indicate an inhalation or exhalation. For example, the algorithm can output a metric representing the respiratory rate, as further discussed below. Thus, the diagnostic platform 610 can not only discover patterns in the values of voice data relevant to diagnosis, but also generate visualizations of those patterns in a way that helps understand the patient's current health state. For illustrative purposes, embodiments can be described in the context of voice data generated by an electronic stethoscope system. However, those skilled in the art will recognize that the voice data can be obtained from other sources.

[0057] Typically, the computing device 600 is associated with a medical professional or a healthcare facility. For example, the computing device 600 may be a mobile workstation placed in a hospital operating room, or a mobile phone or tablet computer accessible to a medical professional while providing services to a patient. Alternatively, the computing device 600 may be a hub unit of an electronic stethoscope system.

[0058] Computing device 600 can include a processor 602, a memory 604, a display mechanism 606, and a communication module 608. Each of these components will be discussed in more detail below. Those skilled in the art will recognize that there can be various different combinations of these components depending on the nature of the computing device 600.

[0059] Processor 602 can have general characteristics similar to a general-purpose processor or can be an application-specific integrated circuit (ASIC) that provides control functions for the computing device 600. As shown in FIG. 6, processor 602 can be directly or indirectly connected for communication to all components of the computing device 600.

[0060] Memory 604 can be composed of any suitable type of storage medium such as static random access memory (SRAM), dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, or registers. In addition to storing instructions executable by processor 602, memory 604 can also store data generated by processor 602 (e.g., when executing modules of the diagnostic platform 610). Note that memory 604 is merely an abstract representation of the storage environment. Memory 604 can be composed of actual memory chips or modules.

[0061] The display mechanism 606 can be any component operable to visually communicate information. For example, the display mechanism 606 can be a panel including LEDs, organic LEDs, liquid crystal elements, or electrophoretic elements. In an embodiment where the computing device 600 represents the hub unit of the electronic stethoscope system, the display mechanism 606 can be a display panel (e.g., the display 458 in FIG. 4), or an LED display (e.g., the LED display 462 in FIG. 4). In some embodiments, the display mechanism 606 is touch-sensitive. Thus, an individual may be able to provide input to the diagnostic platform 610 by interacting with the display mechanism 606. In embodiments where the display mechanism 606 is not touch-sensitive, an individual may be able to interact with the diagnostic platform 610 using a control device (not shown) such as a keyboard, a physical element (e.g., a mechanical button or knob), or a pointing device (e.g., a computer mouse).

[0062] The communication module 608 may be responsible for managing communication between components of the computing device 600, and may also be responsible for managing communication with other computing devices (e.g., the network-accessible server system 508 in FIG. 5). The communication module 608 can be a wireless communication circuit designed to establish a communication channel with other computing devices. Examples of wireless communication circuits include an antenna module configured for a cellular network (also referred to as a "mobile network"), and chips configured for NFC, wireless USB, Bluetooth, etc.

[0063] For the sake of cost, the diagnostic platform 610 may be referred to as a computer program that resides in the memory 604 and is executed by the processor 602. However, the diagnostic platform 610 may be composed of software, firmware, or hardware that is implemented on or accessible to the computing device 600. According to the embodiments described herein, the diagnostic platform 610 may include a training module 612, a processing module 614, a diagnostic module 616, an analysis module 618, and a graphical user interface (GUI) module 620.

[0064] The training module 612 may be responsible for training the models used by the diagnostic platform 610. The training can be performed in a supervised, semi-supervised, or unsupervised manner. For example, assume that the training module 612 receives an input indicating a request to train a model for identifying breathing events in audio data. In such a situation, the training module 612 may obtain an untrained model and subsequently train the model using, for example, audio data labeled as "breathing event" or "no breathing event". Thus, the labeled audio data can be provided to the model as training data so that the model learns how to identify breathing events. Typically, the model will learn to identify patterns of values in the audio data that indicate breathing events.

[0065] The processing module 614 can process the voice data acquired by the diagnostic platform 610 into a format suitable for other modules. For example, the processing module 614 can apply rules, heuristics, or algorithms to the voice data in preparation for analysis by the diagnostic module 616. As another example, the processing module 614 can apply rules, heuristics, or algorithms to the output generated by the diagnostic module 616 in preparation for analysis by the analysis module 618. Thus, the processing module 614 can play a role in ensuring that appropriate data is accessible to other modules of the diagnostic platform 610. Further, the processing module 614 can play a role in ensuring that the output generated by other modules of the diagnostic platform 610 is suitable for storage (e.g., into the memory 604) or transmission (e.g., via the communication module 608).

[0066] The diagnostic module 616 can play a role in identifying an appropriate model for application to the voice data acquired by the diagnostic platform 610. For example, the diagnostic module 616 may identify an appropriate model based on the voice data, the attributes of the patient or the stethoscope system. These attributes may be specified within the metadata associated with the voice data. As an example, an appropriate model may be identified based on the anatomical region where the internal sound is recorded. Alternatively, the diagnostic module 616 may identify an appropriate model based on the measurement criteria generated by the analysis module 618. As an example, when the analysis module 618 is responsible for calculating the respiratory rate, the diagnostic module 616 may subsequently identify a module capable of detecting respiratory events. The desired measurement criteria may be specified by a medical expert or determined by the diagnostic platform 610.

[0067] Generally, the model applied to the voice data by the diagnostic module 616 is one of a plurality of models maintained in the memory 604. These models can be associated with various different breathing events, diseases, etc. For example, the first model is trained to distinguish inhalation and exhalation after being designed, and the second model can be trained to distinguish instances of wheezing or crackling sounds after being designed.

[0068] Generally speaking, each model can represent a collection of algorithms that generate an output that conveys information that can provide insights into the patient's health status when applied to voice data. For example, when the model applied by the diagnostic module 616 distinguishes inhalation and exhalation, the output can be useful for establishing whether the patient's breathing is normal or abnormal. As another example, when the model applied by the diagnostic module 616 identifies instances of wheezing or crackling sounds, the output can be useful for determining the probability that the patient has a given disease.

[0069] In some situations, the output generated by the model applied by the diagnostic module 616 may not be particularly useful by itself. The analysis module 618 can play a role in considering the context of these outputs in a more holistic sense. For example, the analysis module 618 can generate one or more metrics representing the patient's health based on the output generated by the model applied by the diagnostic module 616. These metrics can provide insights into the patient's health without requiring a complete analysis or understanding of the output generated by the model. For example, assume that the model applied by the diagnostic module 616 identifies breathing events based on the analysis of voice data. While knowledge of the breathing events can be useful, medical experts may be interested in metrics such as the respiratory rate. The analysis module 618 can calculate the respiratory rate based on the breathing events identified by the diagnostic module 616.

[0070] The GUI module 620 can play a role in establishing a way to present information for review on the display mechanism 606. Depending on the nature of the display mechanism 606, various types of information can be presented. For example, for presentation to medical experts, the information derived, inferred or obtained by the diagnostic module 616 and the analysis module 618 can be presented on the interface. As another example, visual feedback can be presented on the interface to show when a patient has experienced a change in health.

[0071] Figure 7 includes an example of a workflow diagram showing how the audio data obtained by the diagnostic platform is processed by the background service before the analysis of the audio data is displayed on the interface for review. As shown in Figure 7, the data is obtained and processed by the "backend" of the diagnostic platform, while the analysis of the data is presented by the "frontend" of the diagnostic platform. Individuals such as medical experts and patients can also interact with the diagnostic platform via the frontend. For example, commands are issued via the interface displayed on the display. Further, the diagnostic platform can generate notifications, as will be discussed further below. For example, if the diagnostic platform determines that a patient's respiratory rate is below a predetermined threshold, the diagnostic platform may generate a notification that functions as an alarm. The notification can be presented visually via the display and / or aurally via the speaker.

[0072] Approach for establishing the respiratory rate There are several drawbacks to the conventional approaches for calculating the respiratory rate.

[0073] Some approaches rely on observing inhalation and exhalation at relatively long time intervals. This "observation interval" often lasts for 60 seconds or more. Due to this very long observation interval, these approaches cannot explain short-lived events such as apnea. As a result, since there is little impact on the respiratory rate over the entire observation interval, medical professionals may not always notice, but there is a risk of hardly noticing a temporary cessation of breathing. Since it may be necessary to immediately address a cessation of breathing, misunderstanding the respiratory rate can pose a significant hazard.

[0074] Other approaches rely on monitoring chest movement or end-tidal carbon dioxide (CO2) in exhaled air. However, these approaches are not practical or appropriate in many situations. For example, chest movement may not be an accurate indicator of respiratory events, especially when the patient is under general anesthesia or experiencing a medical event (such as a seizure), and the composition of exhaled air may be unknown when the patient is not wearing a mask.

[0075] This specification presents an approach to address these drawbacks by calculating the respiratory rate through the analysis of audio data. As described above, this approach relies on inhalation and exhalation being detected in a consistent and accurate manner. In auscultation, these respiratory events are important for medical professionals to make diagnostic decisions.

[0076] Generally speaking, this approach includes two stages, namely, a first stage where inhalation and exhalation are detected and a second stage where these inhalations and exhalations are used to calculate the respiratory rate. The first stage is referred to as the "detection stage" and the second stage is referred to as the "calculation stage".

[0077] A. Detection of Respiratory Events Figure 8 shows an overview of the computational pipeline that can be adopted by the diagnostic platform during the detection phase. One of the advantages of this computational pipeline (also referred to as the "computational framework") is its modular design. Each "unit" can be adjusted to achieve the best overall performance after being individually tested. Furthermore, the output of some units can be used for multiple purposes. For example, the spectrogram generated during preprocessing can be provided as input to the model and / or posted on an interface for real-time review.

[0078] For simplicity, this framework is divided into three parts: preprocessing, analysis, and postprocessing. Preprocessing may include not only processing the audio data but also adopting feature engineering techniques. Analysis may include adopting a model trained to identify respiratory events. As described above, the model may include a neural network designed to generate a series of detections (e.g., respiratory events) rather than a single detection (e.g., classification or diagnosis) as output. Each detection may represent an individual prediction made independently by the model. Finally, postprocessing may include examining and / or segmenting the detections generated by the model. Usually, preprocessing is performed by a processing module (e.g., processing module 614 in Figure 6), analysis is performed by a diagnostic module (e.g., diagnostic module 616 in Figure 6), and postprocessing is performed by an analysis module (e.g., analysis module 618 in Figure 6).

[0079] Further details regarding each part are provided in the context of an example below. Those skilled in the art will recognize that the numbers provided below are intended for illustration only. One important aspect of the framework is its flexibility, so other numerical values may be applicable or appropriate in other situations.

[0080] I. Preprocessing In this example, audio data representing the recording of internal sounds generated by the lungs was processed at a sampling frequency corresponding to 4 kilohertz (kHz). Next, a high-pass filter was applied to this audio data with an order of 10 and a cut-off frequency of 80 Hz to remove electrical interference (about 60 Hz) and internal sounds generated by the heart (about 1 - 2 Hz) or another internal organ. Thereby, the diagnostic platform can apply a high-pass filter with a cut-off frequency sufficient to filter out sounds generated by another internal organ that is not the target. This filtered audio data was processed using the short-time Fourier transform (STFT). In this example, the STFT had a Hamming window with a window size of 256 and an overlap rate of 0.25. Thereby, a signal of about 15 seconds could be converted into a corresponding spectrogram of size 938×129. To utilize the spectral information of the target internal sound, the diagnostic platform extracted (i) the spectrogram, (ii) the Mel-frequency cepstral coefficients (MFCC), and (iii) the total energy. In this example, the spectrogram was a logarithmic amplitude spectrogram of 129 bins. Regarding MFCC, the diagnostic platform extracted 20 static coefficients, 20 delta coefficients, and 20 acceleration coefficients. To do this, the diagnostic platform used 40 Mel bands within the frequency range of 0 - 4,000 Hz. The width used for the calculation of the delta coefficients and acceleration coefficients was 9 frames. Thereby, a vector of 60 bins per frame was generated. On the other hand, the diagnostic platform calculated the total energy of three different frequency bands of 0 - 250 Hz, 251 - 500 Hz, and 501 - 1,000 Hz, thereby obtaining three values per frame.

[0081] After extracting these features, the diagnostic platform concatenated these features to form a feature matrix of 938×193. Next, the diagnostic platform applied min-max normalization to each feature so that the values of the normalized features ranged between 0 and 1.

[0082] II. Analysis As described above, the features extracted during preprocessing were provided as input to several models trained to identify respiratory events. To establish the optimal model, six models were tested using the same extracted features. These models included a unidirectional recurrent neural network (Uni-RNN), a unidirectional long short-term memory neural network (uni-LSTM), a unidirectional gated recurrent unit neural network (Uni-GRU), a bidirectional recurrent neural network (Bi-RNN), a bidirectional long short-term memory neural network (Bi-LSTM), and a bidirectional gated recurrent unit neural network (Bi-GRU). Collectively, these models are referred to as the "baseline models."

[0083] The architectures of these baseline models are shown in FIG. 9. The baseline models were designed to be able to detect respiratory events based on the analysis of audio data. The first and second layers are regression layers, which can be RNN, LSTM, or GRU. These regression layers handle the temporal information of the features. Since breathing is usually periodic, these regression layers can learn the nature of the breathing cycle from labeled examples (also referred to as "training data"). To detect the start and end times of respiratory events in the audio data, the diagnostic platform used a time-distributed fully-connected layer as the output layer. This approach resulted in a model that could output a series of detections (e.g., inhalation or no action) rather than a single detection. The sigmoid function was used as the activation function for the time-distributed fully-connected layer.

[0084] Each output generated by each baseline model was a detection vector of size 938x1. Each element of this vector was set to 1 if its value exceeded a threshold, indicating the presence of an inhalation or exhalation in the corresponding time segment, and 0 otherwise. In this example, a single-task learning approach was used for the baseline models, but a multi-task learning approach was also possible.

[0085] In the benchmark model, Adaptive Moment Estimation (ADAM) was used as the optimizer. ADAM is a method for calculating the adaptive learning rate of parameters using stochastic optimization, which is an important process in deep learning and machine learning. When the validation loss did not decrease during 10 epochs, the initial learning rate was set to 0.0001 with step decay (0.2x). This learning process stopped when no improvement was seen for 50 consecutive times.

[0086] III. Post - processing The detection vectors generated by the baseline model can be further processed for other purposes. For example, a diagnostic platform can convert each prediction vector from frame to time for real - time monitoring by medical experts. Further, as shown in Figure 9, domain knowledge can be applied. Since breathing is known to be performed by a living body, the duration of breathing events is usually within a certain range. When the detection vector indicates that there are consecutive breathing events (e.g., inhalation) at short intervals, the diagnostic platform can determine whether to merge these breathing events after examining their continuity. More specifically, the diagnostic platform can calculate the frequency difference of the energy peak (|p_j - p_i|) between the j - th and i - th breathing events when the interval between breathing events is less than T seconds. If this difference is less than a predetermined threshold P, these breathing events are merged as a single breathing event. In this example, T was set to 0.5 seconds and P was set to 25 Hz. When a breathing event is shorter than 0.05 seconds, the diagnostic platform can then simply delete the breathing event completely. For example, the diagnostic platform can adjust the (one or more) labels applied to the corresponding (one or more) segments to indicate that no breathing event occurred.

[0087] IV. Task Definition and Evaluation In this example, two different tasks were performed by the diagnostic platform.

[0088] The first task was the classification of segments of voice data. To achieve this, the recording of each breathing event was first converted into a spectrogram. The temporal resolution of the spectrogram depends on the window size and the overlap rate of the STFT, as described above. For convenience, these parameters were fixed so that each spectrogram would be a matrix of size 938x128. Thus, each recording was divided into 938 segments, and as shown in FIG. 10, each segment was automatically labeled based on the ground truth label.

[0089] After the preprocessing and analysis of the recordings are performed, the diagnostic platform can access the corresponding output by means of sequential detection of size 938x1. This output may be referred to as the "inference result". By comparing the sequential detection with the ground truth segment, the diagnostic platform can define true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). Subsequently, the sensitivity and specificity of the model used to classify these segments can be calculated.

[0090] The second task was the detection of respiratory events in the audio data. After obtaining the sequential detections made by the model, the diagnostic platform can assemble the connected segments with the same label into corresponding respiratory events. For example, the diagnostic platform can identify that the connected segments correspond to inhalation, exhalation, or no action. Further, the diagnostic platform can derive the start time and end time of each assembled respiratory event. In this example, the Jaccard Index (JI) was used to determine whether one respiratory event predicted by the model correctly matched the ground truth event. When the value of JI exceeded 0.5, the diagnostic platform designated the assembled respiratory event as a TP event. When the value of JI was greater than 0 but less than 0.5, the diagnostic platform designated the assembled respiratory event as an FN event. When the value of JI was 0, the diagnostic platform designated the assembled respiratory event as an FP event.

[0091] To evaluate the performance, the diagnostic platform examined the accuracy, sensitivity, positive predictive value, and F1 score of each baseline model. All bidirectional models were superior to the unidirectional models. This result was mainly due to the increase in the complexity of the bidirectional models (and the number of trainable parameters). Overall, Bi-GRU was the most promising among these baseline models, but there could also be situations where other baseline models were more suitable than Bi-GRU.

[0092] B. Calculation of Respiratory Rate A reliable estimation of the respiratory rate plays an important role in the early detection of various diseases. Abnormal respiratory rates may suggest various different diseases and other pathological or psychogenic causes. It has been found that exploring clinically acceptable accurate techniques for continuously establishing the respiratory rate is elusive. As described above, several approaches have been developed to fill this clinical gap, but none have received sufficient credit from medical experts to become the standard of care.

[0093] Figure 11A includes a rough diagram of an algorithm approach for estimating the respiratory rate. As shown in Figure 11A, this algorithm approach depends on the respiratory events predicted by the model adopted by the diagnostic platform. Since appropriately identifying respiratory events is the key to estimating the respiratory rate, the detection stage and the calculation stage need to be completed with consistently high accuracy. Furthermore, the detection stage and the calculation stage can be continuously executed so that the respiratory rate is estimated almost in real time (for example, every few seconds instead of every 30 to 60 seconds).

[0094] Generally speaking, this algorithm approach depends on tracking the occurrence of respiratory events and subsequently continuously calculating the respiratory rate using a sliding window. First, the algorithm executed by the diagnostic platform defines the starting point of a window that includes a portion of the audio data acquired for analysis. Next, the diagnostic platform can monitor the respiratory events inferred by the model as described above to define the ending point of the window. The starting point and the ending point of the window can be defined such that a predetermined number of respiratory events are included within its boundaries. Here, for example, the window includes three respiratory events, but in other embodiments, more or fewer than three respiratory events may be included within the window. In an embodiment programmed such that the window includes three respiratory events, the starting point of the window corresponds to the start of the first respiratory event 1102 included in the window, and the ending point of the window corresponds to the start of the fourth respiratory event 1108. While the fourth respiratory event 1108 defines the ending point of the window, the fourth respiratory event 1108 is not initially included within the window.

[0095] As shown in FIG. 11A, the diagnostic platform can calculate the respiratory rate based on the intervals between inhalations included in the window. Here, the window includes a first respiratory event 1102, a second respiratory event 1104, and a third respiratory event 1106. As described above, the fourth respiratory event 1108 can be used to define the end point of the window. The first period represented by "I" is defined by the time interval between the start of the first respiratory event 1102 and the start of the second respiratory event 1104. The second period represented by "II" is defined by the time interval between the start of the second respiratory event 1104 and the start of the third respiratory event 1106. The third period represented by "III" is defined by the time interval between the start of the third respiratory event 1106 and the start of the fourth respiratory event 1108.

[0096] Each period corresponds to the number of frames (i.e., the time interval). For example, assume that the audio data obtained by the diagnostic platform is a 15-second recording, which is divided into 938 segments during processing, and the duration of each segment is approximately 0.016 seconds. The term "segment" may be used interchangeably with the term "frame". Here, the lengths of the first, second, and third periods are approximately 5 seconds, but the first, second, and third respiratory events 1102, 1104, 1106 have different lengths with respect to segments and seconds. However, those skilled in the art will recognize that the periods between respiratory events do not have to be the same as each other.

[0097] As shown in FIG. 11A, the respiratory rate can be calculated for each frame based on these periods. As an example, consider 15 seconds represented by "A". At this point, since the third respiratory event 1106 has not yet started, the third period is not yet defined. The respiratory rate can be calculated by dividing 120 by the sum of the first and second periods in seconds. Starting from the 16th second represented by "B", the diagnostic platform recognizes the third respiratory event 1106. Thus, starting from the 16th second, the respiratory rate can be calculated by dividing 120 by the sum of the second and third periods in seconds.

[0098] Typically, the diagnostic platform calculates the respiratory rate in a continuous manner when additional respiratory events are identified. In other words, the respiratory rate can be calculated in a "rolling" manner when a new period is identified. Each time a respiratory event is identified (and thus a new period is defined), the diagnostic platform calculates the respiratory rate using the most recent periods. As an example, assume that the first and second periods correspond to 5 seconds and 7 seconds, respectively. In this situation, the respiratory rate is 10 breaths per minute (i.e., 120 / (I + II)=120 / (5 + 7)). As another example, assume that the first and second periods correspond to 4 seconds and 4 seconds, respectively. In this situation, the respiratory rate is 15 breaths per minute (i.e., 120 / (I + II)=120 / (4 + 4)). After another period (i.e., the third period) is identified, the respiratory rate can be calculated using the second and third periods instead of the first and second periods.

[0099] As described above, since breathing is known to be performed by a living body, the duration of a breathing event is usually within a certain range. Therefore, the algorithm used by a diagnostic platform to calculate the respiratory rate can be programmed to increase the denominator if no new breathing event is inferred within a specific number of frames or seconds. In FIG. 11A, the algorithm is programmed to increase the denominator in response to the determination that no new breathing event has been found within 6 seconds. Thus, at 22 seconds represented by "C", the diagnostic platform can calculate the respiratory rate by dividing 120 by the sum of the second period + 6. On the other hand, at 23 seconds represented by "D", the diagnostic platform can calculate the respiratory rate by dividing 120 by the sum of the second period + 7. As shown in FIG. 11A, the diagnostic platform can continue to increase the denominator every second until a new breathing event is found.

[0100] Such an approach ensures that the respiratory rate is updated almost in real time, thereby reflecting changes in the patient's health status. If the patient completely stops breathing, the respiratory rate calculated by the diagnostic platform may not immediately become 0, but the respiratory rate will almost instantaneously show a decreasing trend, which serves as a warning to medical professionals.

[0101] To address situations where the respiratory rate calculated by the diagnostic platform is abnormally low or high, the algorithm can be programmed using a first threshold (also referred to as the "lower threshold") and a second threshold (also referred to as the "upper threshold"). If the respiratory rate falls below the lower threshold, the algorithm may indicate that a warning should be presented by the diagnostic platform. Further, if the diagnostic platform determines that the respiratory rate is below the lower threshold, the diagnostic platform may function as if the respiratory rate were actually 0. Such an approach ensures that the diagnostic platform can easily handle situations where respiratory events have not been detected over a long period without the need for the respiratory rate to actually reach 0. Similarly, if the respiratory rate exceeds the upper threshold, the algorithm may indicate that a warning should be presented by the diagnostic platform. The lower threshold and the upper threshold can be adjusted manually or automatically based on the patient, the services provided to the patient, etc. For example, in the case of a pediatric patient, the lower threshold and the upper threshold may be 4 and 60 respectively, and in the case of an adult patient, the lower threshold and the upper threshold may be 4 and 35 respectively.

[0102] For simplicity, the respiratory rate posted for review (e.g., on the interfaces shown in FIGS. 13A - B) may be an integer value, but the respiratory rate may be calculated to the first or second decimal place (i.e., up to one - tenth or one - hundredth). This can help avoid confusion for medical professionals and patients, especially in emergency situations. As described above, the respiratory rate posted for review may be visually modified if it falls below the lower threshold or exceeds the upper threshold. For example, if the respiratory rate falls below the lower threshold, the diagnostic platform may post an alternative element (e.g., "-" or "--") indicating that the respiratory rate is undetectable. If the respiratory rate exceeds the upper threshold, the diagnostic platform may post another alternative element (e.g., "35+" or "45+") indicating that the respiratory rate is abnormally high. In many situations, medical professionals are more interested in knowing when the respiratory rate is abnormally high than in the exact respiratory rate.

[0103] Generally, the diagnostic platform processes a stream of audio data that is acquired at approximately real-time during recording. In other words, the diagnostic platform can continuously calculate the respiratory rate when a new respiratory event is detected as described above. However, in some embodiments, the diagnostic platform does not calculate the respiratory rate for each single frame in order to conserve processing resources. Instead, the diagnostic platform can calculate the respiratory rate periodically, but even in that case, the respiratory rate is typically calculated frequently (e.g., every few seconds) so that medical experts can reliably recognize changes in the patient's health status.

[0104] Figure 11B shows how the diagnostic platform can calculate the respiratory rate using a sliding window over a recording of a predetermined length (e.g., 12, 15, or 20 seconds) that is updated at a predetermined frequency (e.g., every 2, 3, or 5 seconds). In Figure 11B, the bounding boxes represent the boundaries of the sliding window over which the audio data is processed. After a predetermined amount of time has elapsed, the boundaries of the sliding window shift. Here, for example, the sliding window covers 15 seconds and is shifted by 3 seconds, so 12 seconds of the audio data that was included in the previous sliding window is retained. Such an approach makes it possible to calculate the respiratory rate frequently (e.g., every time the boundaries of the sliding window shift) without consuming excessive amounts of available processing resources.

[0105] Another approach to calculating the respiratory rate relies on autocorrelation rather than ML or AI to perform auscultation. Broadly speaking, the term "autocorrelation" refers to the process by which a signal correlates with a delayed copy of the signal as a function of the delay. Since the analysis of autocorrelation is a common mathematical tool for identifying repeating patterns, it is often used in digital signal processing to derive or infer information about repeating events.

[0106] In embodiments where the diagnostic platform uses autocorrelation for detection, the diagnostic platform uses the STFT to convert the recording of the audio data into a spectrogram as described above. Next, as shown in FIG. 12, in order to determine the interval between inhalation and exhalation, the autocorrelation coefficient is calculated and plotted for the recording. First, the diagnostic platform may normalize the autocorrelation coefficient to reduce noise. Next, a high-pass filter may be applied to the autocorrelation coefficient to exclude those below a threshold (e.g., 15 Hz). Further, the diagnostic platform may perform a trend removal correction to further improve the autocorrelation coefficient.

[0107] After processing the autocorrelation coefficient, the diagnostic platform can calculate the respiratory rate by dividing 60 by the respiratory interval (RI) between the primary peak and the secondary peak. The respiratory rate index can determine which peak is selected. For example, a high respiratory rate index is preferred. Generally, a respiratory rate index of less than about 0.6 may be discarded by the diagnostic platform because it indicates an unstable respiratory rate.

[0108] FIGS. 13A - B include examples of interfaces that may be generated by the diagnostic platform. FIG. 13A shows how a spectrogram generated by the diagnostic platform is presented in near real-time to facilitate a diagnostic decision by a medical professional. On the other hand, FIG. 13B shows how digital elements (also referred to as "graphic elements") overlay the spectrogram to provide additional insights into the patient's health. Here, for example, vertical bands are overlapping on the portion of the spectrogram that the diagnostic platform has determined corresponds to a respiratory event. Note that in some embodiments, these vertical bands cover the entire respiratory cycle (i.e., inhalation and exhalation), while in other embodiments, these vertical bands cover only a part of the respiratory cycle (e.g., inhalation only).

[0109] The respiratory rate can be posted (e.g., in the upper right corner) for review by a medical professional. If the diagnostic platform determines that the respiratory rate is abnormal, the posted value can be visually modified in some way. For example, the posted value can be drawn in a different color, drawn in a different (e.g., larger size), or reposted after being periodically erased so as to "blink" to draw the medical professional's attention. The diagnostic platform can determine that the respiratory rate is abnormal if the value falls below a lower threshold (e.g., 4, 5, or 8) or exceeds an upper threshold (e.g., 35, 45, or 60). Alternatively, the diagnostic platform can determine that the respiratory rate is abnormal if no respiratory events are detected within a predetermined time (e.g., 10, 15, or 20 seconds).

[0110] As shown in FIG. 13B, the interface may include various icons associated with various different functions supported by the diagnostic platform. The record icon 1302, when selected, may initiate the recording of the content shown on the interface. For example, the recording may include a spectrogram along with any digital elements and may also include the corresponding audio data. The freeze icon 1304, when selected, may freeze the interface in its current display. Thus, when the freeze icon 1304 is selected, no additional content may be posted to the interface. The content may be redisplayed on the interface after the freeze icon 1304 is selected again. The control icon 1306, when selected, may control various aspects of the interface. For example, the control icon 1306 may permit a medical professional to lock the interface (e.g., to prevent further changes, either semi - permanently or temporarily), change the layout of the interface, change the color scheme of the interface, change the input mechanism for the interface (e.g., touch or voice), etc. Supplementary information may also be displayed on the interface. For example, in FIG. 13B, the interface includes information regarding active noise cancellation (ANC), playback, and the status of the connected electronic stethoscope system. Thus, it may be readily observable by a medical professional whether the computing device on which the diagnostic platform exists (and on which the interface is shown) is communicatively connected to an electronic stethoscope system responsible for generating audio data.

[0111] Method for calculating respiratory rate FIG. 14 shows a flowchart of a process 1400 for detecting breath events through analysis of audio data and subsequently calculating a respiratory rate based on those breath events. First, the diagnostic platform can obtain audio data representing recordings of sounds generated by a patient's lungs (step 1401). In some embodiments, the diagnostic platform directly receives audio data from an electronic stethoscope system. In other embodiments, the diagnostic platform obtains audio data from a storage medium. For example, a patient may be permitted to upload audio data that they have recorded themselves to a storage medium. As another example, a medical professional may be permitted to upload audio data that they have recorded themselves to a storage medium.

[0112] Next, the diagnostic platform can process the audio data in preparation for analysis by a trained model (step 1402). For example, the diagnostic platform can apply a high-pass filter to the audio data to generate, based on the filtered audio data, (i) a spectrogram, (ii) a series of MFCCs, and (iii) a series of values representing the total energy summed across various frequency bands of the spectrogram. The diagnostic platform can concatenate these features into a feature matrix. More specifically, the diagnostic platform can concatenate these spectrograms, series of MFCCs, and series of values into a feature matrix that can be provided as an input to a trained model. In some embodiments, min-max normalization is performed on the feature matrix such that each entry has a value between 0 and 1.

[0113] Subsequently, the diagnostic platform can apply the trained model to the voice data (or the analysis of the voice data) to generate a vector containing entries arranged in chronological order (step 1403). More specifically, the diagnostic platform can apply the trained model to the feature matrix as described above. Each entry in the vector can represent a detection indicating whether the corresponding segment of the voice data, generated by the trained model, represents a respiratory event. Each entry in the vector can correspond to a different segment of the voice data, although all entries in the vector can correspond to segments of voice data of the same duration. For example, assume that the voice data obtained by the diagnostic platform represents a recording with a total duration of 15 seconds. As part of the preprocessing, this recording can be divided into 938 segments of the same duration. For each of these segments, the trained model can generate an output representing a detection as to whether the corresponding portion of the voice data represents a respiratory event.

[0114] Thereafter, the diagnostic platform can perform post-processing of the entries in the vector as described above (step 1404). For example, the diagnostic platform can examine the vector and thereby identify a pair of entries corresponding to the same type of respiratory event delimited by fewer than a predetermined number of entries, and then merge the pair of entries to indicate that the pair of entries represents a single respiratory event. This can be done to avoid missing respiratory events due to the presence of segments where the model did not indicate a respiratory event. Additionally or alternatively, the diagnostic platform can examine the vector and thereby identify a series of consecutive entries corresponding to the same type of respiratory event shorter than a predetermined length, and then adjust the labels associated with each entry in the series of consecutive entries to delete the respiratory event represented by the series of consecutive entries. This can be done to ensure that respiratory events shorter than a predetermined length (e.g., 0.25, 0.50, or 1.00 seconds) are not recognized.

[0115] Next, the diagnostic platform can identify (i) a first respiratory event, (ii) a second respiratory event, and (iii) a third respiratory event by examining the vector (step 1405). The second respiratory event may follow the first respiratory event, and the third respiratory event may follow the second respiratory event. Each respiratory event may correspond to at least two consecutive entries in the vector indicating that the corresponding segment of the audio data represents a respiratory event. The number of consecutive entries may correspond to the minimum length of the respiratory event imposed by the diagnostic platform.

[0116] Subsequently, the diagnostic platform can determine (a) a first period between the first respiratory event and the second respiratory event, and (b) a second period between the second respiratory event and the third respiratory event (step 1406). As described above, the first period may extend from the start of the first respiratory event to the start of the second respiratory event, and the second period may extend from the start of the second respiratory event to the start of the third respiratory event. Thereafter, the diagnostic platform can calculate the respiratory rate based on the first and second periods (step 1407). For example, the diagnostic platform may divide 120 by the sum of the first and second periods to establish the respiratory rate.

[0117] FIG. 15 represents a flowchart of a process 1500 for calculating a respiratory rate based on an analysis of audio data including sounds produced by a patient's lungs. First, the diagnostic platform can obtain a vector including entries arranged in chronological order (step 1501). Broadly speaking, each entry in the vector may indicate a detection as to whether the corresponding segment of the audio data represents a respiratory event. As described above, the vector may be generated as an output by a model applied to the audio data or by an analysis of the audio data by the diagnostic platform.

[0118] Next, the diagnostic platform can identify (i) a first respiratory event, (ii) a second respiratory event, and (iii) a third respiratory event by examining the entries of the vector (step 1502). Typically, this is accomplished by examining the vector and thereby identifying consecutive entries associated with the same type of respiratory event. For example, the diagnostic platform can analyze the vector and thereby identify consecutive entries with the same detection (e.g., "inspiration" or "no action" as shown in FIG. 9). Thus, each respiratory event can correspond to a series of consecutive entries that (i) exceed a predetermined length and (ii) indicate that the corresponding segment of the voice data represents a respiratory event. As described above, the diagnostic platform can perform post-processing to ensure that false positives and negatives do not affect the ability to identify respiratory events.

[0119] Subsequently, the diagnostic platform can determine (a) a first period between the first respiratory event and the second respiratory event, and (b) a second period between the second respiratory event and the third respiratory event (step 1503). Step 1503 of FIG. 15 can be similar to step 1406 of FIG. 14. Thereafter, the diagnostic platform can calculate the respiratory rate based on the first and second periods (step 1504). For example, the diagnostic platform can divide 120 by the sum of the first and second periods to establish the respiratory rate.

[0120] In some embodiments, the diagnostic platform is configured to display an interface that includes a spectrogram corresponding to the audio data (step 1505). An example of such an interface is shown in FIGS. 13A - B. In such embodiments, the diagnostic platform may post the respiratory rate on the interface to facilitate a diagnostic decision by a medical professional (step 1506). Further, the diagnostic platform may compare the respiratory rate to a threshold as described above (step 1507), and then take an appropriate action based on the result of the comparison (step 1508). For example, if the diagnostic platform determines that the respiratory rate is below a lower threshold, the diagnostic platform may generate a notification indicating that. Similarly, if the diagnostic platform determines that the respiratory rate is above an upper threshold, the diagnostic platform may generate a notification indicating that. Typically, the notification is displayed on the interface where the respiratory rate is posted. For example, the respiratory rate may be visually modified in response to a determination that it is below the lower threshold or above the upper threshold.

[0121] It is assumed that, unless contrary to physical possibility, the above steps can be executed in various orders and combinations.

[0122] For example, process 1500 may be repeatedly executed such that the respiratory rate is continuously recalculated when additional respiratory events are detected. However, as described above, the diagnostic platform may be programmed to increase the denominator (thereby decreasing the respiratory rate) when no respiratory events are detected.

[0123] As another example, the respiratory rate may be calculated using more than two periods. Generally, calculating the respiratory rate using a single period is not desirable because it changes too much in a relatively short period of time to be useful to medical professionals. However, the diagnostic platform may calculate the respiratory rate using three or more periods. For example, assume that the diagnostic platform is tasked with calculating the respiratory rate using three periods defined by four respiratory events. In such a situation, the diagnostic platform may perform most of the steps as described above with reference to FIGS. 14-15. However, the diagnostic platform will divide 180 by the sum of the first, second, and third periods. While most of the process remains the same, to calculate the respiratory rate, the diagnostic platform will divide 60 times "N" (where "N" is the number of periods) by the sum of those "N" periods.

[0124] Other steps may also be included in some embodiments. For example, insights related to a diagnosis, such as the presence of a respiratory abnormality, may be posted to an interface (e.g., the interface of FIG. 13B) or stored within a data structure representing a profile associated with the patient. Thus, insights related to a diagnosis may be stored within a data structure encoded with the corresponding patient's audio data.

[0125] Example Use of the Electronic Stethoscope System and Diagnostic Platform The need for moderate to severe anesthesia has been gradually increasing compared to traditional intubation procedures, and anesthesia has become a mainstream treatment method in medical facilities. However, the risks of respiratory arrest and airway obstruction still cannot be eliminated. There is a lack of a method for directly and continuously monitoring the patient's respiratory state, and the above-described auscultation-led approach provides a solution to this problem.

[0126] According to the surgical safety guidance issued by the World Health Organization (WHO), moderate non-intubated anesthesia requires consistent monitoring by the attending medical professionals. This has been achieved by several methods, including traditionally auscultation-based confirmation of ventilation, end-tidal CO2 monitoring, and verbal communication with the patient. However, these traditional approaches are unrealistic in many situations, if not impossible.

[0127] The above-mentioned auscultation-led approach can be adopted in situations where these traditional approaches are not suitable. For example, the auscultation-led approach can be used in situations such as plastic surgery, gastrointestinal endoscopy, and dental procedures where sedatives are often used to reduce pain. As another example, the auscultation-led approach can be used in situations where there is a high risk of disease transmission (e.g., treatment of patients with respiratory diseases). One of the most important uses of the approach described here is that airway obstruction can be quickly detected by auscultation. By combining an electronic stethoscope system with a diagnostic platform, medical professionals can continuously monitor the patient without the need to manually evaluate the patient's breath sounds.

[0128] A. Plastic Surgery In certain procedures such as rhinoplasty, the patient is unable to wear the necessary mask, which prevents the use of end-tidal CO2 to monitor the respiratory state. The peripheral oxygen concentration meter has an inherent delay and can only generate a notification when the oxygen saturation drops below a certain level.

[0129] B. Gastrointestinal Endoscopy Traditionally, a peripheral blood oxygen concentration meter and an end-tidal CO2 monitor have been used to monitor the respiratory state of patients during gastrointestinal endoscopy. These two approaches have an inherent delay and can only provide a notification when the oxygen saturation decreases or the end-tidal volume changes by at least a certain amount.

[0130] C. Dental Procedures During dental procedures, it is difficult for medical professionals to continuously monitor vital signs. During dental procedures, not only is a high level of concentration required, but it is also not uncommon for there to be only one or two medical professionals. Additionally, patients may be easily nauseated or choked by their own saliva, and this risk is particularly high when these patients are under the influence of anesthesia.

[0131] D. Epidemic prevention To prevent the spread of disease, it is common for patients to be isolated. When isolation is necessary, it can be dangerous for medical professionals to perform traditional auscultation. Additionally, medical professionals may not be able to perform auscultation while wearing personal protective equipment (PPE) such as protective gowns and masks. Simply put, the combination of isolation and PPE can make it difficult to appropriately evaluate a patient's respiratory status.

[0132] The electronic stethoscope system and diagnostic platform can not only provide a means to visualize respiratory events, but can also be used to start the playback of auscultation sounds and generate instant notifications. As shown in FIG. 16, medical professionals can listen to the auscultation sounds to confirm respiratory events and use this to establish the point at which the airway begins to become partially blocked or the respiratory rate begins to become unstable.

[0133] Processing system FIG. 17 is a block diagram showing an example of a processing system 1700 capable of implementing at least some of the operations described herein. For example, the components of the processing system 1700 can be hosted on a computing device that executes the diagnostic platform. Examples of computing devices include electronic stethoscope systems, mobile phones, tablet computers, personal computers, and computer servers.

[0134] The processing system 1700 includes a processor 1702, a main memory 1706, a non-volatile memory 1710, a network adapter 1712 (e.g., network interface), a video display 1718, an input / output device 1720, a control device 1722 (e.g., mechanical input such as a keyboard, pointing device, or button), a drive unit 1724 including a storage medium 1726, and a signal generation device 1730 communicably connected to a bus 1716. The bus 1716 is shown as an abstraction representing one or more physical buses and / or point-to-point connections connected by appropriate bridges, adapters, or controllers. Thus, the bus 1716 can include a system bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express bus, a HyperTransport bus, an Industry Standard Architecture (ISA) bus, a Small Computer System Interface (SCSI) bus, a USB, an Inter-Integrated Circuit (I2C) bus, or a bus compliant with Institute of Electrical and Electronics Engineers (IEEE) standard 1394.

[0135] The processing system 1700 can share the same computer processor architecture as a computer server, router, desktop computer, tablet computer, mobile phone, video game console, wearable electronic device (e.g., watch or fitness tracker), network-connected ("smart") device (e.g., television or home assistant device), an augmented reality or virtual reality system (e.g., head-mounted display), or another electronic device capable of executing a series of instructions (sequential or otherwise) that specify (one or more) actions performed by the processing system 1700.

[0136] Main memory 1706, non-volatile memory 1710, and storage medium 1726 are shown as a single medium, but the terms "storage medium" and "machine-readable medium" should be construed to include a single medium or multiple media that store one or more instruction sets 1728. The terms "storage medium" and "machine-readable medium" should also be construed to include any medium that can store, encode, or carry a series of instructions for execution by processing system 1700.

[0137] Generally, routines executed to implement embodiments of the present disclosure may be implemented as part of an operating system, or of a particular application, component, program, object, module, or series of instructions (collectively referred to as a "computer program"). A computer program typically includes one or more instructions (e.g., instructions 1704, 1708, 1728) set at various memories and storage devices within a computing device at various times. When read and executed by processor 1702, these instructions cause processing system 1700 to perform operations for implementing various aspects of the present disclosure.

[0138] Embodiments have been described in the context of fully functional computing devices, but those skilled in the art will understand that various embodiments can be distributed as various forms of program products. The present disclosure applies regardless of the specific type of machine or computer-readable medium used for actual distribution. Further examples of machine-readable media and computer-readable media include volatile memory, non-volatile memory 1710, removable disks, hard disk drives (HDDs), recordable media such as optical disks (e.g., compact disc read-only memory (CD-ROM) and digital versatile disc (DVD)), cloud-based storage media, and media of transmission types such as digital and analog communication links.

[0139] Network adapter 1712 enables processing system 1700 to mediate data within network 1714 with entities external to processing system 1700 via any communication protocol supported by processing system 1700 and external entities. Network adapter 1712 may include a network adapter card, a wireless network interface card, a switch, a protocol converter, a gateway, a bridge, a hub, a receiver, a repeater, or a transceiver including an integrated circuit (e.g., enabling communication via Bluetooth or Wi-Fi).

[0140] Remarks The foregoing description of the various embodiments of the present technology has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the claimed subject matter to the precise forms disclosed.

[0141] Numerous modifications and variations will be apparent to those skilled in the art. Embodiments have been selected and described in order to best explain the principles of the technology and its practical application, thereby enabling others skilled in the relevant art to understand the claimed subject matter, the various embodiments, and the various modifications suitable for particular applications.

Claims

1. A method for calculating a respiratory rate, comprising: obtaining audio data representing a recording of sounds generated by a patient's lungs; preparing the audio data for analysis by a trained model by applying a high-pass filter to the audio data; generating, based on the audio data, (i) a spectrogram, (ii) a series of Mel-frequency cepstral coefficients (MFCCs), and (iii) a series of values representing the energy summed over different frequency bands of the spectrogram; concatenating the spectrogram, the series of MFCCs, and the series of values to form a feature matrix; processing by; applying the trained model to the feature matrix to generate a vector containing entries arranged in chronological order, wherein each entry in the vector indicates whether the corresponding segment of the audio data represents a respiratory event; inspecting the vector to identify (i) a first respiratory event, (ii) a second respiratory event following the first respiratory event, and (iii) a third respiratory event following the second respiratory event; determining (a) a first period between the first respiratory event and the second respiratory event and (b) a second period between the second respiratory event and the third respiratory event; calculating a respiratory rate based on the first period and the second period; and a method.

2. The method of claim 1, wherein each entry in the vector corresponds to a different segment of the audio data, and all entries in the vector correspond to segments of equal duration of the audio data.

3. The method according to claim 1, wherein each respiratory event corresponds to at least two consecutive entries in the vector, and the at least two consecutive entries indicate that the corresponding segment of the audio data represents a respiratory event. **Claim 4** The step of processing comprises performing min-max normalization on the feature matrix such that each entry has a value between 0 and 1 The method according to claim 1, further comprising. **Claim 5** The method according to claim 1, wherein sounds generated by another internal organ of the patient are filtered from the audio data by the cut-off frequency of the high-pass filter. **Claim 6** examining the vector to identify a pair of entries corresponding to the same type of respiratory event separated by less than a predetermined number of entries; and merging the pair of entries so as to indicate that the pair of entries represents a single respiratory event The method according to claim 1, further comprising. **Claim 7** The method according to claim 6, wherein the pair of entries indicates that the corresponding segment of the audio data represents an inhalation. **Claim 8** The method according to claim 6, wherein the pair of entries indicates that the corresponding segment of the audio data represents an exhalation. **Claim 9** examining the vector to identify a series of consecutive entries corresponding to the same type of respiratory event having a length less than a predetermined length; and adjusting the labels associated with each entry in the series of consecutive entries so as to delete the respiratory event represented by the series of consecutive entries The method according to claim 1, further comprising. **Claim 10** The first period extends from the start of the first respiratory event to the start of the second respiratory event, the second period extends from the start of the second respiratory event to the start of the third respiratory event, and the calculating step is The method according to claim 1, comprising the step of establishing the respiratory rate by dividing 120 by the sum of the first period and the second period.

Citation Information

Patent Citations

  • Vital sign measurement apparatus and body motion detection apparatus

    JP2012105762A

  • Vital measuring instrument

    JP2014061223A

  • Respiration measurement program and respiration measurement apparatus

    JP2018075278A

  • Method and apparatus for training and evaluating artificial neural networks used to determine lung pathology

    US20190088367A1