Obtaining breathing-related sounds from audio recordings
By recording audio in the sleep environment of target patients and utilizing respiratory tracing and classifier technologies, the problems of microphone interference with sleep and sound confusion were solved, enabling accurate identification and screening of respiratory-related sounds of target patients and improving the accuracy of sleep analysis.
Patent Information
- Application Number
- CN202180078851.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-24
- Filing Date
- 2021-09-23
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2041-09-23
AI Technical Summary
Existing technologies for obtaining respiratory sounds from target patients have problems such as microphones interfering with patients' sleep and being unable to accurately identify the target patient's voice, especially in multi-person environments where the source of the sound can be easily confused.
By obtaining audio recordings of the sleep environment of target patients, using respiratory trace recognition and training a classifier, selecting respiratory-related sounds with high or low probability, and eliminating sounds from non-target patients, a computer-implemented method is used for accurate screening.
It enables accurate identification and screening of respiratory sounds from target patients without disturbing their sleep, reducing environmental noise interference and improving the accuracy of sleep analysis.
Smart Images

Figure CN116471988B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to a method for obtaining respiratory-related sounds (RRS) originating from a target patient. Background Technology
[0002] In the field of sleep analysis, one of the elements to be studied is the respiratory-related sounds (RRS). RRS are short audio segments originating from the patient's sounds during sleep analysis, such as snoring, sighing, heavy breathing, or groaning. Further analysis of such sounds can then be used to diagnose sleep disorders, such as sleep apnea. Further research is expected to include the duration of each RRS, the frequency of RRS, the total number of RRS, and to analyze various aspects of the RRS.
[0003] RRS and related metrics can be obtained from the audio recordings of sleeping patients.
[0004] One way to obtain such an audio recording is by attaching a recording microphone to the patient's face, as close as possible to their nose and mouth. This has the advantage of mitigating external sounds and noise through its design. However, the presence of such a microphone can negatively affect the patient's sleep, and therefore the detected RRS may not accurately reflect the patient's natural sleep patterns.
[0005] Alternatively, the audio recording device (e.g., a digital audio recording device such as a mobile phone or a dedicated audio recording device) can be placed further near the target patient. This way, the patient is not obstructed by a microphone or any other device on or near their face, resulting in a more natural sleep. However, a disadvantage in this case is that if the patient is not sleeping alone in the room, another person's RRS may be recorded on the audio recording.
[0006] US2020261687A1 discloses a solution for dynamically shielding audible breathing noise identified as being generated by one or more sleep partners. According to various aspects, it protects a subject's sleep by detecting audible breathing noise in the sleep environment, determining that the audible breathing noise is not generated by the subject, and mitigating the perception of audible breathing noise identified as originating from another subject, such as a bed partner, pet, etc. Dynamic shielding reduces the subject's exposure to unwanted sounds and reduces the opportunity to shield sounds that might disturb the subject's sleep.
[0007] Therefore, the object of the present invention is to solve or at least alleviate one or more of the problems described above. In particular, this disclosure aims to provide a method for identifying the RRS of a target patient in a relatively comfortable manner without interfering with the patient's natural sleep. Summary of the Invention
[0008] To this end, according to a first aspect, a computer-implemented method for obtaining respiratory-related sounds (RRS) originating from a target patient is provided, the method comprising the following steps:
[0009] - Obtain the input audio recordings of the target patient's sleep environment;
[0010] - Obtain a respiratory trace of the target patient's breathing, which represents the patient's breathing during the audio recording period;
[0011] - Identify RRS in the input audio recording;
[0012] - Select RRSs from target patients based on respiratory traces;
[0013] And this option includes:
[0014] - Identify a first subset and / or a second subset of RRSs with correspondingly high probabilities of originating from the target patient and / or low probabilities of originating from the target patient;
[0015] - A classifier is trained based on a first subset and / or a second subset to select RRS derived from the target patient; and
[0016] RRS derived from the target patient are selected using a trained classifier.
[0017] The input audio recording covers the target patient's sleep environment; that is, in addition to the target patient's respiratory rhythm (RRS), it may also include RRS from other people or animals and other environmental sounds. Therefore, the input audio recording includes multiple RRS from the target patient. These RRS are then selected, either entirely or partially, during the selection step. To distinguish the RRS originating from the target patient from other sounds, RRS sounds are selected based on respiratory traces (i.e., a representation of the target patient's breathing as a function of the duration covering the input audio recording). Since the RRS originating from the target patient are related to the target patient's breathing, there is a relationship between these RRS and breathing. Therefore, the RRS originating from the target patient can be distinguished from other sounds in the input audio recording.
[0018] This produces a sound set free of other sounds that could negatively impact the analysis, thus allowing for accurate sleep analysis. Furthermore, audio recording does not require being very close to the patient's mouth or chest when filtering out other sounds. This means the microphone does not suppress RRS from the target patient, or induce unwanted RRS itself.
[0019] Regarding the first subset, only those RRSs originating from the target patients with a probability higher than a certain threshold can be selected, such as those with a probability higher than 90%. This ensures low output error. Furthermore, selecting RRSs with high probabilities is generally easy to determine, i.e., requires low computational power and / or memory capacity.
[0020] Regarding the second subset, only those RRS originating from the target patients with a probability below a certain threshold, such as a probability below 10%, can be selected. The second subset can then be further discarded from the results.
[0021] The results obtained from the first and / or second subsets can be further refined by adding additional RRSs that were not assigned to the first and / or second subsets, using a trained classifier. To achieve this, the classifier is first trained using one or both subsets to classify RRSs as belonging to or not belonging to the target patient. In other words, the first and / or second subsets are used as labeled data. The trained classifier is then used to further classify other RRSs, resulting in a broader selection of RRSs originating from the target patient.
[0022] The respiratory trace can be further obtained using techniques available in the art, such as by obtaining the trace from signals obtained by polysomnography, electrocardiography, electromyography or photoplethysmography (PPG).
[0023] One step is to identify RRS. According to one embodiment, this step further includes identifying breathing-related sounds and non-breathing-related sounds, and discarding non-breathing-related sounds.
[0024] In other words, sounds unrelated to breathing are first discarded from the audio recording, resulting in a subset of sounds that serve as RRS but do not necessarily originate solely from the target patient. Based on the respiratory trace, RRS originating from the target patient are then selected from this subset.
[0025] According to one embodiment, the identification includes determining a set of sounds; wherein the sounds in the set originate from the same source; and wherein the selection further includes selecting RRS originating from the target patient from the set of sounds based on respiratory traces.
[0026] In other words, the sounds are first divided into sets or clusters based on their origin. At this point, it is not yet known which sets originate from the target patient. By referencing the respiratory traces, the RRS of a particular set can be assigned to the target patient. Optionally, non-RRS identification and discarding can be performed before or after set determination.
[0027] For example, sounds can be clustered into sets based on their respective sources using a trained classifier.
[0028] Alternatively, classifier training can be performed only when the number of undetermined RRSs is too high (i.e., there are still many identified RRSs that have neither a high nor a low probability of originating from the target patient). In such cases, performing a more computationally intensive classification operation may be useful.
[0029] According to one embodiment, determining the first subset includes determining an audio timestamp associated with the RRS from the input audio recording, and determining a respiratory timestamp associated with the RRS from the respiratory trace; and determining the first subset based on the audio timestamp and the respiratory timestamp.
[0030] In other words, the audio timestamp indicates the occurrence of the corresponding RRS in the input audio recording, and the respiratory timestamp indicates the occurrence of the corresponding respiratory cycle in the target patient. Since the target patient's RRS is related to the patient's breathing, selection can be performed based on these defined timestamps. For this purpose, the timestamps can be characterized by any detectable temporal feature, such as, for example, a start, local maximum, or local minimum. Thus, the selection operation is simplified to first identifying the temporal features and then performing the operation on these features.
[0031] One operation could be determining the time difference between an audio timestamp and its corresponding respiratory timestamp. Since the RRS of the target patient is related to their breathing, the time difference associated with the patient will be fairly constant, while the time differences associated with other sources will expand more randomly.
[0032] Then, by determining the histogram of time differences, those RRS with a high probability of belonging to the target patient will appear relatively more often in the peak of the histogram, and those RSS with a low probability will appear relatively more often in the tail of the histogram.
[0033] According to a second aspect, a controller is disclosed, the controller including at least one processor and at least one memory, the at least one memory including computer program code, the at least one memory and the computer program code being configured to use the at least one processor to cause the controller to perform the method according to the first aspect.
[0034] According to a third aspect, a computer program product is disclosed, the computer program product including computer-executable instructions for performing the method according to the first aspect when the program is run on a computer.
[0035] According to a fourth aspect, a computer-readable storage medium is disclosed, the computer-readable storage medium comprising a computer program product according to a third aspect. Attached Figure Description
[0036] Figure 1The illustration shows the steps performed according to an example embodiment for selecting respiratory-related sounds from an audio recording originating from a patient;
[0037] Figure 2 The illustration shows the steps performed according to an example embodiment for selecting respiratory sounds originating from a patient from a plurality of respiratory sounds and respiratory traces;
[0038] Figure 3 The illustration depicts the steps performed according to an example embodiment for an expanded set of respiratory-related sounds derived from a patient's selection;
[0039] Figure 4 The illustration shows the steps performed according to an example embodiment for selecting respiratory-related sounds from an audio recording originating from a patient;
[0040] Figure 5 An illustrative graph of an audio recording with defined breathing-related sounds is shown, as well as a graph of a breathing trace with breathing-related timestamps and breathing-related sound timestamps.
[0041] Figure 6 Another illustrative graph of an audio recording with defined breathing-related sounds is shown, as well as a graph of a breathing trace with breathing-related timestamps and breathing-related sound timestamps.
[0042] Figure 7A A histogram showing the time difference when all RRS originate from the target patient is presented;
[0043] Figure 7B A histogram showing the time difference when no RRS originates from the target patient is displayed;
[0044] Figure 7C This shows a histogram illustrating the time difference when RRS originates from different sources; and
[0045] Figure 8 A computing system suitable for performing the various steps according to the example embodiments is shown. Detailed Implementation
[0046] Figure 1Different steps of a computer-implemented method 100 for identifying respiratory-related sounds (RRS) 160 originating from a target patient (i.e., the monitored patient) in an input audio recording 110 are illustrated. RRS correspond to audible events generated by breathing during sleep. Such RRS may correspond, for example, to snoring sounds, sighing sounds, heavy breathing sounds, groaning sounds, or sounds generated during sleep apnea events. RRS occur within the respiratory cycle, such as during inspiration, during expiration, or both. Snoring patients thus generate RRS sequences at time intervals (e.g., seconds, minutes, or even hours). Traces of RRS originating from the monitored patient are valuable for performing sleep analysis because they can reveal or interpret different types of health conditions.
[0047] The method begins by obtaining an audio track 110 or audio recording 110, and then identifying or selecting an RRS 160 originating from the patient from the audio track 110 or audio recording 110. The audio track is recorded within an audible distance from the target patient, i.e., within the patient's sleep environment. This can be achieved, for example, by placing the audio recording device beside the patient's bed or elsewhere in the patient's bedroom. An illustrative example of such an audio recording is also shown in Figure 111, where the amplitude 112 of the recorded audio signal is presented as a function of time.
[0048] Based on the audio recording 110, different RRS 131-134 are identified in step 120 of method 100. These identified RRS may relate to a specific type of RRS, such as relating only to snoring, or to several or even all possible RRS. By identifying RRS, other sounds or noises (e.g., sounds from outside the room) are excluded from further steps. RRS can be identified, for example, by indicating its start time, its end time, and / or the time period within the audio recording 110 that allows it to be uniquely identified.
[0049] RRS identification can be performed, for example, by executing one or more of the following steps:
[0050] a) For example, the acoustic envelope of signal 112 can be determined by calculating the analytical signal of signal 112, by calculating the moving average (e.g., root mean square (RMS) value) of signal 112, or by calculating the peak value of signal 112.
[0051] b) Determine the threshold representing the effective sound segment. This can be done, for example, by calculating the local signal energy value and establishing the lower percentile of the local signal energy to define the baseline threshold.
[0052] c) Calculate when the sound envelope exceeds the threshold.
[0053] d) Mark all segments whose envelope exceeds the threshold as valid segments.
[0054] e) Combine or remove effective segments according to a set of decision rules, for example, to avoid unlikely large or small effective segments.
[0055] f) The effective segment thus obtained is characterized by calculating a set of features such as Mel frequency cepstral coefficients (MFCC), signal power within a specific frequency range, time features such as signal mean and standard deviation, features characterizing the entropy of the signal, features characterizing formants and pitch.
[0056] g) Identify RRS from valid segments, for example, by classifying all valid segments as RRS or non-RRS using a pre-trained classifier, thereby obtaining a set of RRS segments that may originate from one or more sources.
[0057] The identified RRS130s are not necessarily all derived from the target patient. For example, some may originate from another person sleeping next to the patient or in the same room. Furthermore, some RRSs may originate from animals, such as from a dog sleeping in the same room. Therefore, in the subsequent selection step 140, a subset 160 of RRS130 is selected as originating from the monitored patient. For this purpose, subset 160 is selected using respiratory traces 150 from the patient. Such respiratory traces characterize the patient's breathing during the audio recording period 110. Graph 151 illustrates such traces of the patient as a function of time. The rising edge can then correspond to inspiration, and the falling edge to expiration, or vice versa. The respiratory traces can also correspond to discrete timestamps characterizing different respiratory cycles. An observable temporal relationship exists between trace 150 and the RRSs derived from the patient, while other RRSs will not show such a temporal relationship. Based on this, RRS160 derived from the patient is selected as the output of step 140.
[0058] Breathing traces can be obtained directly or indirectly from measurements taken on a patient. For example, traces can be obtained from signals acquired by polysomnography, electrocardiography, electromyography, photoplethysmography (PPG), or accelerometers.
[0059] According to one embodiment, the selection 140 of RRS160 can be determined by, for example... Figure 2Step 200 of the diagram is performed. First, in steps 201 and 202, timestamps 203 and 204 are identified for RRS 130 and respiratory track 150, respectively. For RRS 130, RRS timestamp 203 can characterize the start, end, or any predetermined time reference within the occurrence of an RRS. For respiratory track 150, respiratory timestamp 204 identifies the respiratory cycle, such as the start, end, or any predetermined time reference during a respiratory cycle (inspiration or expiration). Then, in step 205, the difference 206 between timestamps 203 and 204 is determined; that is, for each RRS timestamp 203, the time difference is determined using nearby respiratory timestamps 204 (e.g., using the next or previous respiratory timestamp). Thus, a sequence of time differences 206 is obtained, where each time difference is associated with a corresponding RRS. Based on these time differences 206, a histogram 208 is constructed in the next step 207. Histogram 208 represents the occurrence of a certain time difference or time difference interval. In such a histogram 208, time differences with high incidence rates indicate a strong temporal correlation between the associated RRS and the respiratory trace, and therefore a high probability of originating from the patient. Similarly, time differences with low incidence rates indicate little temporal correlation between the associated RRS and the respiratory trace, and therefore a low probability of originating from the patient. Therefore, RRS 212 with an incidence rate above a certain first threshold are then selected as having a high probability of originating from the patient and added to the selection 160 of patient RRSs. Further RRS 210 with an incidence rate below a certain second threshold can then be selected as having a low probability of originating from the patient. The remaining RRS 211 then remain unassigned. The unassigned RRS 211 can still be used to further expand the patient RRS 160 set, as referenced. Figure 3 and Figure 4 Further examples are described below.
[0060] Another way to select patient RRS160 is by calculating the coherence of one or more RRS130 with the respiratory trace 150, i.e., the degree of synchronization between the audio and respiratory signals of one or more RRSs during the same time interval. In this case, one or more RRSs with high coherence are considered to have a high probability of originating from the patient, and one or more RRSs with low coherence are considered to have a low probability of originating from the patient, thus again obtaining a similar set of RRSs 210, 211, 212. Similar to... Figure 2 The method then selects RRS212 with a high probability as originating from the patient.
[0061] By probability (e.g., by Figure 2The steps for selecting RRSs from patients can depend on whether the results are further expanded. For example, a considerable number of RRSs211 may still remain unassigned, i.e., neither having a low probability of originating from a patient nor a high probability of originating from a patient. In such cases, it is possible to perform... Figure 3 The illustrated step 300. In the first step 301, by selecting RRSs with high and / or low probabilities (e.g., by performing as shown in the reference) Figure 2 The initial selection 302 is performed using step 200. Then, in step 303, further RRSs are identified as originating from patients based on sets of RRSs with high and / or low probabilities (e.g., sets 210 and 212). Based on these sets, some of the unassigned RRSs are further assigned as originating from or not originating from patients. Step 303 can be performed in different ways. According to a first example, step 303 includes training a classifier to classify RRSs based on whether they originate from patients. For training, RRSs with high and / or low probabilities are used as labeled training data. The trained classifier is then used to add unassigned RRSs (e.g., RRS 211) to selection 160. According to a second example, an unsupervised clustering method is used to select unassigned RRSs with similar feature content having similar temporal coherence to the RRSs from either the high-probability set or the low-probability set. The unassigned RRSs clustered with the high-probability set are then added to selection 160.
[0062] Figure 5 , Figure 6 Figure 7 also illustrates step 200. Figure 5 It shows an audio recording 510 and, for example, a device made of... Figure 1 Step 120 yields the first curve of the identified RRS 511. Figure 5 A second graph with a respiratory trace 520 is also shown. In the respiratory trace 520, the start of RRS 511 is indicated by a circle 521, and an RRS timestamp 524 is represented. In the respiratory trace 520, the periodic minimum of the trace is indicated by a cross 522, and a respiratory-related timestamp 525 is represented. Then, the time difference 526 is represented by the space between the dashed line representing the RRS timestamp and the dotted line representing the previous or next RR timestamp. Figure 5 All RRS 511 shown are derived from patients. Therefore, there is a strong temporal relationship between RR timestamp 525 and RRS timestamp 524, which can be observed through a nearly constant time difference 526. Figure 7A This shows that from only sources such as Figure 5 The figure shows a histogram of the time difference obtained from the RRS of the patient. (710)
[0063] Similar to Figure 5 , Figure 6 It shows an audio recording 610 and, for example, a device made of... Figure 1 Step 120 yields the first curve of the identified RRS 611. Figure 6 A second graph with a respiratory trace 620 is also shown. In the respiratory trace 620, the start of RRS 611 is indicated by a circle 621, and an RRS timestamp 624 is represented. In the respiratory trace 620, the periodic minimum of the trace is indicated by a cross 622, and a respiratory-related timestamp 625 is represented. Then, the time difference 626 is represented by the space between the dashed line representing the RRS timestamp and the nearest dotted line representing the RR timestamp. Figure 6 The RRS 611 shown is not from the patient. Therefore, there is a weak temporal relationship between the RR timestamp 625 and the RRS timestamp 624, which can be observed through the highly variable time difference 626. Figure 7B Then it shows the source only from such Figure 6 The figure shows a histogram of the time difference obtained from the RRS of the patient. (720)
[0064] Figure 7C Then it was shown that the data was based on... Figure 5 and Figure 6 The histogram 730 showing the time difference between the two is a combination of histograms 710 and 720. Thus, the data in histogram 730 corresponds to the histogram data 208 of method 200. (See reference...) Figure 2 As explained in step 209, a first threshold 731 can then be defined to select RRS 735 with a high probability, and a second threshold 732 can then be defined to select RRS 733, 737 with a low probability. The remaining RRSs then remain unassigned, as illustrated in regions 734 and 736.
[0065] According to one embodiment, it is possible to, for example Figure 1 Further clustering steps are performed in method 100 as illustrated. This will refer to... Figure 4 The method is further explained below. In the first step 420, corresponding to step 120, RRS 430 are identified from the input audio recording 410. Then, an additional clustering step 470 is performed. In this step 470, RRSs with a high probability of belonging to the same source are grouped into clusters.
[0066] One approach to clustering 470 is to first determine a set of features characterizing the RRS, such as Mel-frequency cepstral coefficients (MFCC), signal power within a specific frequency range, temporal features such as signal mean and standard deviation, features characterizing the entropy of the RRS, and features characterizing formants and tones. Additionally or complementaryly, RRSs occurring in time-repeating patterns can be identified, thus obtaining different RRS chains. The RRSs are then clustered into distinct, seemingly real sources based on their association with the time chains and / or based on the similarity between the different obtained features. Feature-based clustering can be performed, for example, by clustering algorithms such as K-means clustering and Gaussian mixture model (GMM) clustering. Time-chain-based clustering can be performed, for example, by identifying repeating RRS patterns with specific time intervals between occurrences. Through clustering, the RRSs can still remain unassigned, i.e., highly probable not belonging to a particular source. In such cases, further supervised clustering steps can be performed. A classifier is then trained to classify the RRSs into clusters using the already clustered RRSs as labeled training data. For the classifier, a support vector machine (SVM) or a neural network can be used.
[0067] The resulting RRS clusters 471 are then used as input to a further selection step 440, in which clusters with high and / or low probabilities originating from patients are identified. The clusters with high probabilities are then selected as output 160. Step 440 can be performed in the same manner as step 140 or step 200, but based on RRS clusters instead of individual RRS clusters. Further, an additional step 403 can be performed, in which unassigned RRS clusters are added to output 160 in the same manner as step 303, but based on RRS clusters instead of individual RRS clusters.
[0068] The steps of the embodiments described above can be performed by any suitable computing circuitry, such as a mobile phone, tablet, desktop computer, laptop computer, and local or remote server. The steps of the embodiments described above can be performed on the same device as the audio recording apparatus. For this purpose, audio recording can also be performed by, for example, a mobile phone, tablet, desktop computer, or laptop computer. The steps of the embodiments described above can also be performed by suitable circuitry located away from the patient environment. In such cases, the audio recording can be provided to the circuitry via a communication network such as the Internet or a dedicated network.
[0069] Figure 8A suitable computing system 800 is illustrated, comprising circuitry that implements the execution of steps according to the described embodiments. The computing system 800 can generally be configured as a suitable general-purpose computer and includes a bus 810, a processor 802, local memory 804, one or more optional input interfaces 814, one or more optional output interfaces 816, a communication interface 812, a storage element interface 806, and one or more storage elements 808. The bus 810 may include one or more wires allowing communication between components of the computing system 800. The processor 802 may include any type of conventional processor or microprocessor that interprets and executes programmed instructions. The local memory 804 may include random access memory (RAM) or another type of dynamic storage device storing information and instructions for execution by the processor 802, and / or read-only memory (ROM) or another type of static storage device storing static information and instructions for use by the processor 802. The input interface 814 may include one or more conventional mechanisms (such as a keyboard 820, mouse 830, pen, voice recognition) and / or biometric devices, cameras, etc., allowing an operator or user to input information into the computing device 800. Output interface 816 may include one or more conventional mechanisms, such as display 840, for outputting information to an operator or user. Communication interface 812 may include any transceiver-like mechanism, such as one or more Ethernet interfaces, enabling computing system 800 to communicate with other devices and / or systems, such as other computing devices 881, 882, 883. The communication interface 812 of computing system 800 may be connected to another computing system via a local area network (LAN) or wide area network (WAN) (such as the Internet). Storage element interface 806 may include a storage interface (such as a Serial Advanced Technology Attachment (SATA) interface or a Small Computer System Interface (SCSI)) for connecting bus 810 to one or more storage elements 808 (such as one or more local disks, such as SATA disk drives), and for controlling the reading of data from and / or writing of data to these storage elements 808. Although the storage element 808 is described as a local disk, any other suitable computer-readable medium may be used, such as a removable disk, optical storage media (such as a CD-ROM or DVD-ROM), solid-state drive, flash memory card, etc.
[0070] As used in this application, the term "circuit" may refer to one or more or all of the following:
[0071] (a) Hardware circuit implementation only, such as implementations in analog and / or digital circuits only, and
[0072] (b) Combinations such as hardware circuitry and software (if applicable):
[0073] (i) A combination of analog and / or digital hardware circuitry with software / firmware, and
[0074] (ii) Any part of a hardware processor (including a digital signal processor) with software, software, and memory, which work together to enable a device (such as a mobile phone or server) to perform various functions, and
[0075] (c) Hardware circuitry and / or processors that require software (e.g., firmware) to operate, such as a microprocessor or a portion thereof, but which may be absent when the software is not required to operate.
[0076] This definition of "circuit" applies to all uses of the term in this application, including in any claim. As a further example, as used herein, the term "circuit" also covers only hardware circuitry or a processor (or processors) or a portion thereof and its accompanying software and / or firmware implementation. The term "circuit" also covers (e.g., and if applicable to elements of a particular claim) baseband integrated circuits or processor integrated circuits for mobile devices or similar integrated circuits in servers, cellular network devices, or other computing or networking devices.
[0077] Although the invention has been described with reference to specific embodiments, it will be apparent to those skilled in the art that the invention is not limited to the details of the foregoing illustrative embodiments, and that various changes and modifications can be made to implement the invention without departing from its scope. Therefore, the embodiments are to be considered illustrative in all respects and not restrictive, and the scope of the invention is indicated by the appended claims rather than by the foregoing description; and thus, all variations within the meaning and scope of equivalents of the claims are intended to be included therein. In other words, the invention is intended to cover any and all modifications, variations, or equivalents that fall within the scope of the basic principles and whose essential properties are claimed in this patent application. The reader of this patent application will also understand that the words “comprising” or “including” do not exclude other elements or steps, the words “a” or “an” do not exclude a plurality, and a single element such as a computer system, processor, or other integrated unit may perform the functions of the plurality of means listed in the claims. Any reference numerals in the claims should not be construed as limiting the individual claims. When used in the specification or claims, the terms “first,” “second,” “third,” “a,” “b,” “c,” etc., are introduced to distinguish similar elements or steps and do not necessarily describe a sequence or chronological order. Similarly, the terms "top," "bottom," "above," "below," etc., are introduced for descriptive purposes and do not necessarily indicate relative positions. It should be understood that such terms are interchangeable where appropriate, and embodiments of the invention can operate in other orders or orientations different from those described or illustrated above, according to the invention.
Claims
1. A computer-implemented method for obtaining breathing-related sounds originating from a target patient, the method comprising: - obtaining an input audio recording of a sleep environment of the target patient; - obtaining a breathing trace of respiration of the target patient, the breathing trace characterizing respiration of the patient during a time period of the audio recording; - identifying breathing-related sounds in the input audio recording; and - selecting, based on the breathing trace, breathing-related sounds originating from the target patient from the breathing-related sounds; and wherein the selecting comprises: - determining a first subset and / or a second subset of the breathing-related sounds having a respective high probability of originating from the target patient and / or a low probability of originating from the target patient; - training a classifier based on the first subset and / or the second subset to select breathing-related sounds originating from the target patient; and - selecting, by the trained classifier, the breathing-related sounds originating from the target patient.
2. The method according to claim 1, wherein the identifying comprises determining breathing-related sounds and non-breathing-related sounds, and discarding the non-breathing- related sounds.
3. The method according to claim 1, wherein the identifying comprises determining a set of sounds; wherein the sounds of one set originate from the same source; and wherein the selecting further comprises selecting, based on the breathing trace, breathing-related sounds from the set of sounds originating from the target patient.
4. The method according to any one of claims 1 to 3, wherein the selecting further comprises discarding the second subset from the breathing-related sounds.
5. The method according to any one of claims 1 to 3, wherein the selecting comprises performing the training depending on an amount of breathing-related sounds not assigned to the first subset and the second subset. determining, from the input audio recording, audio timestamps associated with the breathing- related sounds, and determining, from the breathing trace, breathing timestamps associated with the breathing-related sounds; and determining the first subset based on the audio timestamps and the breathing timestamps.
6. The method of any of claims 1-3, wherein determining the first subset comprises:
7. The method according to claim 6, wherein determining the first subset further comprises determining a time difference between the audio timestamps and the respective breathing timestamps.
8. The method according to claim 7, wherein determining the first subset further comprises determining a histogram of the time differences; and identifying the first subset from the histogram.
9. The method according to any one of claims 1 to 3, wherein the breathing trace is derived from a signal obtained by a polysomnograph, an electrocardiograph, an electromyograph, or a photoplethysmograph.
10. A controller comprising at least one processor and at least one memory including computer program code, the controller being configured to: - obtain an input audio recording of a sleep environment of a target patient; - obtain a breathing trace of respiration of the target patient, the breathing trace characterizing respiration of the target patient during a time period of the input audio recording; - identify breathing-related sounds in the input audio recording; and - select, based on the breathing trace, breathing-related sounds originating from the target patient from the breathing-related sounds. - selecting, from the respiration-related sounds, the respiration-related sounds originating from the target patient based on the respiration pattern by: - determining a first subset and / or a second subset of the respiration-related sounds having a respective high probability of originating from the target patient and / or a low probability of originating from the target patient; - training a classifier based on the first subset and / or the second subset to select respiration-related sounds originating from the target patient; and - selecting, by the trained classifier, the respiration-related sounds originating from the target patient.
11. The controller of claim 10, wherein to identify the respiration-related sounds, the controller is configured to determine respiration-related sounds and non-respiration- related sounds, and to discard the non-respiration-related sounds.
12. The controller of claim 10, wherein to identify the respiration-related sounds, the controller is configured to determine a set of sounds; wherein sounds of one set originate from the same source; and wherein to select from the respiration-related sounds, the controller is further configured to select, based on the respiration pattern, respiration-related sounds from the set of sounds originating from the target patient.
13. The controller of any one of claims 10 to 12, wherein the selecting further comprises discarding the second subset from the respiration-related sounds.
14. The controller of any one of claims 10 to 12, wherein to select from the respiration-related sounds, the controller is configured to perform the training depending on an amount of respiration-related sounds not assigned to the first subset and the second subset. determining, from the input audio recording, audio timestamps associated with the respiration-related sounds, and determining, from the respiration pattern, respiration timestamps associated with the respiration-related sounds; 15. The controller of any one of claims 10 to 12, wherein determining the first subset comprises: and determining the first subset based on the audio timestamps and the respiration timestamps.
16. The controller of claim 15, wherein determining the first subset further comprises determining a time difference between the audio timestamps and the respective respiration timestamps.
17. The controller of claim 16, wherein determining the first subset further comprises determining a histogram of the time differences; and identifying the first subset from the histogram.
18. The controller of any one of claims 10 to 12, wherein the respiration pattern is derived from signals obtained by a polysomnograph, an electrocardiograph, an electromyograph, or a photoplethysmograph.
19. The controller of claim 18, further comprising the polysomnograph, the electrocardiograph, the electromyograph, or the photoplethysmograph.
20. A computer-readable storage medium comprising computer-executable instructions for performing the method according to any one of claims 1 to 9 when a program comprising the method is run on a computer.
Citation Information
Patent Citations
Dynamic masking depending on source of snoring
US20200261687A1
Systems and methods for screening, diagnosis, detection, monitoring and / or treatment
CN119031880A