Audio processing method, apparatus and computing device for lung function indicators
By acquiring user identity information and audio clips, and using machine learning models to predict lung function indicators, this method solves the problems of large size and high price of traditional lung function instruments and equipment, enabling convenient and economical lung function testing, and improving the accuracy and convenience of testing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BOJIANG LIFE SCI (SHANGHAI) CO LTD
- Filing Date
- 2023-04-04
- Publication Date
- 2026-07-21
AI Technical Summary
Traditional pulmonary function instruments are bulky, cumbersome to operate, inconvenient to carry, and expensive. There is a lack of convenient and economical methods for testing pulmonary function indicators.
By obtaining user identity information and audio clips, including cough sounds, blowing sounds, or vowels, machine learning models are used to predict lung function indicators. By combining identity information and audio processing technology, the prediction of lung function indicators can be achieved.
This provides a convenient and economical method for testing lung function indicators, which can provide reference for medical workers without the need for doctor intervention, thereby improving the accuracy and convenience of testing.
Smart Images

Figure CN116312551B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an audio processing method, apparatus, and computing device for lung function indicators. Background Technology
[0002] Traditional pulmonary function testing equipment is bulky and cumbersome to use; on the other hand, portable pulmonary function testing equipment is expensive. Therefore, there is a need for methods, devices, or systems that can use audio to predict pulmonary function indicators. Summary of the Invention
[0003] According to one aspect of this disclosure, an audio processing method is provided, comprising: obtaining at least one piece of identity information related to a first user; obtaining at least one audio segment related to the first user, the at least one audio segment including at least one of the following: a first audio segment containing a cough sound emitted by the first user, a second audio segment containing a blowing sound emitted by the first user, or a third audio segment containing a vowel emitted by the first user; processing the at least one audio segment to obtain a processed audio segment; determining at least one predicted value of at least one lung function indicator of the first user based on the at least one piece of identity information; and determining a corresponding final predicted value of the at least one lung function indicator of the first user by means of a machine learning model based on the at least one predicted value and the processed audio segment.
[0004] According to another aspect of this disclosure, an audio processing apparatus is provided, comprising: an identity information acquisition unit for acquiring at least one piece of identity information related to a first user; an audio segment acquisition unit for acquiring at least one audio segment related to the first user, the at least one audio segment including at least one of the following: a first audio segment containing a cough sound emitted by the first user, a second audio segment containing a blowing sound emitted by the first user, or a third audio segment containing a vowel emitted by the first user; an audio segment processing unit for processing the at least one audio segment to obtain a processed audio segment; a first determination unit for determining at least one predicted value of at least one lung function indicator of the first user based on the at least one piece of identity information; and a second determination unit for determining a corresponding final predicted value of the at least one lung function indicator of the first user based on the at least one predicted value and the processed audio segment using a machine learning model.
[0005] According to another aspect of this disclosure, a computing device is provided, comprising: a memory, a processor, and a computer program stored on the memory, wherein the processor is configured to execute the computer program to implement the steps of the method according to embodiments of this disclosure.
[0006] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of the method according to embodiments of this disclosure.
[0007] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, it implements the steps of the method according to embodiments of this disclosure.
[0008] These and other aspects of this disclosure will be apparent from the embodiments described below, and will be elucidated with reference to the embodiments described below. Attached Figure Description
[0009] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0010] Figure 1 This is a schematic diagram illustrating an example system in which various methods described herein may be implemented according to exemplary embodiments;
[0011] Figure 2 This is a flowchart illustrating an audio processing method according to an exemplary embodiment;
[0012] Figure 3 This is a schematic diagram illustrating the cough envelope according to an exemplary embodiment;
[0013] Figure 4 This is a flowchart illustrating an audio processing method according to another exemplary embodiment;
[0014] Figure 5 This is a schematic block diagram illustrating an audio processing apparatus according to an exemplary embodiment;
[0015] Figure 6 This is a block diagram illustrating an exemplary computer device that can be applied to an exemplary embodiment. Detailed Implementation
[0016] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0017] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. As used herein, the term "multiple" means two or more, and the term "based on" should be interpreted as "at least partially based on". Furthermore, the terms "and / or" and "at least one of..." cover any one of the listed items and all possible combinations thereof.
[0018] Exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0019] Figure 1 This is a schematic diagram illustrating an example system 100 in which various methods described herein may be implemented according to exemplary embodiments.
[0020] refer to Figure 1 The system 100 includes a client device 110, a server 120, and a network 130 that communicatively couples the client device 110 and the server 120.
[0021] Client device 110 includes a display 114 and a client application (APP) 112 that can be displayed on the display 114. Client application 112 can be an application that needs to be downloaded and installed before running, or a lightweight application (liteapp). If client application 112 is an application that needs to be downloaded and installed before running, client application 112 can be pre-installed on client device 110 and activated. If client application 112 is a mini-app, user 102 can directly run client application 112 on client device 110 without installing it, by searching for client application 112 in the host application (e.g., by the name of client application 112) or scanning the graphic code of client application 112 (e.g., barcode, QR code, etc.). In some embodiments, client device 110 can be any type of mobile computing device, including mobile computers, mobile phones, wearable computing devices (e.g., smartwatches, head-mounted devices including smart glasses, etc.), or other types of mobile devices. In some embodiments, the client device 110 may alternatively be a fixed computer device, such as a desktop computer, server computer, or other type of fixed computer device.
[0022] Server 120 is typically a server deployed by an Internet Service Provider (ISP) or Internet Content Provider (ICP). Server 120 can represent a single server, a cluster of multiple servers, a distributed system, or a cloud server providing basic cloud services (such as cloud databases, cloud computing, cloud storage, and cloud communications). It will be understood that, although... Figure 1 The diagram shows that server 120 communicates with only one client device 110, but server 120 can provide background services to multiple client devices simultaneously.
[0023] Examples of network 130 include combinations of local area networks (LANs), wide area networks (WANs), personal area networks (PANs), and / or communication networks such as the Internet. Network 130 can be wired or wireless. In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc., are used to process data exchanged through network 130. Furthermore, encryption technologies such as Secure Sockets Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some of the links. In some embodiments, custom and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0024] For the purposes of this disclosure's embodiments, Figure 1 In the example, client application 112 can be a browser or other dedicated application for audio processing. Correspondingly, server 120 can be a server used in conjunction with the client application. Server 120 can provide data processing, storage, and other services to the client application 112 running on client device 110. Alternatively, server 120 can also provide corresponding computing power to client device 110.
[0025] Figure 2 This is a flowchart illustrating an audio processing method 200 according to an exemplary embodiment. Method 200 can be implemented on a client device (e.g., ...). Figure 1 The execution is performed at the client device 110 shown, that is, the execution entity of each step of method 200 can be... Figure 1 The client device 110 shown. In some embodiments, method 200 can be performed on a server (e.g., Figure 1 The method 200 is executed at server 120 (as shown in the figure). In some embodiments, the method 200 may be executed in combination by a client device (e.g., client device 110) and a server (e.g., server 120).
[0026] In the following text, see references Figure 2 Describe each step of method 200 in detail.
[0027] In step 210, at least one piece of identity information related to the first user is obtained.
[0028] In step 220, at least one audio segment associated with the first user is obtained, the at least one audio segment including at least one of the following: a first audio segment containing a cough sound made by the first user, a second audio segment containing a blowing sound made by the first user, or a third audio segment containing a vowel made by the first user.
[0029] In step 230, the at least one audio segment is processed to obtain a processed audio segment.
[0030] In step 240, based on the at least one piece of identity information, at least one predicted value for at least one lung function indicator of the first user is determined.
[0031] In step 250, based on the at least one predicted value and the processed audio segment, a machine learning model is used to determine the corresponding final predicted value of the at least one lung function index of the first user.
[0032] The above method can obtain predicted lung function index values based on the user's identity information and cough sounds. It should be understood that due to individual differences, without professional analysis and confirmation by a doctor, the predicted values obtained solely from Method 200 cannot be used to directly draw a diagnosis of the user's disease or health status. Method 200 is essentially an audio processing method that processes audio signals through a machine learning model. The entire process does not require the involvement of a doctor, and the resulting predicted values can only provide a reference for doctors and medical professionals. Therefore, such a method does not prevent doctors from freely choosing treatment plans and should not be considered a disease diagnosis method.
[0033] In some examples, user identification information may include, but is not limited to, the user's gender, age, height, weight, medical history, etc., and may take the form of numerical values, enumerations, boolean values, text fields, or other forms, and it is understood that this disclosure is not limited thereto.
[0034] According to some embodiments, the at least one audio segment includes the first audio segment, and processing the at least one audio segment to obtain a processed audio segment may include: determining a first background noise threshold based on the first audio segment; identifying at least one valid cough peak from the first audio segment; and obtaining at least one valid cough segment based on the first background noise threshold and the at least one valid cough peak, wherein the processed audio segment includes the at least one valid cough segment.
[0035] In such examples, cough sound segments can be processed in terms of both background noise and whether the cough is effective. This allows for more accurate extraction of the parts of the cough sound segments that are more effective in predicting lung function indicators, resulting in more accurate predictions. Such results are more helpful in assisting users such as doctors and healthcare workers.
[0036] For example, obtaining at least one valid cough segment based on the first background noise threshold and the at least one valid cough peak may include determining at least one cough sound segment corresponding to each of the at least one valid cough based on the first background noise threshold. Specifically, such a process may include determining a cough envelope, identifying valid cough peaks, and determining the portion of each valid cough peak above the background noise threshold based on the background noise threshold.
[0037] refer to Figure 3 An exemplary cough envelope 300 is shown, including exemplary peaks 301, 302, and 303. For example, among peaks 301, 302, and 303, peaks 301 and 303 are identified as valid cough peaks, while peak 302 is a throat clearing sound, noise, or other peak unsuitable for identifying lung function indicators. Therefore, only the cough audio associated with peaks 301 and 303 can be extracted for identification. Further exemplarily, Figure 3 A first background noise threshold TH1 is also shown. Therefore, portions of the envelope of valid cough peaks that are above the background noise threshold, such as portions of envelopes t1-t2 and t5-t6, can be truncated to include them in the processed audio segment. As an example, identified valid cough segments can be spliced, for example, such that envelope t5-t6 immediately follows envelope t1-t2 in the spliced segment, and so on. For the portion of valid cough peak 301 before t1 and after t2, although it is a valid cough, it may contain insufficient valid information because it is below the background noise threshold; therefore, it may be more advantageous to remove it. Similarly, for valid cough peak 303, the portion before t5 and after t6 in its envelope can be removed. For the envelope portion t3-t4, although it is above the background cough threshold, it can be disregarded because 302 is identified as not a valid cough, thus making the predicted value more accurate and representative.
[0038] Understandably, although Figure 3 The image shows three peaks, and shows that the peak values of effective cough peaks 301 and 303 and ineffective cough peak 302 are all above the first background noise threshold TH1, but this disclosure is not limited thereto. For example, there may be more or fewer effective cough peaks and more or fewer ineffective cough peaks in the audio waveform; or, there may be one or more peaks below the first background noise threshold in the audio waveform, and so on.
[0039] According to some embodiments, identifying at least one valid cough peak from the first audio segment may include: calculating the frequency domain energy distribution of the first audio segment; identifying amplitude points exceeding a predetermined valid cough threshold from the frequency domain energy distribution, the valid cough threshold being used to distinguish between cough sounds and throat clearing sounds; and determining the amplitude points exceeding the valid cough threshold as the at least one valid cough peak.
[0040] Based on such examples, it is possible to distinguish between effective coughs with high energy and ineffective coughs with low energy, thereby making the prediction results more accurate.
[0041] As an example, the frequency domain energy distribution can be a root mean square (RMS) energy curve. For instance, a first frequency domain curve can be obtained by performing a short-time Fourier transform on the first audio segment, and then a first RMS energy curve can be obtained based on the first frequency domain curve. It is understood that this disclosure is not limited to this, and the frequency domain energy distribution can take other energy representation forms that can be understood by those skilled in the art.
[0042] According to some embodiments, the at least one audio segment may include a second audio segment containing a blowing sound emitted by the first user, and wherein processing the at least one audio segment to obtain a processed audio segment may further include: determining a second background noise threshold based on the second audio segment; and truncating the second audio segment based on the second background noise threshold to obtain a processed second audio segment, wherein the processed audio segment includes the processed second audio segment.
[0043] In such an embodiment, in addition to cough sounds, the user's breathing sounds can also be considered, thereby obtaining more accurate and helpful prediction results.
[0044] According to some embodiments, the at least one audio segment may include a third audio segment containing a vowel emitted by the first user, and wherein processing the at least one audio segment to obtain a processed audio segment may further include: determining a third background noise threshold based on the third audio segment; and truncating the third audio segment based on the third background noise threshold to obtain a processed third audio segment, wherein the processed audio segment includes the processed third audio segment.
[0045] In such an embodiment, in addition to cough sounds, the user's vowels can also be considered, thereby obtaining more accurate and helpful prediction results.
[0046] It is understood that, according to various embodiments of this disclosure, at least one audio segment may include any one, two, or three of cough sounds, blowing sounds, and vowels.
[0047] According to some embodiments, the method may further include, before obtaining at least one audio segment associated with the first user: obtaining a fourth audio segment; determining a background noise level based on the fourth audio segment; and, in response to determining that the background noise level exceeds a fourth background noise threshold, outputting a prompt to change the measurement environment.
[0048] In such an embodiment, the background noise level where the user is located can be determined in advance, and the user can be prompted to change their location if the background noise is too loud, so as to prevent the noise from affecting the audio quality during the actual recording.
[0049] For example, after obtaining the fourth audio segment, the audio segment can be divided into multiple windows, and the sound pressure level (SPL) of the audio within each window can be calculated. As a non-limiting example, with the sampling rate of the fourth audio segment set to 16000Hz, the audio can be divided into windows of 0.05 seconds. Then, the background noise level can be calculated by averaging all the SPL values. For example, the SPL can be calculated using the following formula:
[0050]
[0051] Among them, P ref As the reference sound pressure level, P e This is the sound pressure level.
[0052] For example, in the air, P can be ref Take 2x10 -5 Pa, or other values as needed.
[0053] Sound pressure level P e The following formula can be used for calculation:
[0054]
[0055] Where N is the number of time-domain sampling points, and x(n) is the standardized value of the nth time-domain sampling point.
[0056] It is understood that, in this document, the first, second, third, and fourth background noise thresholds may be predetermined and may be the same as, at least partially the same as, or different from each other. As a non-limiting example, the fourth background noise threshold may be lower than the first, second, and third background noise thresholds to prompt the user to move to a quieter area before testing. As another non-limiting example, any one of the first, second, and third background noise thresholds may be lower than the fourth background noise threshold to obtain, for example, relatively more audio information emitted by the first user under certain noise interference. It is understood that these noise thresholds can be adjusted according to desired audio quality, model parameters, recognition performance, etc., and this disclosure is not limited thereto.
[0057] In some exemplary embodiments, the background noise threshold may also be determined based on the recorded sound. For example, according to exemplary embodiments of this disclosure, an audio extraction algorithm may also be provided to extract valid vocal segments to prevent excessively long silences at the beginning and end of the recording or the influence of background noise on the model. For example, after obtaining a breath sound segment, the amplitude curve of the envelope can be obtained by performing audio processing operations such as normalization and filtering on the signal, and the amplitude curves can be sorted in ascending order. The first certain percentage (e.g., 15%, 20%, etc.) of amplitude points are extracted, and the mean is used as the estimated noise value, which may be a second background noise threshold, for example, in the case of a breath sound segment. Exemplarily, if the determined background noise threshold is greater than a threshold percentage of the audio waveform amplitude (e.g., the breath sound amplitude), the corresponding background noise threshold can be set to that threshold percentage. As a non-limiting example, the threshold percentage may be 0.1, and when the amplitude point calculated based on 20% is greater than 0.1 of the largest amplitude point, the corresponding background noise threshold is set to 0.1 of the largest amplitude; otherwise, the threshold is set to the noise value. The first point below the threshold can be found before and after the point of maximum amplitude, respectively, as the start and end points of the segment. It is understood that the background noise threshold can be calculated separately for each audio segment such as cough sound, blowing sound, vowel, etc., that is, the first background noise threshold, the second background noise threshold, and the third background noise threshold as described above, but this disclosure is not limited to this.
[0058] According to some embodiments, the at least one lung function indicator includes at least one of the following: forced vital capacity (FVC), forced expiratory volume in one second (FEV1), or the ratio of forced expiratory volume in one second to forced vital capacity (FEV1 / FVC). It is understood that although the architecture of the model is described hereinafter using FVC, FEV1, or their ratio as examples, this disclosure is not limited to these lung function indicators, and predicted values can be generated for other lung function indicators that can be understood by those skilled in the art.
[0059] According to some embodiments, determining at least one predicted value of at least one lung function indicator of the first user based on the at least one piece of identity information may include: fitting the at least one piece of identity information to obtain the at least one predicted value by means of a linear regression model.
[0060] In such an embodiment, the relationship between user identity information and lung function indicators can be linearly fitted using known statistical data to reflect the impact of different user identities on the corresponding lung function indicators. This allows for the acquisition of population-average predicted values based on population statistics, such as predicted FVC and FEV1 values.
[0061] As an example, identity information may include height, weight, age, and gender. A dataset can be composed of identity information and corresponding known lung function index values (e.g., FVC and FEV1 measurements from a hospital pulmonary function testing machine). For example, users can be categorized by gender, with height, weight, age, etc., as variables, and the corresponding measurements (e.g., FVC and FEV1 measurements) as target values, fitted using a linear regression model. A linear regression model formula of the form:
[0062] FVC est =A1*Age+B1*Height+C1*weight+D1
[0063] FEV1 est =A2*Age+B2*Height+C2*weight+D2
[0064] Among them, FVC est and FEV1 estHere, FVC and FEV1 are the predicted values, Age is age, Height is height, and weight is weight, and A1, B1, C1, D1, A2, B2, C2, and D2 are the corresponding parameters of the fitted linear regression model. It is understood that the above formulas and variables are illustrative. For example, the model can use more or less identity information. In one example, additionally or alternatively, identity information may include one or more of the following: BMI, medical history, smoking history, history of neck trauma, occupation, place of origin, etc., or may include other identity information that may be correlated with lung function indicators. Furthermore, the model may be a nonlinear fitting model, and this disclosure is not limited thereto.
[0065] According to some embodiments, the machine learning model may include an encoder portion and a projection portion. In such an example, determining the final predicted value of the at least one lung function indicator of the first user based on the at least one predicted value and the processed audio segment may include: encoding the processed audio segment using the encoder portion to obtain a first encoded vector; concatenating the first encoded vector with the vectorized representation of the at least one predicted value to obtain a second encoded vector; and determining the corresponding final predicted value of the at least one lung function indicator based on the second encoded vector using the projection portion.
[0066] As a specific, non-limiting example, the machine learning model can be an Audio Spectrogram Transformer (AST) model, which may include a transformer encoding part and a linear projection part. In some examples, an AST model modified based on the actual task can be used. Compared to the original AST model, a variant AST according to embodiments of this disclosure can modify the linear layer after the Transformer encoder by taking the class embedding vector output by the Transformer encoder and concatenating it with the vectorized representation of the predicted value. As a more specific, non-limiting example, when at least one pulmonary function index is two or more pulmonary function indices (e.g., FVC, FEV1, and FEV1 / FVC), the method can employ three separately trained variant AST models. For each model, for different tasks, the class embedding vector output by the Transformer encoder is concatenated with the normalized corresponding predicted value (e.g., predicted values of FVC, FEV1, and FEV1 / FVC) to form a new embedding, and the concatenated vector is then input into the corresponding linear layer. In such a specific, non-restrictive example, ReLU can be used as the activation function for an AST model trained for FVC and FEV1 tasks, while Sigmoid can be used for a model trained for FEV1 / FVC tasks. The model can then output the final result. For example, the mean squared error (MSE) can be used as the loss function during training. As a more specific, non-restrictive example, features can be masked (e.g., 50%) during training to improve the robustness of the trained model.
[0067] It is understood that, in the exemplary embodiment, personal information can be anonymized by encoding data, including personally identifiable information. In other examples, other encoding methods, de-identification algorithms, or other data processing methods may be used to process personal information, enabling the removal of personally identifiable information.
[0068] It is understood that although the various operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring that these operations be performed in the specific order shown or in chronological order, nor should it be construed as requiring that all the shown operations be performed to obtain the desired result. For example, step 210 may be performed after step 230, or concurrently with step 230. Furthermore, it is understood that the above method 200 and its variations may be performed by a client device, by a server, or by a combination of client and server devices.
[0069] The following is combined with Figure 4 Another exemplary, non-limiting embodiment according to this disclosure is described.
[0070] In a 401 redirect, a detection request is initiated. For example, the first node can send a detection request via a peer-to-peer network. The first node can be a user client, such as an application on a mobile phone or other nodes that can interact with the user. The first node can connect to the server at the data processing end via the peer-to-peer network, which will be referred to as the second node below for convenience.
[0071] In a 402 redirect, a basic information questionnaire is displayed to the user. For example, after receiving a request, the second node sends a basic information questionnaire to the first node, which may include height, weight, age, gender, BMI, occupation, place of residence, and other information. Upon receiving the questionnaire, the first node can present it to the user for completion. Alternatively, the questionnaire could be pre-stored or cached on the second node.
[0072] In a 403 error, user information data may be sent or stored. For example, after a user completes their input, such as after the first node receives the user's input by clicking "Next," the first node may send the data to the second node. In another example, the first node may also store or temporarily store the user information data and send it to the second node, for example, in a subsequent step, along with an audio file. Yet another example, in an example where the calculation is performed by the first node itself, the first node may also store the user information data without sending it to the second node.
[0073] In a 404 error, microphone access requests can be initiated. For example, after receiving data, the second node returns the reception result and sends a request for microphone access permissions. Alternatively, the first node can initiate the microphone access permission request.
[0074] In step 405, ambient sound recording is performed. After the first node grants microphone permissions, it displays an ambient sound recording prompt and a record button. Clicking the "Record" button initiates audio recording, for example, recording 10 seconds of ambient sound. After recording, the ambient noise level can be analyzed. This analysis can be performed using a built-in algorithm or sent back to the second node for analysis. If the ambient noise level exceeds a set maximum noise level threshold, the first node can output a prompt to the user, such as displaying the text "The environment is noisy; would you like to change to a quieter location for testing?" The first node can provide "Retest" and "Continue" options for the user to choose from. If "Retest" is received, step 405 is re-executed; otherwise, the method proceeds to step 406.
[0075] In node 406, a breath sound segment is recorded. Exemplarily, the first node can display an animated tutorial and text guidance for recording the breath sound, allowing the user to record the segment by clicking "Record" and "Pause" requests. After recording is complete, the first node sends the audio data to the second node. Exemplarily, after recording is complete, "Upload" and "Re-record" options can be output, with data uploaded only if the user selects "Upload," and the breath sound segment recording re-executed when "Re-record" is received. Exemplarily, sending the audio data to the second node can include packaging the Pulse Code Modulation (PCM) data stream into an audio file and sending it to the second node.
[0076] In step 407, the first node records a cough sound segment, which can be similar to step 406. As a specific, non-limiting example, as described above, for a cough sound, the cough audio can be STFT converted, and then the RMS curve along the frequency dimension can be calculated. A threshold that distinguishes a cough sound from a throat clearing sound can be statistically determined based on known sample data, such as a tagged dataset, and amplitude points exceeding this threshold are considered cough sound segments. Subsequently, the first point before the first amplitude point that is lower than a first background noise threshold can be taken as the starting point, and the first point after the last amplitude point that is lower than the first background noise threshold can be taken as the ending point. The first background noise threshold can be predetermined, or it can be calculated based on the amplitude curve of the audio waveform as described above, and this disclosure is not limited thereto.
[0077] In step 408, the first node records a vowel segment, which can be similar to step 406.
[0078] In step 409, lung function indicators are predicted. For example, a linear regression model can be used to fit the information from the questionnaire into predicted values for FVC and FEV1. Then, spectral features are extracted for blowing, coughing, and vowels, and multiple audio features are concatenated. Three different AST models with varying parameters are used to predict the values of FVC, FEV1, and FEV1 / FVC for each feature to obtain the final predicted values. These steps can be performed by the second node, or in some examples, by the first node.
[0079] Understandably, concatenating multiple audio features can include, for example, extracting blowing and continuous coughing audio into fixed durations (e.g., 2.5 seconds, padded with zeros if necessary), and then concatenating the blowing, coughing, and vowel sounds. For instance, during training, only the first 10 seconds of the entire audio can be used. The concatenated audio can be converted to a Mel spectrum and normalized. As a specific, non-limiting example, the spectrum dimension can be [128, 313], where 128 is the frequency dimension and 313 is the time dimension.
[0080] In 410, the final predicted value is displayed. For example, the first node can display one or more predicted values, such as one or more of the predicted FVC, FEV1, and FEV1 / FVC values, or the second node can return one or more of the predicted FVC, FEV1, and FEV1 / FVC values to the first node for display.
[0081] Figure 5 This is a schematic block diagram illustrating an audio processing apparatus 500 according to an exemplary embodiment. The audio processing apparatus 500 may include: an identity information acquisition unit 510, an audio segment acquisition unit 520, an audio segment processing unit 530, a first determination unit 540, and a second determination unit 550. The identity information acquisition unit 510 may be used to acquire at least one piece of identity information associated with a first user. The audio segment acquisition unit 520 may be used to acquire at least one audio segment associated with the first user, the at least one audio segment including at least one of the following: a first audio segment containing a cough sound emitted by the first user, a second audio segment containing a blowing sound emitted by the first user, or a third audio segment containing a vowel emitted by the first user. The audio segment processing unit 530 may be used to process the at least one audio segment to obtain a processed audio segment. The first determination unit 540 may be used to determine at least one predicted value for at least one lung function indicator of the first user based on the at least one piece of identity information. The second determination unit 550 may be used to determine the corresponding final predicted value for the at least one lung function indicator of the first user using a machine learning model based on the at least one predicted value and the processed audio segment.
[0082] It should be understood that Figure 5 The various modules of the device 500 shown can be connected to the reference. Figure 2 The steps in method 200 described correspond to each other. Therefore, the operations, features, and advantages described above for method 200 also apply to apparatus 500 and its included modules. For the sake of brevity, some operations, features, and advantages will not be repeated here.
[0083] According to embodiments of the present disclosure, a computing device is also disclosed, including a memory, a processor, and a computer program stored on the memory, wherein the processor is configured to execute the computer program to implement the steps of the audio processing method and variations thereof according to embodiments of the present disclosure.
[0084] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium is also disclosed, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of the audio processing method and variations thereof according to embodiments of the present disclosure.
[0085] According to embodiments of the present disclosure, a computer program product is also disclosed, including a computer program, wherein when the computer program is executed by a processor, it implements the steps of the audio processing method and variations thereof according to embodiments of the present disclosure.
[0086] While specific functions have been discussed above with reference to specific modules, it should be noted that the functions of the modules discussed herein can be divided into multiple modules, and / or at least some functions of multiple modules can be combined into a single module. The specific module performing an action discussed herein includes the specific module itself performing the action, or alternatively, the specific module calling or otherwise accessing another component or module that performs the action (or performs the action in conjunction with the specific module). Therefore, a specific module performing an action can include the specific module performing the action itself and / or another module that the specific module calls or otherwise accesses to perform the action. As used herein, the phrase "entity A initiates action B" can mean that entity A issues an instruction to perform action B, but entity A itself does not necessarily perform action B. For example, an expression such as "first node displays predicted value" can mean that the first node instructs a display or display device (not shown) to present the predicted value, without the first node itself performing the "display" action.
[0087] It should also be understood that this article can describe various technologies in the general context of software and hardware components or program modules. The above regarding... Figure 5 The various modules described can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these modules can be implemented as computer program code / instructions configured to execute in one or more processors and stored in a computer-readable storage medium. Alternatively, these modules can be implemented as hardware logic / circuit. For example, in some embodiments, one or more of the units or modules described in the embodiments of this disclosure can be implemented together in a System on Chip (SoC). The SoC may include an integrated circuit chip (which includes a processor (e.g., a Central Processing Unit (CPU), microcontroller, microprocessor, digital signal processor (DSP), etc.), memory, one or more communication interfaces, and / or one or more components of other circuitry) and may optionally execute received program code and / or include embedded firmware to perform functions.
[0088] According to one aspect of this disclosure, a computing device is provided, including a memory, a processor, and a computer program stored in the memory. The processor is configured to execute the computer program to implement the steps of any of the method embodiments described above.
[0089] According to one aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of any of the method embodiments described above.
[0090] According to one aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of any of the method embodiments described above.
[0091] The collection, acquisition, storage, use, processing, transmission, provision, and public application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0092] In the following text, combined with Figure 6 Illustrative examples describing such computer devices, non-transitory computer-readable storage media, and computer program products.
[0093] Figure 6 An example configuration of a computer device 600 that can be used to implement the methods described herein is shown. For example, Figure 1 The server 120 and / or client device 110 shown may include an architecture similar to computer device 600. The aforementioned audio processing device / apparatus may also be implemented wholly or at least partially by computer device 600 or similar devices or systems.
[0094] Computer device 600 can be a variety of different types of devices, such as a service provider's server, a device associated with a client (e.g., a client device), a system-on-a-chip, and / or any other suitable computer device or computing system. Examples of computer device 600 include, but are not limited to: desktop computers, server computers, laptop or netbook computers, mobile devices (e.g., tablets, cellular or other wireless phones (e.g., smartphones), notebook computers, mobile stations), wearable devices (e.g., glasses, watches), entertainment devices (e.g., entertainment appliances, set-top boxes communicatively coupled to a display device, game consoles), televisions or other display devices, automotive computers, and so on. Therefore, the range of computer device 600 can be from full-resource devices with large amounts of memory and processor resources (e.g., personal computers, game consoles) to low-resource devices with limited memory and / or processing resources (e.g., traditional set-top boxes, handheld game consoles).
[0095] Computer device 600 may include at least one processor 602, memory 604, multiple communication interfaces 606, display device 608, other input / output (I / O) devices 610, and one or more mass storage devices 612 capable of communicating with each other, such as via system bus 614 or other suitable connections.
[0096] Processor 602 may be a single processing unit or multiple processing units, and all processing units may include single or multiple computing units or multiple cores. Processor 602 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any device that manipulates signals based on operating instructions. Among other capabilities, processor 602 may be configured to acquire and execute computer-readable instructions stored in memory 604, mass storage device 612, or other computer-readable media, such as program code of operating system 616, program code of application program 618, program code of other program 620, etc.
[0097] Memory 604 and mass storage device 612 are examples of computer-readable storage media for storing instructions executed by processor 602 to perform the various functions described above. For example, memory 604 may generally include both volatile and non-volatile memory (e.g., RAM, ROM, etc.). Furthermore, mass storage device 612 may generally include hard disk drives, solid-state drives, removable media, including external and removable drives, memory cards, flash memory, floppy disks, optical disks (e.g., CDs, DVDs), storage arrays, network-attached storage, storage area networks, etc. Both memory 604 and mass storage device 612 may be collectively referred to herein as memory or computer-readable storage media, and may be non-transitory media capable of storing computer-readable, processor-executable program instructions as computer program code, which may be executed by processor 602 as a specific machine configured to perform the operations and functions described in the examples herein.
[0098] Multiple program modules may be stored on mass storage device 612. These programs include operating system 616, one or more application programs 618, other programs 620, and program data 622, and they may be loaded into memory 604 for execution. Examples of such application programs or program modules may include, for example, computer program logic (e.g., computer program code or instructions) for implementing the following components / functions: method 200 and / or method 400 (including any suitable steps of method 200, 400), and / or other embodiments described herein.
[0099] Although Figure 6The modules 616, 618, 620, and 622, or portions thereof, are illustrated as being stored in memory 604 of computer device 600; however, modules 616, 618, 620, and 622 may be implemented using any form of computer-readable medium accessible by computer device 600. As used herein, “computer-readable medium” includes at least two types of computer-readable media: computer storage media and communication media.
[0100] Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, DVD, or other optical storage devices, magnetic cassettes, magnetic tapes, disk storage devices or other magnetic storage devices, or any other non-transfer medium that can be used to store information for access by computer equipment.
[0101] In contrast, communication media can embody computer-readable instructions, data structures, program modules, or other data within modulated data signals such as carrier waves or other transmission mechanisms. Computer storage media as defined herein do not include communication media.
[0102] Computer device 600 may also include one or more communication interfaces 606 for exchanging data with other devices, such as via a network, direct connection, etc., as discussed above. Such communication interfaces can be one or more of the following: any type of network interface (e.g., a network interface card (NIC)), wired or wireless (such as IEEE 802.11 Wireless LAN (WLAN)) wireless interface, Wi-MAX interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth. TM Interfaces, near field communication (NFC) interfaces, etc. Communication interface 606 can facilitate communication across various network and protocol types, including wired networks (e.g., LAN, cable, etc.) and wireless networks (e.g., WLAN, cellular, satellite, etc.), the Internet, etc. Communication interface 606 can also provide communication with external storage devices (not shown) such as storage arrays, network-attached storage, storage area networks, etc.
[0103] In some examples, a display device 608, such as a monitor, may be included for displaying information and images to the user. Other I / O devices 610 may be devices that receive various inputs from the user and provide various outputs to the user, and may include touch input devices, gesture input devices, cameras, keyboards, remote controls, mice, printers, audio input / output devices, and so on.
[0104] Although this disclosure has been described and illustrated in detail in the accompanying drawings and the foregoing description, such description and illustration should be considered illustrative and suggestive, not restrictive; this disclosure is not limited to the disclosed embodiments. By studying the drawings, the disclosure, and the appended claims, those skilled in the art will be able to understand and implement variations of the disclosed embodiments in practicing the claimed subject matter. In the claims, the word "comprising" does not exclude other elements or steps not listed, and the words "a" or "an" do not exclude a plurality. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be beneficial.
Claims
1. An audio processing method, comprising: Obtain at least one piece of identity information related to the first user; Obtain at least one audio segment associated with the first user, the at least one audio segment including at least one of the following: a first audio segment containing a cough sound made by the first user, a second audio segment containing a blowing sound made by the first user, or a third audio segment containing a vowel made by the first user. The at least one audio segment is processed to obtain a processed audio segment; Determining at least one predicted value for at least one lung function indicator of the first user based on the at least one piece of identity information includes: fitting the at least one piece of identity information to a linear regression model to obtain the at least one predicted value; and Based on the at least one predicted value and the processed audio segment, a machine learning model is used to determine the corresponding final predicted value of the at least one lung function indicator for the first user. The machine learning model includes an encoder part and a projection part, and the final predicted value of the at least one lung function index of the first user is determined by the machine learning model based on the at least one predicted value and the processed audio segment, which includes: encoding the processed audio segment by the encoder part to obtain a first encoded vector; concatenating the first encoded vector with the vectorized representation of the at least one predicted value to obtain a second encoded vector; and determining the corresponding final predicted value of the at least one lung function index based on the second encoded vector by the projection part.
2. The method according to claim 1, wherein, The at least one audio segment includes the first audio segment, and wherein processing the at least one audio segment to obtain a processed audio segment includes: A first background noise threshold is determined based on the first audio segment; Identify at least one valid cough peak from the first audio segment; and Based on the first background noise threshold and the at least one effective cough peak, at least one effective cough segment is obtained, wherein the processed audio segment includes the at least one effective cough segment.
3. The method according to claim 2, wherein, Identifying at least one valid cough peak from the first audio segment includes: Calculate the frequency domain energy distribution of the first audio segment; Identify amplitude points exceeding a predetermined effective cough threshold from the frequency domain energy distribution, the effective cough threshold being used to distinguish between cough sounds and throat clearing sounds; and The amplitude point exceeding the effective cough threshold is determined as the at least one effective cough peak.
4. The method according to any one of claims 1-3, wherein, The at least one audio segment includes the second audio segment, and wherein processing the at least one audio segment to obtain a processed audio segment further includes: Determine a second background noise threshold based on the second audio segment; and The second audio segment is truncated based on the second background noise threshold to obtain a processed second audio segment, wherein the processed audio segment includes the processed second audio segment.
5. The method according to any one of claims 1-3, wherein, The at least one audio segment includes the third audio segment, and wherein processing the at least one audio segment to obtain the processed audio segment further includes: A third background noise threshold is determined based on the third audio segment; and The third audio segment is truncated based on the third background noise threshold to obtain a processed third audio segment, wherein the processed audio segment includes the processed third audio segment.
6. The method according to any one of claims 1-3, further comprising, before obtaining at least one audio segment associated with the first user: Obtain the fourth audio segment; Determine the background noise level based on the fourth audio segment; and In response to determining that the background noise level exceeds a fourth background noise threshold, a prompt to change the measurement environment is output.
7. The method according to any one of claims 1-3, wherein, The at least one lung function indicator includes at least one of the following: forced vital capacity (FVC), forced expiratory volume in one second (FEV1), or the ratio of forced expiratory volume in one second to forced vital capacity (FEV1 / FVC).
8. An audio processing apparatus, comprising: An identity information acquisition unit is used to acquire at least one piece of identity information related to the first user; An audio segment acquisition unit is configured to acquire at least one audio segment related to the first user, the at least one audio segment including at least one of the following: a first audio segment containing a cough sound emitted by the first user, a second audio segment containing a blowing sound emitted by the first user, or a third audio segment containing a vowel emitted by the first user. An audio segment processing unit is used to process the at least one audio segment to obtain a processed audio segment; The first determining unit is configured to determine at least one predicted value for at least one lung function indicator of the first user based on the at least one piece of identity information, including: fitting the at least one piece of identity information to a linear regression model to obtain the at least one predicted value; and The second determining unit is configured to determine, based on the at least one predicted value and the processed audio segment, the corresponding final predicted value of the at least one lung function index of the first user using a machine learning model. The machine learning model includes an encoder part and a projection part, and the final predicted value of the at least one lung function index of the first user is determined by the machine learning model based on the at least one predicted value and the processed audio segment, which includes: encoding the processed audio segment by the encoder part to obtain a first encoded vector; concatenating the first encoded vector with the vectorized representation of the at least one predicted value to obtain a second encoded vector; and determining the corresponding final predicted value of the at least one lung function index based on the second encoded vector by the projection part.
9. A computing device, comprising: Memory, processor, and computer program stored on said memory, The processor is configured to execute the computer program to implement the steps of the method according to any one of claims 1-7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
11. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.