Breathing sound recognition method and device, equipment and storage medium

By collecting and analyzing respiratory audio data on consumer-grade devices and using AI models to identify gender, age and health status, it solves the problem of difficulty in collecting and diagnosing respiratory audio data in the home environment, and achieves low-cost and high-convenience respiratory disease measurement.

CN120199281APending Publication Date: 2025-06-24SHANGHAI JINSHENG COMM TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302257.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively collect and diagnose respiratory audio data in a home environment, and requires professional medical equipment, which is costly and inconvenient to popularize.

Method used

Consumer-grade equipment is used to collect respiratory audio data of the target object, and the gender, age characteristics and health assessment results of the target object are obtained through feature extraction and training AI model identification.

Benefits of technology

The collection and diagnosis of respiratory audio data in a home environment without professional medical equipment is required, reducing the cost of measuring respiratory diseases and improving convenience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199281A_ABST
    Figure CN120199281A_ABST
Patent Text Reader

Abstract

The invention provides a breath sound recognition method and device, equipment and a storage medium. The method is applied to consumption-level equipment, and the method comprises the following steps: collecting first breathing audio data of a target object; performing feature extraction on the first breathing audio data to determine a feature matrix; inputting the feature matrix into a trained AI model to recognize the physiological state of the target object to obtain a recognition result; wherein the recognition result comprises at least one of the gender of the target object, the age feature of the target object and the health assessment result of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to electronic technology, including but not limited to methods and devices for respiratory sound recognition, equipment, and storage media. Background Art

[0002] Respiratory sounds, abbreviated as lung sounds in medicine, are sounds generated through structures such as airways and alveoli during the gas circulation process. They are extremely important biological signals and can be used as important indicators to reflect the physiological and pathological characteristics of the lungs and diagnose lung-related diseases. Traditionally, doctors listen to the sounds of the human lungs through a stethoscope and identify abnormal conditions of human lung sounds based on experience. However, this method requires the participation of professional medical equipment and is not suitable for the collection and diagnosis of respiratory sounds in a home environment. Summary of the Invention

[0003] In a first aspect, an embodiment of this application provides a method for respiratory sound recognition. The method is applied to a consumer device and includes: collecting first respiratory audio data of a target object; extracting features from the first respiratory audio data to determine a feature matrix; inputting the feature matrix into a trained AI model to recognize the physiological state of the target object and obtain a recognition result; where the recognition result includes at least one of the gender of the target object, the age characteristics of the target object, and the health assessment result of the target object.

[0004] In a second aspect, an embodiment of this application provides a device for respiratory sound recognition. The device includes: a collection module configured to collect first respiratory audio data of a target object; a determination module configured to extract features from the first respiratory audio data to determine a feature matrix; an identification module configured to input the feature matrix into a trained AI model to recognize the physiological state of the target object and obtain a recognition result; where the recognition result includes at least one of the gender of the target object, the age characteristics of the target object, and the health assessment result of the target object.

[0005] In a third aspect, an embodiment of this application provides a consumer device, including a memory and a second processor. The memory stores a computer program that can run on the second processor, and when the second processor executes the program, it implements the method described in the first aspect.

[0006] In a fourth aspect, an embodiment of this application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the second processor, it implements the method described in the first aspect.

[0007] In a fifth aspect, an embodiment of this application provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by the second processor, it implements the method described in the first aspect of this application.

[0008] In a sixth aspect, an embodiment of the present application provides a computer program that causes a second processor to execute the method described in the first aspect.

[0009] In an embodiment of the present application, a consumer device is used to collect first respiratory audio data of a target object; feature extraction is performed on the first respiratory audio data to determine a feature matrix; then the feature matrix is input into a trained AI model to identify the physiological state of the target object, and the gender, age characteristics, and health assessment results of the target object are obtained. In this way, there is no need for the target object to collect professional medical equipment, nor is it necessary to deploy professional medical equipment in a home environment. Instead, a consumer device is used to collect the respiratory audio data of the target object, and based on the collected respiratory audio data, the gender, age characteristics, and health assessment results of the target object are obtained using a consumer device, thereby benefiting in a home environment by reducing the cost of measuring respiratory diseases and improving the convenience of measuring respiratory diseases.

[0010] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solutions of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0012] The flowcharts shown in the drawings are only exemplary illustrations, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0013] Figure 1 Schematic diagram of the implementation process of the respiratory sound recognition method provided by the embodiment of the present application;

[0014] Figure 2 Schematic diagram of the implementation process of determining the feature matrix provided by the embodiment of the present application;

[0015] Figure 3 Schematic diagram of the implementation process of determining the second respiratory audio data provided by the embodiment of the present application;

[0016] Figure 4 Schematic diagram of the implementation process of step 202 provided by the embodiment of the present application;

[0017] Figure 5 Schematic diagram of respiratory audio data provided by an embodiment of the present application;

[0018] Figure 6 Schematic diagram of a deep learning lung sound classification model provided by an embodiment of the present application;

[0019] Figure 7 Engineering framework flowchart of a respiratory sound recognition method provided by an embodiment of the present application;

[0020] Figure 8 Schematic structural diagram of a respiratory sound recognition device provided by an embodiment of the present application;

[0021] Figure 9 Schematic structure of a consumer device provided by an embodiment of the present application Figure 1 ;

[0022] Figure 10 Schematic structure of a consumer device provided by an embodiment of the present application Figure 2 。 Detailed implementation manners

[0023] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0025] In the following descriptions, references to "some embodiments", "this embodiment", "embodiments of the present application" and examples, etc., describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0026] The descriptions such as "first, second, third" etc. that appear in the embodiments of the present application are only for the purpose of indicating and distinguishing the described objects, without order, and do not represent special limitations on the number of devices in the embodiments of the present application, and cannot constitute any limitation to the embodiments of the present application.

[0027] To facilitate the understanding of the technical solutions of the embodiments of the present application, the following explains the relevant technologies or terms of the embodiments of the present application. The following relevant technologies or relevant terms can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.

[0028] Respiratory sounds, also known as lung sounds, are mainly used in the medical field to determine whether a patient has various respiratory diseases. Using respiratory sounds can measure the human body state. In a related technology, a professional medical device is placed outside the human lungs to listen to and record respiratory audio, and diagnose whether a person has lung diseases, colds, pneumonia, etc. In another related technology, multiple or single microphones are placed near the human head to simply identify whether a person is in a sleeping state. In yet another related technology, multiple microphones are used to identify the physiological state of a specific type of population, for family guardianship of children or the elderly, etc.

[0029] However, through research and analysis, the inventors of the present application found that the above-mentioned related technologies have the following problems:

[0030] (1) Most technologies require professional equipment, which is costly and limits the simple determination of respiratory diseases for ordinary consumer users at home.

[0031] (2) Using multiple microphones to detect the human sleeping or physiological state is technically feasible, but in a home environment, deploying multiple microphones around the bed itself does not have wide life feasibility.

[0032] Based on this, an embodiment of the present application provides a respiratory sound recognition method. Figure 1 It is a schematic flowchart of the implementation process of the respiratory sound recognition method provided by the embodiment of the present application. As Figure 1 shown, the method includes the following steps 101 to step 103:

[0033] Step 101, collect the first respiratory audio data of the target object;

[0034] Step 102, perform feature extraction on the first respiratory audio data to determine a feature matrix;

[0035] Step 103, input the feature matrix into a trained AI model to identify the physiological state of the target object, and obtain an identification result; wherein, the identification result includes at least one of the gender of the target object, the age feature of the target object, and the health assessment result of the target object.

[0036] It can be understood that in the embodiments of the present application, a consumer device is used to collect the first respiratory audio data of the target object; feature extraction is performed on the first respiratory audio data to determine a feature matrix; then the feature matrix is input into a trained AI model to identify the physiological state of the target object, and the gender, age characteristics, and health assessment results of the target object are obtained. In this way, there is no need for the target object to collect professional medical equipment, nor is it necessary to deploy professional medical equipment in a home environment. Instead, a consumer device is used to collect the respiratory audio data of the target object, and based on the collected respiratory audio data, the gender, age characteristics, and health assessment results of the target object are obtained using the consumer device. This is beneficial for reducing the cost of measuring respiratory diseases and improving the convenience of measuring respiratory diseases in a home environment.

[0037] The further optional implementation manners of the above steps, related terms, etc. are described below.

[0038] In step 101, the first respiratory audio data of the target object is collected.

[0039] It should be understood that in the embodiments of the present application, there is no limit to the length of the first respiratory audio data, nor is there a limit to the timing of collecting the first respiratory audio data of the target object. In some embodiments, the first respiratory audio data of the target object can be collected at any time; in other embodiments, the first respiratory audio data of the target object is collected when a trigger condition is met.

[0040] In the embodiments of the present application, there is also no limit to the trigger condition for collecting the first respiratory audio data of the target object. In some embodiments, collecting the respiratory audio data of the target object includes: collecting the respiratory audio data of the target object when a trigger condition is met; where the trigger condition includes at least one of the following: the consumer device approaches the respiratory part of the target object; the distance between the consumer device and the respiratory part of the target object is less than or equal to a first threshold; the gaze distance between the target object and the consumer device is less than or equal to a second threshold.

[0041] It can be understood that in the embodiments of the present application, when the consumer device approaches the respiratory part of the target object, or when the distance between the consumer device and the respiratory part of the target object is less than or equal to the first threshold, or when the gaze distance between the target object and the consumer device is less than or equal to the second threshold, the respiratory audio data of the target object is collected. That is to say, in the embodiments of the present application, the respiratory audio data of the target object is collected only when a certain trigger condition is met, rather than at any time. In this way, it is beneficial to reduce the power consumption of the consumer device for collecting the first respiratory audio data.

[0042] It should be understood that in the embodiments of the present application, the consumer devices are not limited. Consumer devices are relative to professional devices and refer to electronic devices designed, produced, and sold for ordinary consumers, which are usually used in fields such as daily life, entertainment, and health management. In some embodiments, the consumer devices include, but are not limited to, at least one of the following: smart phones, tablets, wearable devices, smart home devices, etc.

[0043] In the embodiments of the present application, the first threshold and the second threshold are not limited. The first threshold and the second threshold may be the same or different. In some embodiments, the second threshold is 25 cm. In the embodiments of the present application, the breathing part of the target object is also not limited. The breathing part refers to the organs and tissues of the human body involved in the breathing process. In some embodiments, the breathing part includes, but is not limited to, at least one of the following: nasal cavity, pharynx, larynx, trachea, etc.

[0044] In step 102, feature extraction is performed on the first breathing audio data to determine a feature matrix.

[0045] In some embodiments, the first breathing audio data is preprocessed to obtain fifth breathing audio data, and feature extraction is performed on the fifth breathing audio data to determine a feature matrix.

[0046] It should be understood that in the embodiments of the present application, the specific implementation manner of the preprocessing is not limited. In some embodiments, the preprocessing includes: performing signal enhancement and noise reduction on the first breathing audio data to obtain fifth breathing audio data. In the embodiments of the present application, the specific implementation manner of performing feature extraction on the first breathing audio data or the fifth breathing audio data to determine a feature matrix is not limited. In some embodiments, the first breathing audio data or the fifth breathing audio data is subjected to feature extraction through methods such as short-time fast Fourier transform and triangular filters, Mel Frequency Cepstral Coefficients (MFCC), spectrogram, and Q-factor wavelet transform to obtain a feature matrix.

[0047] In some embodiments, Figure 2 is a schematic diagram of the implementation process for determining the feature matrix provided by the embodiments of the present application. As Figure 2 shown, the feature matrix can be determined through the following steps 201 to 202:

[0048] Step 201: Filter the first breathing audio data using a filter to determine second breathing audio data;

[0049] Step 202: Perform Fourier transform on the second breathing audio data to determine the feature matrix.

[0050] It should be understood that in the embodiments of the present application, the filter is not limited. In some embodiments, the filter includes but is not limited to at least one of the following: low-pass filter, high-pass filter, band-pass filter, band-stop filter, notch filter, and adaptive filter.

[0051] In the embodiments of the present application, the Fourier transform is not limited. The Fourier transform is a mathematical tool for converting a signal from the time domain to the frequency domain. Its core idea is that any complex signal can be decomposed into a superposition of a series of sine waves or cosine waves with different frequencies. In some embodiments, the Fourier transform may be a discrete Fourier transform. In other embodiments, the Fourier transform may be a short-time Fourier transform.

[0052] In some embodiments, Figure 3 is a schematic diagram of the implementation process for determining the second respiratory audio data provided by the embodiments of the present application. As Figure 3 shown, the second respiratory audio data can be determined through the following steps 301 to 302:

[0053] Step 301, filtering the first respiratory audio data by using a filter with a high Q factor to obtain third respiratory audio data;

[0054] Step 302, filtering the first respiratory audio data by using a filter with a low Q factor to obtain fourth respiratory audio data;

[0055] Step 303, determining the second respiratory audio data based on the third respiratory audio data, the fourth respiratory audio data, and the first respiratory audio data.

[0056] In some embodiments, the determining the second respiratory audio data based on the third respiratory audio data, the fourth respiratory audio data, and the first respiratory audio data includes: determining a first result based on the third respiratory audio data and the fourth respiratory audio data; and determining the second respiratory audio data based on the first respiratory audio data and the first result.

[0057] Exemplarily, in a possible implementation manner, the first respiratory audio data is respectively filtered by using a filter with a high Q factor and a filter with a low Q factor, and the Q-factor residual of the filtering result is solved.

[0058] It should be understood that in the embodiments of the present application, the Q factor is used to define the ratio of the filter bandwidth to the center frequency. The higher the Q value, the narrower the filter bandwidth and the higher the frequency resolution. The frequency response of a filter with a high Q factor is narrow, which can effectively suppress background noise and interference signals, thereby improving the signal-to-noise ratio and clarity of the signal. The frequency response of a filter with a low Q factor is wide, which can smooth the signals in a relatively wide frequency range, making the signals more stable and continuous. The Q factor residual extracts the signals that are not captured by the filters with high Q factors and low Q factors by comparing the outputs of the filters with high Q factors and the outputs of the filters with low Q factors.

[0059] In some embodiments, Figure 4 is a schematic flowchart for implementing step 202 provided by the embodiments of the present application. As Figure 4 shown, step 202 can be implemented through the following steps 401 to step 402:

[0060] Step 401: Perform windowing processing on the second respiratory audio data to obtain multiple windowed respiratory audio data;

[0061] Step 402: Perform Fourier transform on the multiple windowed respiratory audio data respectively to obtain the spectral matrix of the first respiratory audio data;

[0062] Step 403: Stack N of the spectral matrices to obtain the feature matrix; where N is greater than or equal to 1.

[0063] It should be understood that in the embodiments of the present application, no further limitation is imposed on the implementation manner of step 401. In some embodiments, the second respiratory audio data can be windowed through a rectangular window, a Hamming window, a Hanning window, etc. to obtain multiple windowed respiratory audio data. In the embodiments of the present application, no limitation is imposed on N. In a possible implementation manner, N is equal to 3.

[0064] In step 103, the feature matrix is input into the trained AI model to identify the physiological state of the target object, and an identification result is obtained; where the identification result includes at least one of the gender of the target object, the age feature of the target object, and the health assessment result of the target object.

[0065] It should be understood that in the embodiments of the present application, no limitation is imposed on the age feature of the target object. In some embodiments, the age feature may be the age value of the target object, and the age feature may also be the age range of the target object, etc. In the embodiments of the present application, no limitation is imposed on the health assessment result of the target object either. In some embodiments, the health assessment result of the target object includes whether the target object is healthy, the disease type of the target object, etc.

[0066] In the embodiments of the present application, the AI model is not limited. In some embodiments, the AI model includes a machine learning model or a deep learning model; wherein, the machine learning model includes, but is not limited to, at least one of the following: logistic regression model, decision tree, random forest, naive Bayes, K-nearest neighbor, etc.; the deep learning model includes, but is not limited to, at least one of the following: recurrent convolutional neural network, generative adversarial convolutional neural network, long short-term memory network, feedforward neural network, multi-layer perceptron, convolutional neural network including attention mechanism, deep belief network, etc.

[0067] In some embodiments, the trained AI model includes a trained support vector machine.

[0068] In some embodiments, the trained AI model includes a trained lightweight neural network model.

[0069] In some embodiments, the trained lightweight neural network model is obtained by training a lightweight neural network model; in other embodiments, the trained lightweight neural network model is obtained by pruning and / or quantifying a neural network model after training.

[0070] In some embodiments, inputting the feature matrix into the trained AI model to identify the physiological state of the target object and obtain an identification result includes: inputting the feature matrix into a shallow feature extraction module in the trained multi-residual block deep learning lung sound model to perform primary feature extraction on the feature matrix of the first respiratory audio data to obtain shallow features; inputting the shallow features into a multi-residual module in the trained multi-residual block deep learning lung sound model to perform deep feature extraction on the feature matrix of the first respiratory audio data to obtain deep features; inputting the deep features into an identification module in the trained multi-residual block deep learning lung sound model to identify the physiological state of the target object and obtain an identification result.

[0071] It should be understood that in the embodiments of the present application, the multi-residual module is not limited. In some embodiments, the multi-residual module includes an original residual block (ResNetⅠ), an improved residual block (ResNetⅡ), and a residual block combined with an attention mechanism (Attention-Augmented ResNet, AAResNet).

[0072] In some embodiments, the feature matrix is input into a shallow feature extraction module in a trained multi-residual block deep learning lung sound model to perform primary feature extraction on the feature matrix of the first respiratory audio data, obtaining shallow features, including: successively performing convolution, normalization, rectified linear unit (ReLU), and pooling operations on the feature matrix to obtain shallow features.

[0073] In some embodiments, the deep features are input into an identification module in a trained multi-residual block deep learning lung sound model to identify the physiological state of the target object, obtaining an identification result, including: successively performing normalization, rectified linear unit (ReLU), pooling, first fully connected, and second fully connected operations on the deep features to obtain the identification result.

[0074] In some embodiments, the training process of the trained lightweight neural network model includes: training an initial lightweight neural network model based on lung sound sample data to obtain a first neural network model; wherein, the lung sound sample data is obtained by collecting lung audio of multiple objects through a stethoscope or a microphone array; training the first neural network model based on respiratory audio sample data to obtain the trained lightweight neural network model; wherein, the respiratory audio sample data is obtained by collecting respiratory audio of multiple objects through a consumer device.

[0075] It can be understood that in the embodiments of the present application, first, an initial lightweight neural network model is trained based on lung sound sample data obtained by collecting lung audio of multiple objects through a stethoscope or a microphone array, obtaining a first neural network, such that the first neural network learns the general features of the lung sound sample data. Then, based on respiratory audio sample data obtained by collecting respiratory audio of multiple objects through a consumer device, the first neural network is fine-tuned, so that the trained lightweight neural network model has the ability to identify respiratory audio data collected by a consumer device.

[0076] In some embodiments, the training process of the trained lightweight neural network model includes: training an initial neural network model based on lung sound sample data to obtain a second neural network model; wherein, the lung sound sample data is obtained by collecting lung audio of multiple objects through a stethoscope or a microphone array; training the second neural network model based on respiratory audio sample data to obtain the trained neural network model; wherein, the respiratory audio sample data is obtained by collecting respiratory audio of multiple objects through a consumer device; pruning and / or quantizing the trained neural network model to obtain the trained lightweight neural network model.

[0077] In some embodiments, the trained lightweight neural network model is deployed on the first processor of the consumer device; wherein, the first processor includes a low-power processor.

[0078] It can be understood that in the embodiments of the present application, the trained lightweight neural network model is deployed on the low-power processor of the consumer device, thereby helping to reduce the power consumption for determining the recognition result of the target object based on the collected first respiratory audio data.

[0079] In some embodiments, the consumer device includes a large-core processor and a small-core processor; wherein, the power consumption of the large-core processor is greater than that of the small-core processor. The small-core processor focuses on low power consumption and high energy efficiency and is suitable for processing lightweight tasks, while the large-core processor focuses on high performance and is suitable for processing complex tasks.

[0080] The following gives examples to describe possible implementation solutions of the respiratory sound recognition method described in one or more of the above embodiments.

[0081] The embodiments of the present application provide a respiratory sound recognition method. This method uses a consumer device (such as a mobile phone, tablet, earphone microphone, etc.) to collect the user's respiratory audio (i.e., an example of the first respiratory audio data), and analyzes the first respiratory audio data by using methods such as signal processing and machine learning to identify the user's age stage (i.e., an example of age characteristics), gender, whether there is a respiratory disease (i.e., an example of a health assessment result), etc. At the same time, being deployed on a low-power platform, it can be continuously mounted or triggered under certain conditions to identify the result to be recognized with low power consumption.

[0082] In some embodiments, Figure 5 is a schematic diagram of the respiratory audio data provided by the embodiments of the present application, as Figure 5 shown, the respiratory audio data is the respiratory audio data collected by an ordinary consumer device placed at a distance of 20 cm to 25 cm from the human body. Among them, 501 represents the collected respiratory audio, and 502 represents the normal voice conversation audio. In some embodiments, the respiratory audio represented by 501 is used as a research sample, and its length is generally 3 seconds to 5 seconds.

[0083] In some embodiments, after extracting the respiratory audio represented by 501, feature extraction is performed as follows: First, signal enhancement is performed on the upper and lower signals (i.e., an example of the first respiratory audio data) respectively to enhance the intensity of the respiratory audio data; then, the respiratory audio data after signal enhancement preprocessing is subjected to feature extraction through methods such as short-time fast Fourier transform, triangular filters, MFCC, spectrogram, and Q-factor wavelet transform to obtain a feature matrix; a pulmonary sound classification model is used to extract the feature matrix to obtain a classification result; among them, the pulmonary sound classification model (including classical machine learning models or deep learning models, i.e., an example of an AI model) is obtained through pre-training and tuning for different classification targets. The training process of the pulmonary sound classification model includes: pre-training the pulmonary sound classification model through a large number of existing open-source pulmonary sound data (i.e., an example of pulmonary sound sample data). Then, specific classification target respiratory data (such as age classification, gender classification, whether having a cold classification, etc.) collected manually (or crowdsourcing data with personal information removed or removed by differential privacy algorithms in the future) (i.e., an example of respiratory audio sample data) is used to tune the pre-trained pulmonary sound classification model, and finally, a target classification model (i.e., a trained pulmonary sound classification model, also a trained AI model) is obtained.

[0084] Exemplarily, in a possible implementation, the first respiratory audio data is filtered respectively through a high-Q-factor filter and a low-Q-factor filter, and Q-factor residuals of the filtering results are solved, and then short-time Fourier transform is performed three times in parallel to solve the feature map. Using the ordinary Support Vector Machine (SVM) classification method, the classification result of the first respiratory audio data is obtained. For relatively clearly collected respiratory audio data, a good recognition effect can be achieved (for example: in the age task, the recognition accuracy for the elderly in the age range of 70-79 exceeds 90%). Or using a deep learning pulmonary sound classification model (i.e., an example of a neural network model) to obtain the classification result of the respiratory audio data. Among them, the deep learning pulmonary sound classification model needs to be pre-trained using open-source pulmonary sound data (i.e., an example of pulmonary sound sample data), and then fine-tuned using respiratory audio sample data. For respiratory audio with non-standard acquisition paradigms, a good recognition effect can also be achieved (for example: in the age task, the recognition for the elderly in the age range of 70-79).

[0085] Figure 6Schematic diagram of the deep learning lung sound classification model provided by the embodiments of this application. Among them, Conv2D(64) represents a two-dimensional convolutional layer using 64 convolutional kernels, GN represents group normalization, ReLU represents the rectified linear unit, MaxPool2D represents a two-dimensional max pooling layer, ResNetⅠ represents the original residual network, ResNetⅡ represents the improved residual network, AAResNet represents the residual network combined with the attention mechanism, GAP represents global average pooling, FC(64) represents a fully connected layer containing 64 neurons, and FC(4) represents a fully connected layer containing 4 neurons.

[0086] At the same time, the breath sound detection model (i.e., an example of a trained lightweight neural network model) is deployed on a low-power AON platform, which can continuously detect the user's breath sound state with low power consumption. In order to achieve the goal of breath sound detection with even lower power consumption, various preconditions such as raising the hand and looking closely can also be designed to achieve the purpose of fast perception with low power consumption. Among them, AON, short for Always On, is a technology or functional characteristic that enables a device or system to continuously run certain specific tasks or functions without waking up the main processor (such as the CPU). AON technology is usually used in low-power devices to achieve long standby time and instant response.

[0087] Figure 7 Engineering framework flowchart of the breath sound recognition method provided by the embodiments of this application, as Figure 7 shown, the audio receiver inputs the collected original signal into the Sensor Driver, and the Sensor Driver converts the collected original signal into data and inputs this data into the perception model (i.e., an example of an AI model) to obtain a classification result, which is output at the Android APP layer. Among them, the inference framework is the bridge between the model and the hardware, responsible for the deployment and optimization of the model. The NPU is the hardware foundation for efficiently executing model calculations, providing high-energy-efficiency AI computing capabilities; 701 represents the Sensorhub.

[0088] In some embodiments, the trained perception model needs to be deployed and optimized through the inference framework. The inference framework provides tools and interfaces to convert the perception model into a format that can run on the target hardware. The inference framework can also optimize the perception model (such as quantization and pruning) to improve the running efficiency. The inference framework is responsible for mapping the perception model to the NPU for execution. The inference framework needs to support the dedicated instruction set and computing characteristics of the NPU to fully utilize the performance of the NPU. The NPU is more focused on the execution of neural network-related algorithms and tasks, and can provide higher computing efficiency and lower power consumption.

[0089] In some embodiments, the Sensor Driver is a bridge connecting the sensor hardware to the upper-layer application or operating system. It is responsible for initializing the sensor hardware, configuring sensor parameters, reading sensor data, and transmitting this data to the upper layer in a specific format. At the same time, the Sensor Driver is also responsible for preprocessing tasks such as calibrating and filtering sensor data to improve the accuracy and reliability of the data.

[0090] An embodiment of the present application provides a method for recognizing respiratory audio. This method can quickly sense and detect the physiological state reflected by the user's breathing sound using a consumer device, with low cost and simple use. It is an effective means for detecting the physiological state in some scenarios where users do not want to make sounds due to privacy compliance issues.

[0091] It should be noted that although the steps of the method in the present application are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution; or, steps in different embodiments may be combined into a new technical solution.

[0092] Based on the foregoing embodiments, an embodiment of the present application provides a respiratory sound recognition device. The device includes each module included and each unit included in each module, and can be implemented by a processor; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be an AI acceleration engine (such as an NPU, etc.), GPU, central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP), or field programmable gate array (FPGA), etc.

[0093] Figure 8 It is a schematic structural diagram of the respiratory sound recognition device provided by the embodiment of the present application, as Figure 8 shown, the respiratory sound recognition device 800 includes an acquisition module 801, a determination module 802, and a recognition module 803, where:

[0094] The acquisition module 801 is configured to acquire first respiratory audio data of a target object;

[0095] The determination module 802 is configured to perform feature extraction on the first respiratory audio data to determine a feature matrix;

[0096] The recognition module 803 is configured to input the feature matrix into a trained AI model to recognize the physiological state of the target object and obtain a recognition result; wherein, the recognition result includes at least one of the gender of the target object, the age feature of the target object, and the health assessment result of the target object.

[0097] In some embodiments, the determination module 802 is configured to filter the first respiratory audio data using a filter to determine second respiratory audio data; perform a Fourier transform on the second respiratory audio data to determine the feature matrix.

[0098] In some embodiments, the determination module 802 is configured to filter the first respiratory audio data using a filter with a high Q factor to obtain third respiratory audio data; filter the first respiratory audio data using a filter with a low Q factor to obtain fourth respiratory audio data; determine the second respiratory audio data based on the third respiratory audio data, the fourth respiratory audio data, and the first respiratory audio data.

[0099] In some embodiments, the determination module 802 is configured to perform windowing processing on the second respiratory audio data to obtain a plurality of windowed respiratory audio data; perform a Fourier transform on each of the plurality of windowed respiratory audio data to obtain a spectral matrix of the first respiratory audio data; stack N of the spectral matrices to obtain the feature matrix; wherein, N is greater than or equal to 1.

[0100] In some embodiments, the trained AI model includes a trained lightweight neural network model.

[0101] In some embodiments, the training process of the trained lightweight neural network model includes: training an initial lightweight neural network model based on lung sound sample data to obtain a first neural network model; wherein, the lung sound sample data is obtained by collecting lung audio of multiple objects using a stethoscope or a microphone array; training the first neural network model based on respiratory audio sample data to obtain the trained lightweight neural network model; wherein, the respiratory audio sample data is obtained by collecting respiratory audio of multiple objects using a consumer device.

[0102] In some embodiments, the trained lightweight neural network model is deployed on a first processor of the consumer device; wherein, the first processor includes a low-power processor.

[0103] In some embodiments, the acquisition module 801 is configured to acquire the breathing audio data of the target object when a triggering condition is met; wherein, the triggering condition includes at least one of the following: the consumer device approaches the breathing part of the target object; the distance between the consumer device and the breathing part of the target object is less than or equal to a first threshold; the gazing distance between the target object and the consumer device is less than or equal to a second threshold.

[0104] The description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0105] It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, may exist separately physically, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware, or in the form of a software functional unit, or in the form of a combination of software and hardware.

[0106] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a consumer device to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0107] The embodiments of the present application provide a consumer device, Figure 9 which is a schematic structure of the consumer device provided by the embodiments of the present application Figure 1 , as Figure 9 shown, the consumer device 90 includes a memory 901 and a second processor 902. The memory 901 stores a computer program that can run on the second processor 902, and when the second processor 902 executes the program, it implements the steps in the method provided in the above embodiments.

[0108] It should be noted that the memory 901 is configured to store instructions and applications executable by the second processor 902, and can also cache data to be processed or already processed by each module in the second processor 902 and the consumer device 90 (such as image data, audio data, voice communication data, and video communication data), which can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0109] In some embodiments, Figure 10 is a schematic structural diagram of the consumer device provided by the embodiment of the present application Figure 2 , as Figure 10 shown, the consumer device further includes a first processor 1001, the first processor 1001 includes a low-power processor, and a trained lightweight neural network model is deployed on the first processor 1001.

[0110] The embodiment of the present application also provides a computer-readable storage medium for storing a computer program.

[0111] Optionally, the computer-readable storage medium can be applied to the consumer device in the embodiment of the present application, and the computer program causes the second processor to execute the various methods of the embodiment of the present application. For the sake of brevity, it will not be elaborated here.

[0112] The embodiment of the present application also provides a computer program product including computer program instructions.

[0113] Optionally, the computer program product can be applied to the consumer device in the embodiment of the present application, and the computer program instructions cause the second processor to execute the various methods of the embodiment of the present application. For the sake of brevity, it will not be elaborated here.

[0114] The embodiment of the present application also provides a computer program.

[0115] Optionally, the computer program can be applied to the consumer device in the embodiment of the present application. When the computer program runs on the second processor, it causes the second processor to execute the various methods of the embodiment of the present application. For the sake of brevity, it will not be elaborated here.

[0116] It should be pointed out here that: the descriptions of the above consumer device, storage medium, computer program product, and computer program embodiment are similar to the descriptions of the above method embodiment, and have beneficial effects similar to those of the method embodiment. For the technical details not disclosed in the consumer device, storage medium, computer program product, and computer program embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0117] It should be understood that the "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" or "in some embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the sequence of execution, and the execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments. The descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated herein.

[0118] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, object A and / or object B can represent: object A exists alone, object A and object B exist simultaneously, and object B exists alone.

[0119] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0120] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The above-described embodiments are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling or communication connection between the components shown or discussed with each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.

[0121] The modules described above as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network elements; some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0122] In addition, in each embodiment of the present application, all the functional modules may be integrated into one processing unit, or each module may be a separate unit alone, or two or more modules may be integrated into one unit; the above-mentioned integrated modules may be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

[0123] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, read-only memory (ROM), magnetic disks, or optical disks and other various media that can store program codes.

[0124] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application essentially or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a consumer device to execute all or part of the methods described in the various embodiments of the present application. And the foregoing storage medium includes: removable storage devices, ROM, magnetic disks, or optical disks and other various media that can store program codes.

[0125] The methods disclosed in several method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.

[0126] The features disclosed in several product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.

[0127] The features disclosed in several method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0128] As described above, it is only the implementation mode of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claimed rights.

Claims

1. A method for identifying respiratory sounds, characterized in that: The method is applied to a consumer-level device, and the method comprises: Collecting first breathing audio data of the target object; Performing feature extraction on the first breathing audio data to determine a feature matrix; The feature matrix is ​​input into a trained AI model to identify the physiological state of the target object and obtain an identification result; wherein the identification result includes at least one of the gender of the target object, the age characteristic of the target object, and the health assessment result of the target object.

2. The method according to claim 1, characterized in that The step of extracting features from the first breathing audio data to determine a feature matrix includes: Using a filter to filter the first breathing audio data to determine second breathing audio data; Perform Fourier transform on the second breathing audio data to determine the feature matrix.

3. The method according to claim 2, characterized in that The step of filtering the first breathing audio data using a filter to determine the second breathing audio data includes: Using a high Q factor filter to filter the first breathing audio data to obtain third breathing audio data; Using a low Q factor filter to filter the first breathing audio data to obtain fourth breathing audio data; The second breathing audio data is determined based on the third breathing audio data, the fourth breathing audio data, and the first breathing audio data.

4. The method according to claim 2 or 3, characterized in that: The performing Fourier transform on the second breathing audio data to determine the feature matrix includes: Performing windowing processing on the second breathing audio data to obtain a plurality of window breathing audio data; Performing Fourier transform on the plurality of window breathing audio data respectively to obtain a frequency spectrum matrix of the first breathing audio data; The N frequency spectrum matrices are stacked to obtain the feature matrix; wherein N is greater than or equal to 1.

5. The method according to claim 1, characterized in that The trained AI model includes a trained lightweight neural network model.

6. The method according to claim 5, characterized in that The training process of the trained lightweight neural network model includes: Training an initial lightweight neural network model based on lung sound sample data to obtain a first neural network model; wherein the lung sound sample data is obtained by collecting lung audio of multiple subjects through a stethoscope or a microphone array; The first neural network model is trained based on the breathing audio sample data to obtain the trained lightweight neural network model; wherein the breathing audio sample data is obtained by collecting breathing audio of multiple objects through consumer-grade devices.

7. The method according to claim 5 or 6, characterized in that: The trained lightweight neural network model is deployed on a first processor of the consumer device; wherein the first processor includes a low-power processor.

8. The method according to claim 1, characterized in that: The step of collecting breathing audio data of the target object includes: Collecting breathing audio data of the target object when a trigger condition is met; The trigger condition includes at least one of the following: The consumer-grade device approaches the respiratory part of the target object; The distance between the consumer device and the breathing part of the target object is less than or equal to a first threshold; The gaze distance between the target object and the consumer-grade device is less than or equal to a second threshold.

9. A breathing sound recognition device, characterized in that: The device comprises: A collection module configured to collect first breathing audio data of a target object; A determination module, configured to perform feature extraction on the first breathing audio data to determine a feature matrix; The recognition module is configured to input the feature matrix into a trained AI model to identify the physiological state of the target object and obtain a recognition result; wherein the recognition result includes at least one of the gender of the target object, the age characteristic of the target object and the health assessment result of the target object.

10. A consumer-grade device, comprising a memory and a second processor, wherein the memory stores a computer program executable on the second processor, and the second processor implements the method according to any one of claims 1 to 8 when executing the program.

11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a second processor.

12. A computer program product, comprising a computer program or instructions, wherein when the computer program or instructions are executed by a second processor, the method according to any one of claims 1 to 8 is implemented.