A method of creating a lung sound database and related devices

By collecting, preprocessing, labeling, and storing lung sound data, the problems of scarce quantity and low quality of lung sound databases have been solved, and data consistency and usability have been achieved.

CN116578545BActive Publication Date: 2025-12-12JIANGHAN UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310309877.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-27
Publication Date
2025-12-12
Estimated Expiration
2043-03-27

AI Technical Summary

Technical Problem

Traditional lung sound signals cannot be saved, making it difficult to reproduce the state of the respiratory system. Lung sound databases are scarce and of low quality, lacking standardized collection and annotation, and are difficult to use effectively.

Method used

Lung sound data from multiple users were collected, preprocessed for noise reduction and heart sound elimination, labeled and named, and stored according to tissue structure to form a unified lung sound database.

Benefits of technology

It ensured the consistency and searchability of data samples in the lung sound database, standardized the data storage format, and facilitated subsequent use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116578545B_ABST
    Figure CN116578545B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a lung sound database creation method and related equipment, which can not only ensure the consistency of data samples in the lung sound database, but also standardize the information that should be saved and the organization form how to save the lung sound data to facilitate subsequent searching and utilization. The method comprises the following steps: collecting lung sound data corresponding to each user in a plurality of users; preprocessing the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user in a unified format; labeling the target lung sound data corresponding to each user; naming the target lung sound data corresponding to each user after labeling according to a naming rule of data samples in the lung sound database to obtain lung sound data samples corresponding to each user; and storing the lung sound data samples corresponding to each user based on an organization structure of the lung sound database, wherein the lung sound database comprises data samples of lung patients and data samples of a healthy control group.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of database, in particular to a method for creating a lung sound database and related equipment. BACKGROUND

[0002] There is a large amount of physiological information in lung sound, which can reflect the health status of human body, especially the respiratory system. Traditional lung sound auscultation, as an important part of clinical examination, is performed by a doctor through a stethoscope to capture the physiological signals generated by the human body during the breathing process, to knock the patient's chest and back, to evaluate the changes in the vibration of the sound passing through the chest cavity, and to reflect the physiological conditions of the patient's lungs and airways. The results of auscultation provide a reliable basis for the clinical diagnosis of respiratory diseases. However, the lung sound signals obtained by traditional auscultation cannot be saved, which is not convenient for reproducing the respiratory system status at that time, and is not convenient for discussion, analysis, research and summary after auscultation.

[0003] After the emergence of electronic stethoscopes, the collected lung sound signals can be saved. However, the establishment of lung sound database still faces many difficulties and problems. The number of lung sound databases, the number and quality of samples in the database are not only far behind the electrocardiogram database, but also far behind the heart sound database.

[0004] The lung sound signal itself is relatively weak, and is mixed with external environmental noise and human body internal noise, especially the much stronger heart sound, so that the number of valuable lung sound signal samples is small. The collection of lung sound signals and the establishment of lung sound database are not standardized. When collecting lung sound signals, they are easily mixed with other different types of diseases, and it is not easy to find useful data samples. Further, this leads to the scarcity of lung sound databases, and most of the databases contain a small number of cases and samples, and the number of high-quality samples is even smaller. SUMMARY

[0005] The embodiments of the present application provide a method for creating a lung sound database and related equipment, which can not only ensure the consistency of data samples in the lung sound database, but also standardize the information that should be saved and the organization form of how to save the lung sound data to facilitate subsequent searching and utilization.

[0006] The first aspect of the embodiments of the present application provides a method for creating a lung sound database, which comprises:

[0007] Collecting lung sound data corresponding to each user in a plurality of users;

[0008] Preprocessing the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user in a unified format;

[0009] Labeling the target lung sound data corresponding to each user;

[0010] According to a naming rule of data samples in the lung sound database, the target lung sound data corresponding to each user after labeling is named to obtain lung sound data samples corresponding to each user.

[0011] The lung sound data samples corresponding to each user are stored based on an organization structure of the lung sound database to complete creation of the lung sound database, the lung sound database including data samples of lung patients and data samples of a healthy control group.

[0012] The second aspect of the present application provides a lung sound database creation device, comprising:

[0013] A collection unit is configured to collect lung sound data corresponding to each user in a plurality of users.

[0014] A preprocessing unit is configured to preprocess the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user in a uniform format.

[0015] A labeling unit is configured to label the target lung sound data corresponding to each user.

[0016] A naming unit is configured to name the target lung sound data corresponding to each user after labeling according to a naming rule of data samples in the lung sound database to obtain lung sound data samples corresponding to each user.

[0017] A storage unit is configured to store the lung sound data samples corresponding to each user based on an organization structure of the lung sound database to complete creation of the lung sound database, the lung sound database including data samples of lung patients and data samples of a healthy control group.

[0018] The third aspect of the present application provides an electronic device, comprising a memory and a processor, the processor being configured to execute a computer management program stored in the memory to implement the steps of the lung sound database creation method of the first aspect.

[0019] The fourth aspect of the present application provides a computer readable storage medium having a computer management program stored thereon, the computer management program being executed by a processor to implement the steps of the lung sound database creation method of the first aspect.

[0020] In summary, it can be seen that, in the embodiment provided by the present application, the lung sound data corresponding to each user in multiple users is collected, and the target lung sound data in a unified format is obtained by preprocessing the lung sound data, then the target lung sound data is labeled, and the lung sound data is named according to the naming rule of the data sample in the database, and meanwhile, the lung sound data sample corresponding to each user is stored based on the organization structure of the lung sound database. Therefore, the consistency of the data sample in the lung sound database can be ensured, and the information that should be saved and the organization form in which the information should be saved of the lung sound data are standardized to facilitate subsequent searching and utilization. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 A flowchart of a lung sound database creation method provided by the embodiment of the present application is shown in the figure.

[0022] Figure 2 A system framework diagram of a lung sound signal acquisition system provided by the embodiment of the present application is shown in the figure.

[0023] Figure 3 A lung sound signal acquisition site diagram provided by the embodiment of the present application is shown in the figure.

[0024] Figure 4 A comparison diagram of effects before and after denoising of lung sound signals provided by the embodiment of the present application is shown in the figure.

[0025] Figure 5 A diagram of eliminating heart sounds in lung sound signals by using a multi-subdomain product method of wavelet transform coefficients provided by the embodiment of the present application is shown in the figure.

[0026] Figure 6 A diagram of eliminating heart sounds in lung sound signals by using a multi-subdomain product method of short-time Fourier transform coefficients provided by the present application is shown in the figure.

[0027] Figure 7 A diagram of heart sound suppression in lung sound signals provided by the embodiment of the present application is shown in the figure.

[0028] Figure 8 A diagram of results of recovering lung sound after removing heart sound provided by the embodiment of the present application is shown in the figure.

[0029] Figure 9 A diagram of an organization structure in a lung sound database provided by the embodiment of the present application is shown in the figure.

[0030] Figure 10 A virtual structure diagram of a lung sound database creation device provided by the embodiment of the present application is shown in the figure.

[0031] Figure 11 A hardware structure diagram of a lung sound database creation device provided by the embodiment of the present application is shown in the figure.

[0032] Figure 12An embodiment schematic diagram of an electronic device provided by an embodiment of the present application is shown in the figure;

[0033] Figure 13 An embodiment schematic diagram of a computer readable storage medium provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0034] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0035] In the following description, specific embodiments of the present application will be described with reference to steps and symbols executed by one or more computers, unless otherwise specified. Therefore, these steps and operations will be mentioned several times by computers, and computer execution referred to herein includes operations of a computer processing unit represented by electronic signals in a structured form. The operation transforms the data or maintains it at a location in the memory system of the computer, which can reconfigure or otherwise change the operation of the computer in a manner known to those skilled in the art. The data structure maintained by the data is the physical location of the memory, which has specific characteristics defined by the data format. However, the principles of the present application are described in the above description, which does not represent a limitation, and those skilled in the art will understand that the various steps and operations described below can also be implemented in hardware.

[0036] The terms "first", "second", and "third" and the like in the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion.

[0037] The method for creating a lung sound database will be described below from a lung sound database creation device. The lung sound database creation device can be a server or a service unit in the server, and the specific implementation is not limited.

[0038] Please refer to Figure 1 , Figure 1 The flowchart of the method for creating a lung sound database provided by an embodiment of the present application is shown in the figure, which includes:

[0039] 101, collect lung sound data corresponding to each user in a plurality of users.

[0040] In this embodiment, the lung sound database creation device can collect lung sound data corresponding to each user in a plurality of users through a lung sound signal collection system. The lung sound signal collection system will be described below. Figure 2Description of the lung sound signal acquisition system:

[0041] Please see Figure 2 , Figure 2 This is a system framework diagram of the lung sound signal acquisition system provided in an embodiment of the present invention. The lung sound signal acquisition system includes sensors, front-end circuits, signal acquisition and holding circuits, data storage units, and display units. The lung sound signal acquisition system can be a single-channel system or a multi-channel system. The sensor part includes acoustic sensors and electrocardiogram (ECG) sensors. The acoustic sensors are the main sensors and are required, with a quantity of N, where 1 ≤ N ≤ 10. When the system is single-channel, N = 1; when the system is multi-channel, N can be any integer greater than 1 and less than or equal to 10. The ECG sensors are auxiliary sensors and are optional. The purpose of acquiring one ECG signal is to facilitate the location of the starting position of the heart sound signal during subsequent lung sound signal processing, thereby improving the effect of heart sound cancellation.

[0042] In one embodiment, the lung sound database creation device collects lung sound data corresponding to each of multiple users, including:

[0043] Determine the duration and location for lung sound data acquisition;

[0044] Determine the data collection information for each user among multiple users;

[0045] Based on the lung sound data collection time and collection site, lung sound information is collected for each user among multiple users.

[0046] The lung sound information and the data acquisition information are identified as lung sound data.

[0047] In this embodiment, before collecting lung sound data for each user among multiple users, the lung sound database creation device needs to first determine the lung sound data collection duration and collection site. When collecting lung sound signals, for a lung sound signal acquisition system with multiple channels, lung sound signals from 10 sites and one ECG signal are collected simultaneously. The collection duration is 30-180 seconds, depending on the situation of the subject: for those with severe illness and physical weakness or children who have difficulty staying still for a long time, the collection time can be 30 seconds; for those whose physical condition can support it, the collection time can be 180 seconds, so as to record respiratory status changes as completely as possible.

[0048] In addition, due to the large volume of the lungs, the chest and back of the human body can hear lung sounds in a large range, but the lung sounds at different positions are quite different. Referring to the lung disease diagnosis guide, the lung sound signal collection system is used to collect lung sound signals at 10 positions of the human body. The 10 positions are: left middle upper lung 1, left lower lung 2, right middle upper lung 3, right lower lung 4, left axillary middle lung 5, right axillary middle lung 6, back left middle upper lung 7, back left lower lung 8, back right middle upper lung 9, and back right lower lung 10. The specific collection number position is shown in Figure 3 It can be understood that for patients with tracheal intubation or serious illness, who cannot sit up or turn over, the first 1-6 positions are recorded.

[0049] In addition, before the lung sound signal collection of the lung sound database creation device, each user in the plurality of users needs to fill in the data collection information table first. The lung sound database creation device can obtain the data collection information of each user in the plurality of users, which includes but is not limited to data collection date and time, data collection device name, collection object identification number, age, gender, onset time, disease diagnosis, previous major diseases, and auxiliary information (such as height, weight, blood pressure, and pulse information). Then, a professional places a sensor for the collection object, checks the connection, and collects each user in the plurality of users based on the lung sound signal collection time length and the collection position to obtain the lung sound information of each user in the plurality of users. The collected lung sound information corresponds to the collection position one by one. The lung sound information and the data collection information are determined as lung sound data.

[0050] It should be noted that the lung sound database creation device collects lung sound information through the lung sound signal collection system, and the lung sound signal collection system is divided into multi-channel simultaneous sampling and single-channel sampling. Therefore, in addition to the multi-channel lung sound signal collection system which can simultaneously sample through each channel, the single-channel lung sound signal collection system needs to sample one position after another in order, and the lung sound signal collection time length of each position remains consistent.

[0051] 102. Preprocess the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user.

[0052] In this embodiment, after the lung sound database creation device collects the lung sound data corresponding to each user in the plurality of users, the lung sound data corresponding to each user can be preprocessed to obtain target lung sound data corresponding to each user. The preprocessing of the lung sound data is described in detail as follows:

[0053] Step 1. Data arrangement is performed on the lung sound data corresponding to the target user to obtain first lung sound data, wherein the target user is any one of the plurality of users.

[0054] In this step, the lung sound database creation device can arrange the lung sound data corresponding to the target user to obtain first lung sound data, which includes format conversion of the lung sound data file, sampling rate conversion of the lung sound signal, bit conversion, and elimination of obviously failed data samples.

[0055] It can be understood that, in order to facilitate subsequent use of data, the most common wav file format, 4k / s sampling rate, and 16-bit bit number can be used, and of course other file formats, sampling rates, and bit numbers can also be used, such as MP3 format, 8k / s sampling rate, and 8-bit bit number, which are not limited in particular.

[0056] Step 2, denoising the first lung sound data to obtain second lung sound data.

[0057] In this step, the lung sound database creation device can use a filter (such as a Butterworth filter, a Chebyshev filter, a Bessel filter, or an elliptical filter) to weaken the noise in the set frequency band when denoising the first lung sound data, and of course other methods can also be used for denoising, such as wavelet transform or Hilbert transform to realize denoising of lung sound data in the transform domain, which is not limited in particular as long as the lung sound data can be denoised. For example, a fifth-order Butterworth band-pass filter is used, with an attenuation rate of 30 decibels, and a band-pass filtering range of 150 to 2000 Hz. The following will be described in combination with Figure 4 The lung sound signals before and after denoising are described as follows, Figure 4 The effect comparison diagram of the lung sound signal before and after denoising provided by the embodiment of the present application is as follows:

[0058] Referring to Figure 4 , it can be seen that the baseline of the original lung sound signal drifts to a certain extent, and the breathing rhythm is not obvious, especially as shown in a), the abnormal lung sound signal waveform is chaotic, and it is difficult to see the breathing rhythm. After denoising, the two lung sound signals become smooth, the breathing period is obvious, and the denoising effect is significant. Figure 4

[0059] Step 3, eliminating the heart sound segment in the second lung sound data to obtain third lung sound data.

[0060] In this step, the lung sound database creation device can eliminate the heart sound segment in the second lung sound data to obtain third lung sound data.

[0061] ​The lung sound signal collection system is a single-channel system, that is, the lung sound data corresponding to the target user is lung sound data collected through the single-channel system, and the multi-subdomain product method of transform domain coefficients can be used to eliminate the heart sound segments in the lung sound signal.

[0062] The principle of the multi-subdomain product method is that the heart sound, the noise and the lung sound have different performances in the transform domain, the transform domain distribution of the noise and the lung sound is wider than that of the heart sound, and the intensity of the noise (such as AWGN) and the lung sound is usually smaller. In contrast, compared with the noise and the lung sound, the distribution of the heart sound is concentrated in the low end of the transform domain. The intensity of the heart sound is usually much higher than that of the noise and the lung sound. Based on this characteristic in the transform domain, the heart sound position is estimated by using the multi-subdomain product of the transform domain coefficients, and transform domain filtering is first performed to eliminate the signal components outside the desired subdomain, so that only the components in the desired subdomain are left. Since the intensity of the heart sound in the required subdomain is high, and the intensity of the noise and the respiratory sound in these subdomains is low, the multiplication of the transform domain coefficients in the remaining required subdomains can reduce the noise and the lung sound, and thus the segments containing the heart sound can be distinguished; after successfully locating the heart sound segments in the second lung sound data, the heart sound segments can be removed from the second lung sound data to obtain clear lung sound, that is, third lung sound data.

[0063] The transform domain coefficients can be wavelet transform coefficients, short-time Fourier transform coefficients or other transform domain coefficients, and the specific embodiments are as follows:

[0064] 1) The multi-subdomain product method of wavelet transform coefficients is used to eliminate the heart sound segments in the lung sound signal:

[0065] The second lung sound data is subjected to wavelet transform and transform domain filtering to obtain the wavelet transform coefficients in the desired subdomain;

[0066] The transform domain coefficients in the desired subdomain are multiplied to determine the heart sound segments contained in the second lung sound data;

[0067] The heart sound segments contained in the second lung sound data are removed to obtain third lung sound data.

[0068] Please refer to Figure 5 , Figure 5 The schematic diagram of the multi-subdomain product method of wavelet transform coefficients provided by the embodiment of the present application for eliminating the heart sound in the lung sound signal is shown in FIG. 1. Figure 5 As shown in FIG. 1, when the wavelet transform coefficients in different subdomains are multiplied (the absolute value of the product is actually taken), the heart sound is strengthened and the lung sound and other noises are suppressed, so that the position of the heart sound can be accurately determined, and then the heart sound segments are removed from the second lung sound data. The following takes the multi-subdomain product method of short-time Fourier transform coefficients as an example to describe this process in detail:

[0069] 2) using the short-time Fourier transform coefficient of the multi-sub-domain product method to eliminate the heart sound segment in the lung sound signal:

[0070] The second lung sound data is subjected to short-time Fourier transform to obtain transform coefficients in the transform domain;

[0071] The target spectrum corresponding to the second lung sound data is generated according to the transform coefficients;

[0072] The spectrum of the second lung sound data of the expected sub-domain is determined according to the target spectrum filtered in the frequency domain;

[0073] The multi-sub-domain product method processing is performed based on the spectrum of the second lung sound data of the expected sub-domain to locate the heart sound segment in the second lung sound data;

[0074] The heart sound segment in the second lung sound data is removed to obtain third lung sound data.

[0075] The following will be described in combination with Figure 6 The short-time Fourier transform coefficient of the multi-sub-domain product method is used to eliminate the heart sound in the lung sound signal:

[0076] Please refer to Figure 6 , Figure 6 The short-time Fourier transform coefficient of the multi-sub-domain product method provided by the present application is used to eliminate the heart sound in the lung sound signal. The time domain waveform of the real sound signal captured by the sound sensor from the chest wall, i.e. the original lung sound waveform, is shown in 601. Then the signal is transformed (here, short-time Fourier transform) to obtain transform coefficients (i.e. Fourier coefficients) in the transform domain, which is shown in 602. The sound spectrum diagram composed of the transform coefficients is shown in 603. It can be seen that the Fourier coefficient values representing the noise and lung sound components are small. The multi-sub-domain product (the logarithmic square value of the product is actually taken) of the Fourier coefficients of the required sub-domain is shown in 604. It can be seen that the values of the sampling points representing the noise and lung sound in the multi-sub-domain product are small. Therefore, the heart sound segment in the second lung sound data can be accurately located according to the results of the multi-sub-domain product, i.e. the segment with the multi-sub-domain product value greater than the threshold, which is shown in 605. The position of the heart sound is located in the box in the figure. The threshold value of the threshold can be determined by different methods. In this example, the sum of the mean value and the standard deviation of the multi-sub-domain product is used as the threshold value. After the position of the heart sound is accurately determined, the heart sound segment can be removed from the second lung sound data, and a clearer lung sound signal, i.e. the third lung sound signal, can be restored. This process is described in step 4 below.

[0077] II. The lung sound signal acquisition system is a multi-channel system, i.e. the lung sound data corresponding to the target user is the lung sound data acquired by the multi-channel system. One lung sound signal with strong signal and one heart sound signal with strong signal can be selected, and adaptive noise cancellation is used to suppress the heart sound, which is as follows:

[0078] Specifically, the first channel of the lung sound data corresponding to the target user and having a lung sound signal strength greater than a first preset value and the second channel of the lung sound data and having a heart sound signal strength greater than a second preset value can be determined; and the heart sound segment in the second lung sound data is eliminated through adaptive noise based on the lung sound data of the first channel and the lung sound data of the second channel, to obtain third lung sound data. Figure 7 As shown in the following figure, Figure 7 The heart sound suppression in the lung sound signal provided by the embodiment of the present application eliminates the heart sound segment in the second lung sound data through the adaptive noise elimination method to obtain the third lung sound data.

[0079] Step 4, restoring the missing segment in the third lung sound data to obtain fourth lung sound data.

[0080] In this step, the lung sound database creation device can first determine the forward segment and / or backward segment corresponding to the missing segment in the third lung sound data; and restore the missing segment in the third lung sound data through the transform domain coefficient based on the forward segment and / or backward segment to obtain the fourth lung sound data. That is, after removing the heart sound segment in the lung sound signal, the missing part in the lung sound can be restored using forward, backward or bidirectional prediction using the transform domain coefficient. Please refer to Figure 8 , Figure 8 The result diagram of the lung sound recovery after removing the heart sound provided by the embodiment of the present application using the Fourier transform coefficient based on the AR model bidirectional linear prediction. 801 represents the original lung sound with heart sound interference, and 802 represents the clean lung sound after removing the heart sound from the original signal and performing lung sound recovery.

[0081] Step 5, segmenting and adjusting the fourth lung sound data to determine the target lung sound data corresponding to the target user.

[0082] In this step, the lung sound database creation device can segment the fourth lung sound data according to the breathing cycle, and dynamically time adjust the segmented fourth lung sound data to obtain the target lung sound data corresponding to the target user. That is, in order to facilitate subsequent processing and use, the lung sound database creation device can segment the lung sound signal according to the breathing cycle. The lung sound signal segmentation can be automatically segmented using an algorithm, or manually segmented (in this step, the method of first automatic segmentation and then manual adjustment can be used to save manpower and time under the premise of ensuring the accuracy of segmentation). The segmented lung sound sample is normalized by dynamic time warping, which is convenient for subsequent feature extraction and intelligent diagnosis and other processing.

[0083] 103, labeling the target lung sound data corresponding to each user.

[0084] In this embodiment, the lung sound database creation device can label the target lung sound data corresponding to each user. The labeling of data samples in the database is an important factor in determining the quality and value of the database, and must be done by professionals. The labeling content of the data samples by professionals includes the description of the samples by researchers, such as the presence of wheezing or bubbling sound, the diagnosis conclusion such as the type and severity of the disease, the start of the respiratory phase and the end of the respiratory phase, and the effective data duration. When the labels of several professionals are inconsistent, please discuss and reach a consensus. If still cannot reach a consensus, take the principle of majority over minority.

[0085] 104. Name the target lung sound data corresponding to each user after labeling according to the naming rules of the data samples in the lung sound database, to obtain the lung sound data samples corresponding to each user.

[0086] In this embodiment, the lung sound database creation device can predefine the naming rules of the data samples in the lung sound database, wherein the naming rules of the data samples in the database are: source code_sublibrary number_patient code_sampling site number_recording number_pathological sound type number, which will be described below:

[0087] The source code is a 3-digit number representing the collection site of the data signal, such as the source code 001 of the Shanghai Jiaotong Hospital;

[0088] The sublibrary number is a 2-digit number representing the type of disease diagnosed by the doctor, such as 01-common pneumonia, 02-atypical pneumonia, 03-novel coronavirus infection, 04-upper respiratory tract infection, 05-asthma, 06-chronic obstructive pulmonary disease, 07-emphysema, 08-lung water, and 09-lung cancer, etc.;

[0089] The patient code is a 4-digit number representing different patients;

[0090] The sampling site number is a 2-digit number, 01-left middle upper lung, 02-left lower lung, 03-right middle upper lung, 04-right lower lung, 05-left axillary middle lung, 06-right axillary middle lung, 07-back left middle upper lung, 08-back left lower lung, 09-back right middle upper lung, and 10-back right lower lung;

[0091] The recording number is a 4-digit number representing different lung sound segments;

[0092] The pathological sound type number is a 1-digit number, 0-no abnormality, 1-rhonchi, 2-wet rales, 3-wheezing, 4-tussiculation, 5-crackles, 6-bubbling sound, and 7-friction sound, etc.

[0093] 105. Store the lung sound data samples corresponding to each user based on the organization structure of the lung sound database to complete the creation of the lung sound database.

[0094] In this embodiment, the lung sound database creation device first considers the organization structure of the lung sound database, please refer to Figure 9 , Figure 9 The organization structure of the lung sound database provided by the embodiment of the present application is shown in the figure, wherein the database not only includes the data samples of patients, but also includes the data samples of the healthy control group. The data of the patients is constructed into a sub-database according to different disease types. The disease types include common pneumonia, atypical pneumonia, new coronavirus infection, upper respiratory tract infection, asthma, chronic obstructive pulmonary disease, emphysema, pulmonary hydrops, lung cancer, etc. Each sub-database is grouped according to the severity of the patient's condition, such as critical type, severe type, common type, etc. Each sub-database and the healthy control group database should contain not only the original data samples, but also the samples after data preprocessing, to facilitate subsequent research and use. Each sub-database is gradually supplemented and improved according to the principle of first acute and then slow, first common and then rare, for example, the sub-database of new coronavirus pneumonia is constructed first. Then, the lung sound database creation device can be stored in the corresponding sub-database according to the condition of each user.

[0095] It should be noted that, in actual application process, in order to protect the privacy and safety of patients, the patient information should be desensitized to prevent data leakage, which includes retrospective data and anonymous data.

[0096] As can be seen from the above, in the embodiment of the present application, the lung sound data corresponding to each user in a plurality of users is collected, and the target lung sound data with unified format is obtained by preprocessing the lung sound data. Then, the lung sound data is labeled, and the lung sound data is named according to the naming rules of the data samples in the database. At the same time, the lung sound data sample corresponding to each user is stored based on the organization structure of the lung sound database. Therefore, the consistency of the data samples in the lung sound database can be ensured, and the information to be saved and the organization form of how to save the lung sound data are standardized to facilitate subsequent searching and utilization.

[0097] The embodiment of the present application is described above from the lung sound database creation method, and the embodiment of the present application is described below from the lung sound database creation device.

[0098] Please refer to Figure 10 , the virtual structure diagram of the lung sound database creation device in the embodiment of the present application, the lung sound database creation device 1000 comprises:

[0099] The acquisition unit 1001 is used for acquiring the lung sound data corresponding to each user in a plurality of users;

[0100] The preprocessing unit 1002 is configured to preprocess the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user in a uniform format;

[0101] The labeling unit 1003 is configured to label the target lung sound data corresponding to each user;

[0102] The naming unit 1004 is configured to name the target lung sound data corresponding to each user after labeling according to a naming rule of a data sample in a lung sound database to obtain lung sound data samples corresponding to each user.

[0103] The storage unit 1005 is configured to store the lung sound data samples corresponding to each user based on an organization structure of the lung sound database to complete creation of the lung sound database, wherein the lung sound database includes data samples of lung diseases and data samples of a healthy control group.

[0104] In a possible design, the preprocessing unit 1002 is specifically configured to:

[0105] perform data arrangement on lung sound data corresponding to a target user to obtain first lung sound data, wherein the target user is any one of the plurality of users;

[0106] perform denoising on the first lung sound data to obtain second lung sound data;

[0107] remove heart sound segments in the second lung sound data to obtain third lung sound data;

[0108] restore missing segments in the third lung sound data to obtain fourth lung sound data;

[0109] perform segmentation and adjustment on the fourth lung sound data to determine the target lung sound data corresponding to the target user.

[0110] In a possible design, if the lung sound data corresponding to the target user is collected by a single-channel system, the preprocessing unit 1002 removes the heart sound segments in the second lung sound data to obtain the third lung sound data, and the removing includes:

[0111] perform transform domain filtering on the second lung sound data to obtain lung sound signals in a desired subdomain;

[0112] perform multi-subdomain product processing on transform domain coefficients in the desired subdomain to determine the heart sound segments contained in the second lung sound data;

[0113] remove the heart sound segments contained in the second lung sound data to obtain the third lung sound data.

[0114] In a possible design, if the lung sound data corresponding to the target user is collected by a single-channel system, the preprocessing unit 1002 removes the heart sound segments in the second lung sound data to obtain third lung sound data, including:

[0115] performing short-time Fourier transform on the second lung sound data to obtain transform coefficients in a transform domain;

[0116] generating a target spectrum corresponding to the second lung sound data according to the transform coefficients;

[0117] determining a spectrum of the second lung sound data in an expected subdomain according to the target spectrum filtered in the frequency domain;

[0118] performing multi-subdomain product method processing based on the spectrum of the second lung sound data in the expected subdomain to locate the heart sound segments in the second lung sound data;

[0119] removing the heart sound segments in the second lung sound data to obtain the third lung sound data.

[0120] In a possible design, if the lung sound data corresponding to the target user is collected by a multi-channel system, the preprocessing unit 1002 removes the heart sound segments in the second lung sound data to obtain third lung sound data, including:

[0121] determining a first channel in which a lung sound signal strength of the lung sound data corresponding to the target user is greater than a first preset value and a second channel in which a heart sound signal strength is greater than a second preset value;

[0122] removing the heart sound segments in the second lung sound data by adaptive noise cancellation based on the lung sound data of the first channel and the lung sound data of the second channel to obtain the third lung sound data.

[0123] In a possible design, the preprocessing unit 1002 restores missing segments in the third lung sound data to obtain fourth lung sound data, including:

[0124] determining a forward segment and / or a backward segment corresponding to the missing segments;

[0125] restoring the missing segments in the third lung sound data by transform domain coefficient according to the forward segment and / or the backward segment to obtain the fourth lung sound data.

[0126] In a possible design, the preprocessing unit 1002 segments and adjusts the fourth lung sound data to determine target lung sound data corresponding to the target user, including:

[0127] segmenting the fourth lung sound data according to a breathing cycle;

[0128] The segmented fourth lung sound data is dynamically time-adjusted to obtain target lung sound data corresponding to the target user.

[0129] In a possible design, the collection unit 1001 is specifically configured to:

[0130] determine a lung sound signal collection duration and a collection position;

[0131] determine data collection information corresponding to each user in the plurality of users;

[0132] collect each user in the plurality of users based on the lung sound signal collection duration and the collection position, to obtain lung sound information of each user in the plurality of users;

[0133] determine the lung sound information and the data collection information as the lung sound data.

[0134] The above Figure 10 The device for creating a lung sound database in the embodiment of the present application is described from the perspective of a modular functional entity, and the device for creating a lung sound database in the embodiment of the present application is described in detail from the perspective of hardware processing. Please refer to Figure 11 An embodiment of the device for creating a lung sound database 1100 in the embodiment of the present application is shown in the figure, and the device for creating a lung sound database 1100 includes:

[0135] an input device 1101, an output device 1102, a processor 1103, and a memory 1104 (wherein the number of processors 1103 can be one or more, Figure 11 and one processor 1103 is taken as an example in some embodiments of the present application). In some embodiments of the present application, the input device 1101, the output device 1102, the processor 1103, and the memory 1104 can be connected through a communication bus or other means, wherein, Figure 11 and the communication bus is taken as an example in some embodiments of the present application.

[0136] The processor 1103 is configured to execute the following steps by calling operation instructions stored in the memory 1104:

[0137] collect lung sound data corresponding to each user in the plurality of users;

[0138] preprocess the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user in a unified format;

[0139] label the target lung sound data corresponding to each user;

[0140] According to a naming rule of data samples in the lung sound database, the target lung sound data corresponding to each user is named to obtain lung sound data samples corresponding to each user.

[0141] The lung sound data samples corresponding to each user are stored based on an organization structure of the lung sound database to complete creation of the lung sound database, the lung sound database including data samples of lung patients and data samples of a healthy control group.

[0142] By invoking the operation instructions stored in the memory 1104, the processor 1103 is further configured to execute Figure 1 any of the corresponding embodiments.

[0143] Referring to Figure 12 , Figure 12 an embodiment of an electronic device provided by the embodiment of the present application is shown.

[0144] As Figure 12 shown, the embodiment of the present application provides an electronic device, which includes a memory 1210, a processor 1220, and a computer program 1211 stored in the memory 1210 and executable on the processor 1220, and the processor 1220 implements the following steps when executing the computer program 1211:

[0145] Collect lung sound data corresponding to each user in a plurality of users;

[0146] Preprocess the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user in a unified format;

[0147] Label the target lung sound data corresponding to each user;

[0148] According to a naming rule of data samples in the lung sound database, the target lung sound data corresponding to each user is named to obtain lung sound data samples corresponding to each user;

[0149] The lung sound data samples corresponding to each user are stored based on an organization structure of the lung sound database to complete creation of the lung sound database, the lung sound database including data samples of lung patients and data samples of a healthy control group.

[0150] In the specific implementation process, when the processor 1220 executes the computer program 1211, any of the Figure 1 corresponding embodiments can be implemented.

[0151] Since the electronic device introduced in the embodiment is the device used by the device for creating a lung sound database in the embodiment of the present application, based on the method introduced in the embodiment of the present application, the person skilled in the art can understand the specific implementation of the electronic device in the embodiment and its various forms, so the electronic device how to implement the method in the embodiment of the present application is not described in detail here, as long as the device used by the person skilled in the art to implement the method in the embodiment of the present application belongs to the scope of the present application.

[0152] Please refer to Figure 13 , Figure 13 An embodiment of a computer readable storage medium provided in the embodiment of the present application is shown.

[0153] As Figure 13 shown, the embodiment of the present application further provides a computer readable storage medium 1300, which stores a computer program 1311, and the computer program 1311 is executed by a processor to implement the following steps:

[0154] Collect lung sound data corresponding to each user in a plurality of users;

[0155] Preprocess the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user in a unified format;

[0156] Label the target lung sound data corresponding to each user;

[0157] Name the target lung sound data corresponding to each user after labeling according to the naming rules of the data samples in the lung sound database to obtain lung sound data samples corresponding to each user;

[0158] Store the lung sound data samples corresponding to each user based on the organization structure of the lung sound database to complete the creation of the lung sound database, and the lung sound database includes data samples of lung patients and data samples of a healthy control group.

[0159] In the specific implementation process, the computer program 1311 is executed by the processor to implement Figure 1 any embodiment in the corresponding embodiment.

[0160] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0161] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code thereon for use by or in connection with an instruction execution system. For the purposes of this description, a computer-usable or computer readable storage medium can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-usable storage medium can be a computer- readable storage medium that can be any media that can be accessed by the computer. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of computer- readable program code means which can be accessed by the computer. The computer- readable storage medium can also be, for example, a computer-usable side, a computer-usable base, or a computer-usable hub. The computer-readable storage medium can further include a computer program product for practicing a

[0162] The present application is described in reference to the flowchart illustrations and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart and / or block diagram. Figure 1 one or more functions specified in the flowchart and / or block diagram.

[0163] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flowchart and / or block diagrams. Figure 1 one or more functions specified in the flowchart and / or block diagram. Figure 1 one or more functions specified in the flowchart and / or block diagram.

[0164] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart and / or block diagrams. Figure 1 one or more functions specified in the flowchart and / or block diagram. Figure 1 one or more functions specified in the flowchart and / or block diagram.

[0165] Embodiments of the present application also provide a computer program product comprising computer software instructions which, when run on a processing device, cause the processing device to perform the steps of any of the methods. ​ the flowchart of the corresponding embodiment.

[0166] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the flow or function described in the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that the computer can store or be integrated into a data storage device such as a server, data center, etc. containing one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0167] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0168] In several embodiments provided by the present application, it should be understood that the disclosed system, device and method can be implemented by other ways. For example, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and another division mode can be used in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0169] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.

[0170] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0171] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the entire or part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0172] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some technical features. These modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A method of creating a lung sound database, characterized by, The method comprises the following steps: Collecting lung sound data corresponding to each user in a plurality of users; Preprocessing the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user in a unified format; Labeling the target lung sound data corresponding to each user; According to the naming rules of the data samples in the lung sound database, the target lung sound data corresponding to each user after labeling is named to obtain the lung sound data samples corresponding to each user; Based on the organizational structure of the lung sound database, the lung sound data samples corresponding to each user are stored to complete the creation of the lung sound database, which includes data samples of lung diseases and data samples of healthy control groups; The preprocessing of the lung sound data corresponding to each user to obtain the target lung sound data corresponding to each user comprises: Data collation is performed on the lung sound data corresponding to the target user to obtain first lung sound data, wherein the target user is any one of the plurality of users; The first lung sound data is denoised to obtain second lung sound data; The heart sound segment in the second lung sound data is removed to obtain third lung sound data; The missing segment in the third lung sound data is recovered to obtain fourth lung sound data; The fourth lung sound data is segmented and adjusted to determine the target lung sound data corresponding to the target user; The recovery of the missing segment in the third lung sound data to obtain the fourth lung sound data comprises: Determine the forward segment and / or backward segment corresponding to the missing segment; According to the forward segment and / or the backward segment, the missing segment in the third lung sound data is recovered by transform domain coefficient to obtain the fourth lung sound data.

2. The method of claim 1, wherein, If the lung sound data corresponding to the target user is collected by a single-channel system, the removal of the heart sound segment in the second lung sound data to obtain the third lung sound data comprises: Transform domain filtering is performed on the second lung sound data to obtain lung sound signals in the desired sub-domain; The transform domain coefficients in the desired sub-domain are processed by multi-sub-domain multiplication method to determine the heart sound segment contained in the second lung sound data; The heart sound segment contained in the second lung sound data is removed to obtain the third lung sound data.

3. The method of claim 1, wherein, If the lung sound data corresponding to the target user is collected by a single-channel system, the removal of the heart sound segment in the second lung sound data to obtain the third lung sound data comprises: Short-time Fourier transform is performed on the second lung sound data to obtain transform coefficients in the transform domain; According to the transform coefficients, the target frequency spectrum corresponding to the second lung sound data is generated; According to the target frequency spectrum after frequency domain filtering, the frequency spectrum of the second lung sound data in the desired sub-domain is determined; Based on the frequency spectrum of the second lung sound data in the desired sub-domain, the multi-sub-domain multiplication method is processed to locate the heart sound segment in the second lung sound data; The heart sound segment in the second lung sound data is removed to obtain the third lung sound data.

4. The method of claim 1, wherein, If the lung sound data corresponding to the target user is collected by a multi-channel system, the rejecting the heart sound segment in the second lung sound data to obtain third lung sound data comprises: Determining a first channel with a lung sound signal intensity greater than a first preset value and a second channel with a heart sound signal intensity greater than a second preset value in the lung sound data corresponding to the target user; Based on the lung sound data of the first channel and the lung sound data of the second channel, the heart sound segment in the second lung sound data is removed by adaptive noise cancellation to obtain the third lung sound data.

5. The method of claim 1, wherein, The segmenting and adjusting the fourth lung sound data to determine the target lung sound data corresponding to the target user comprises: Segmenting the fourth lung sound data according to the respiratory cycle; Performing dynamic time adjustment on the segmented fourth lung sound data to obtain the target lung sound data corresponding to the target user.

6. The method according to any one of claims 1 to 5, characterized in that, The collecting lung sound data corresponding to each user in the plurality of users comprises: Determining the lung sound signal collection time length and the collection site; Determining the data collection information corresponding to each user in the plurality of users; Based on the lung sound signal collection time length and the collection site, collecting each user in the plurality of users to obtain lung sound information of each user in the plurality of users; Determining the lung sound information and the data collection information as the lung sound data.

7. An apparatus for creating a lung sound database, characterized by comprising: The lung sound database creation device is used to execute the lung sound database creation method of any one of claims 1-6, comprising: A collection unit for collecting lung sound data corresponding to each user in the plurality of users; A preprocessing unit for preprocessing the lung sound data corresponding to each user to obtain target lung sound data corresponding to each user in a uniform format; A labeling unit for labeling the target lung sound data corresponding to each user; A naming unit for naming the labeled target lung sound data corresponding to each user according to the naming rules of the data samples in the lung sound database to obtain the lung sound data samples corresponding to each user; A storage unit for storing the lung sound data samples corresponding to each user based on the organization structure of the lung sound database to complete the creation of the lung sound database, wherein the lung sound database comprises data samples of lung diseases and data samples of a healthy control group.

8. The apparatus of claim 7, wherein, The preprocessing unit is specifically used for: Data arrangement of the lung sound data corresponding to the target user to obtain first lung sound data, wherein the target user is any one of the plurality of users; De-noising the first lung sound data to obtain second lung sound data; Rejecting the heart sound segment in the second lung sound data to obtain third lung sound data; Restoring the missing segment in the third lung sound data to obtain fourth lung sound data; Segmenting and adjusting the fourth lung sound data to determine the target lung sound data corresponding to the target user.