Snore sound recognition method based on multi-feature fusion

By using a multi-feature fusion method for snoring recognition, multiple features of the snoring sample audio are extracted to form a set of positive and confused samples. The target feature comparison order is determined, which solves the problem of low recognition accuracy in existing technologies and achieves higher recognition accuracy.

CN116030840BActive Publication Date: 2026-01-27BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211603944.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2026-01-27
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

In existing technologies, snoring-based recognition methods do not consider the representativeness of different characteristics for different populations, resulting in low recognition accuracy.

Method used

A snoring recognition method based on multi-feature fusion is adopted. By acquiring training and validation sets, multiple audio features of snoring sample audio are extracted to form a positive sample set and a confused sample set. The target feature comparison order is determined, and then the snoring audio of the test patient is compared to determine the triggering site.

Benefits of technology

It improves the accuracy of identifying the location of snoring triggers, avoids using the same feature to identify all patients, and improves the accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116030840B_ABST
    Figure CN116030840B_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to the technical field of audio recognition, and provides a snoring sound recognition method based on multi-feature fusion, which comprises the following steps: acquiring a plurality of snoring sound sample audio training sets and verification sets; extracting at least two types of first audio features corresponding to each snoring sound sample audio; acquiring a positive sample set and a confusion sample set according to the first audio features of different snoring sound sample audios and sample labels; determining a target feature comparison sequence according to the positive sample set, the confusion sample set and the verification set; and comparing the audio features of the snoring sound audio of a to-be-tested patient based on the target feature comparison sequence to determine an excitation position of the snoring sound caused by the to-be-tested patient. The method provided by the application can recognize snoring sound by using different snoring sound audio features for different patients, and the accuracy of recognizing the excitation position of the snoring sound of a patient is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a snoring recognition method based on multi-feature fusion. Background Technology

[0002] Obstructive sleep apnea-hypopnea syndrome (OSAHS) is a common clinical condition characterized by snoring at night, accompanied by pauses in breathing and shallow breathing. Patients with OSAHS often experience daytime sleepiness and are highly susceptible to various cardiovascular and cerebrovascular diseases, seriously endangering their physical and mental health.

[0003] In the diagnosis of obstructive sleep apnea-hypopnea syndrome, it is usually necessary to identify the triggering site (e.g., neck, base of tongue, chest cavity, lungs, etc.) that causes snoring. Currently, doctors often identify the triggering site by analyzing the audio recordings of snoring produced during sleep. Specifically, feature extraction is performed on the snoring audio recordings to obtain a specific type of feature (e.g., MFCC (Mel-Frequency Cepstral Coefficients), GFCC (Gammatone feature cochleagram), or frequency domain features), and the triggering site is identified based on this type of feature.

[0004] However, the above-mentioned one-size-fits-all identification method does not take into account the representativeness of different features for different groups of people, and its identification accuracy is low. Summary of the Invention

[0005] This application provides a snoring recognition method based on multi-feature fusion, which can combine the characteristics of different features of snoring audio and recognize the patient's snoring audio through different features, with a high recognition accuracy.

[0006] The first aspect of this application provides a snoring recognition method based on multi-feature fusion, the method comprising:

[0007] A training set and a validation set are obtained. The training set includes: multiple snoring sample audios, and sample labels corresponding to each of the snoring sample audios. The sample labels are used to indicate the first excitation part of the snoring sample audios and to indicate the first target feature for determining the first excitation part. The validation set includes: multiple snoring verification audios, and validation labels corresponding to each of the snoring verification audios. The validation labels include the second excitation part of the snoring verification audios and the second target feature for determining the second excitation part.

[0008] Extract at least two types of first audio features corresponding to each of the snoring sample audios;

[0009] Based on the first audio features of different snoring sample audios and the sample labels, obtain positive sample sets and confused sample sets corresponding to each type of first audio feature. The positive sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is greater than or equal to a preset threshold. The confused sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is less than a preset threshold.

[0010] Based on the positive sample set, the confused sample set, and the validation set, determine the target feature comparison order when performing snoring recognition;

[0011] Based on the target feature comparison order, the audio features of the snoring audio of the test patient are compared, and the triggering site of the snoring of the test patient is determined by the comparison results.

[0012] A second aspect of this application provides a snoring recognition device based on multi-feature fusion, the device comprising:

[0013] The first acquisition module is used to acquire a training set and a validation set. The training set includes: multiple snoring sample audios, and sample labels corresponding to each of the snoring sample audios. The sample labels are used to indicate the first excitation part of the snoring sample audios and to indicate the first target feature for determining the first excitation part. The validation set includes: multiple snoring verification audios, and verification labels corresponding to each of the snoring verification audios. The verification labels include the second excitation part of the snoring verification audios and the second target feature for determining the second excitation part.

[0014] The extraction module is used to extract at least two types of first audio features corresponding to each of the snoring sample audios;

[0015] The second acquisition module is used to acquire, based on the first audio features of different snoring sample audios and the sample labels, a positive sample set and a confused sample set corresponding to each type of first audio feature, respectively. The positive sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is greater than or equal to a preset threshold. The confused sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is less than a preset threshold.

[0016] The determination module is used to determine the target feature comparison order when performing snoring recognition based on the positive sample set, the confused sample set, and the verification set;

[0017] The comparison and determination module is used to compare the audio features of the snoring audio of the patient under test with the target feature comparison order, and determine the triggering site of the snoring of the patient under test through the comparison results.

[0018] A third aspect of this application provides a computer device, the computer device comprising:

[0019] Processor; memory used to store processor-executable instructions.

[0020] A processor for reading executable instructions from memory and executing the instructions to implement the snoring recognition method provided in the first aspect.

[0021] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the snoring recognition method based on multi-feature fusion provided in the first aspect.

[0022] The fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the snoring recognition method based on multi-feature fusion provided in the first aspect.

[0023] The technical solution provided in this application can achieve at least the following beneficial effects:

[0024] The snoring recognition method based on multi-feature fusion provided in this application includes: acquiring a training set and a validation set of multiple snoring sample audios; extracting at least two types of first audio features corresponding to each snoring sample audio; acquiring a positive sample set and a confused sample set corresponding to each type of first audio feature based on the first audio features and sample labels of different snoring sample audios; determining the target feature comparison order for snoring recognition based on the positive sample set, the confused sample set, and the validation set; comparing the audio features of the snoring audio of the test patient based on the target feature comparison order, and determining the excitation site of the snoring in the test patient through the comparison results. The snoring recognition method provided in this application can compare different features of the snoring audio sequentially according to a predetermined feature comparison order, so as to determine one audio feature from multiple features to determine the excitation site of the snoring in the test patient. Different patients can be identified based on different snoring audio features, avoiding and improving the situation where the same feature is used to identify the snoring excitation site for all patients, thus improving the accuracy of identifying the snoring excitation site of the patient. Attached Figure Description

[0025] Figure 1 This is an application environment diagram illustrating an exemplary embodiment of the snoring recognition method based on multi-feature fusion.

[0026] Figure 2 This is a flowchart illustrating an exemplary embodiment of a snoring recognition method based on multi-feature fusion.

[0027] Figure 3 This is a flowchart illustrating an exemplary embodiment of a snoring recognition method based on multi-feature fusion.

[0028] Figure 4 This is a flowchart illustrating an exemplary embodiment of a snoring recognition method based on multi-feature fusion.

[0029] Figure 5 This is a flowchart illustrating an exemplary embodiment of a snoring recognition method based on multi-feature fusion.

[0030] Figure 6 This is a flowchart illustrating an exemplary embodiment of a snoring recognition method based on multi-feature fusion.

[0031] Figure 7 This is a schematic diagram of a snoring recognition device based on multi-feature fusion, as illustrated in an exemplary embodiment of this application;

[0032] Figure 8 This is a schematic diagram of the internal structure of a computer device shown in an exemplary embodiment of this application. Detailed Implementation

[0033] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0034] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0035] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0036] like Figure 1 As shown in Figure 1, this is a schematic diagram of the system structure of the snoring recognition method based on multi-feature fusion according to an embodiment of this application. The system includes a server 102 and at least one snoring acquisition device 104. Figure 1 The implementation of server 102 and snoring sound collection device 104 presented herein is merely illustrative and not intended to limit their application. The server 102 and snoring sound collection device 104 can be wirelessly connected. Optionally, server 102 can communicate with snoring sound collection device 104 via a mobile network. The mobile network standard can be any of 2G (GSM), 2.5G (GPRS), 3G (WCDMA, TD-SCDMA, CDMA2000, UTMS), 4G (LTE), 4G+ (LTE+), WiMax, 5G, or other future network standards. Optionally, server 102 can also communicate with snoring sound collection device 104 via Bluetooth, WiFi, infrared, or other methods. Snoring sound collection device 104 transmits at least one segment of snoring audio to server 102 via mobile network, Bluetooth, WiFi, infrared, or other methods. Server 102 extracts and identifies features from the snoring audio to determine the triggering site for the snoring.

[0037] In this embodiment, the physical form of the server is not limited. Server 102 can be a single server device, a server array, or a virtual machine (VM) running in a cloud-based server array. Alternatively, server 102 can also refer to other computing devices with corresponding service capabilities, such as computers, smartphones, or other terminals 103 running server programs.

[0038] The configuration of server 102 may include, but is not limited to: processor 1021, memory 1022, interface device 1023, communication device 1024, input device 1025, and output device 1026. Processor 1022 may include, but is not limited to, a central processing unit (CPU), a microprocessor (MCU), etc. Memory 1022 may include, but is not limited to, ROM (Read-Only Memory), RAM (Random Access Memory), and non-volatile memory such as a hard disk. Interface device 1023 may include, but is not limited to, a USB interface, a serial interface, and a parallel interface. Communication device 1024 may be capable of wired or wireless communication, specifically including WiFi communication, Bluetooth communication, 2G / 3G / 4G / 5G communication, etc. Input device 1025 may include, but is not limited to, a keyboard, mouse, touchscreen, microphone, etc. Output device 1026 may include, but is not limited to, a display screen, speakers, etc.

[0039] However, current technologies, for different patients, extract features from the snoring audio to obtain a specific feature (such as MFCC (Mel-Frequency Cepstral Coefficients), GFCC (Gammatone feature cochleagram), or frequency domain features), and then identify the excitation site based on this single feature. This one-size-fits-all identification method does not consider the representativeness of different features for different populations, resulting in low accuracy.

[0040] like Figure 2 As shown, Figure 2 This application provides a method for snoring recognition based on multi-feature fusion, which is applied to... Figure 1 Taking server 102 as an example, the method includes the following steps:

[0041] Step S100: Obtain a training set and a validation set. The training set includes: multiple snoring sample audios, and sample labels corresponding to each snoring sample audio. The sample labels are used to indicate the first excitation part of the snoring sample audio, and to indicate the first target feature for determining the first excitation part. The validation set includes: multiple snoring validation audios, and validation labels corresponding to each snoring validation audio. The validation labels include the second excitation part of the snoring validation audio, and the second target feature for determining the second excitation part.

[0042] The training set includes multiple snoring sample audio files. The number of these files can be, for example, 100, 1000, or 3000, and is not limited here. These multiple snoring sample audio files can come from different patients. The first triggering site has been identified in the multiple snoring sample audio files based on the first target feature. It should be noted that the first target feature corresponding to the multiple snoring sample audio files can be different or the same, and the first triggering site corresponding to the multiple snoring sample audio files can also be different or the same.

[0043] For example, the training set includes 1000 snoring audio samples:

[0044] 1-MFCC characteristic - Palatal snoring;

[0045] 2-GFCC characteristic - oropharyngeal snoring;

[0046] 3-Frequency domain characteristics-Snoring at the base of the tongue;

[0047] 4-MFCC characteristic - epiglottic snoring;

[0048] 5-Frequency Domain Characteristics-Snoring at the Base of the Tongue;

[0049] 6-GFCC characteristic - oropharyngeal snoring;

[0050] 7-Frequency Domain Characteristics-Palate Snoring;

[0051] 8-MFCC characteristic - oropharyngeal snoring;

[0052] 9-Frequency Domain Characteristics-Palate Snoring;

[0053] 10-Frequency Domain Characteristics-Snoring at the Base of the Tongue;

[0054] 11-GFCC characteristic - epiglottic snoring;

[0055] 12-MFCC characteristic - palatal snoring;

[0056] 13-MFCC characteristic - oropharyngeal snoring;

[0057] 14-Frequency Domain Characteristics-Epiglottic Snoring;

[0058] 15-GFCC characteristic - palatal snoring;

[0059] 16-Frequency Domain Characteristics-Oropharyngeal Snoring;

[0060] 17-MFCC characteristic - epiglottic snoring;

[0061] 18-Frequency Domain Characteristics-Snoring at the Base of the Tongue;

[0062] 19-Frequency Domain Characteristics-Palate Snoring;

[0063] 20-GFCC characteristic - oropharyngeal snoring;

[0064] ...

[0065] The validation set includes multiple snoring verification audio files. The number of these files can be, for example, 100, 1000, 3000, etc., and is not limited here. These multiple snoring verification audio files can come from different patients. The second triggering site has been identified in each of the multiple snoring verification audio files based on the second target feature. It should be noted that the second target features corresponding to the multiple snoring verification audio files can be different or the same, and the second triggering sites corresponding to the multiple snoring verification audio files can be different or the same. The number of snoring verification audio files in the validation set can be the same as or different from the number of snoring sample audio files in the training set; this application does not limit this.

[0066] For example, the verification set includes 100 snoring verification audio clips:

[0067] 1-MFCC characteristic - Palatal snoring;

[0068] 2-GFCC characteristic - oropharyngeal snoring;

[0069] 3-Frequency domain characteristics-Snoring at the base of the tongue;

[0070] 4-MFCC characteristic - epiglottic snoring;

[0071] 5-Frequency Domain Characteristics-Snoring at the Base of the Tongue;

[0072] 6-GFCC characteristic - epiglottic snoring;

[0073] 7-Frequency Domain Characteristics-Oropharyngeal Snoring;

[0074] 8-MFCC characteristic - epiglottic snoring;

[0075] 9-Frequency Domain Characteristics-Palate Snoring;

[0076] 10-Frequency Domain Characteristics-Snoring at the Base of the Tongue;

[0077] ...

[0078] The training set is used for subsequent partitioning into positive and confused sample sets, while the validation set is used to determine the alignment order of the targets. Both the training and validation sets can be obtained from historical databases.

[0079] Step S200: Extract at least two types of first audio features corresponding to each snoring sample audio.

[0080] After obtaining the training set, it is necessary to divide the audio samples of each snoring sound in the training set. The division is based on the similarity between the audio features of each snoring sound sample and the label. Therefore, it is necessary to extract at least two types of first audio features corresponding to each snoring sound sample. The first audio feature is at least two of the first feature, second feature and third feature. The first feature is a frequency domain feature, the second feature is an MFCC feature and the third feature is a GFCC feature.

[0081] For example, MFCC features, GFCC features, and frequency domain features can be extracted from all snoring audio samples to obtain a first feature set, a second feature set, and a third feature set. The first feature set contains 1000 MFCC features, the second feature set contains 1000 GFCC features, and the third feature set contains 1000 frequency domain features.

[0082] The following section explains how to extract MFCC features, GFCC features, and frequency domain features from snoring audio samples:

[0083] First, this application can extract the frequency domain characteristics of each snoring sample audio by fractional Fourier transform, and obtain the first feature corresponding to each snoring sample audio, where the first feature is a frequency domain feature.

[0084] Specifically, the audio samples of each snoring sound can be preprocessed first to obtain snoring sound samples with better quality and the same amplitude, which can facilitate subsequent batch extraction of audio features from each snoring sound sample audio and improve the efficiency of recognition.

[0085] Preprocessing can involve amplitude normalization of each snoring audio sample, i.e., processing using the following formula:

[0086]

[0087] Where, x(i) ′ is the amplitude-normalized snoring audio data, x(i) is the original snoring audio data, and max(|x(i)|) is the maximum value of the absolute value of the amplitude of the original snoring audio data.

[0088] Based on the preprocessed snoring audio samples described above, the first feature can be extracted using the following formula:

[0089]

[0090]

[0091] Where j is the complex number symbol and p is the Fourier order.

[0092] Secondly, this application can be made through, for example Figure 3 The steps shown extract the second feature of each snoring audio sample, which is the MFCC feature.

[0093] Step S201: Perform frame segmentation on the audio of each snoring sample to obtain the first audio corresponding to each snoring sample audio.

[0094] Step S202: Perform Hamming windowing on each of the first audio samples to obtain multiple second audio samples;

[0095] Step S203: Perform fast Fourier transform on each second audio signal to obtain multiple discrete power spectra;

[0096] Step S204: Squaring each discrete power spectrum yields multiple energy spectra, and filtering each energy spectrum is performed using a pre-constructed first filter group.

[0097] Step S205: Perform cepstral analysis on each energy spectrum after filtering to obtain the second feature corresponding to each snoring sample audio.

[0098] In the process of frame splitting, the "frame length" can be 64ms, the "frame shift" can be 32ms, and the snoring sample audio of "one frame" is regarded as one sample.

[0099] The first filter is a Mel filter bank, which can be constructed in the following way:

[0100] Construct a filter bank with m triangular filters, specifically as follows:

[0101]

[0102] In the above formula, the center frequency f(m) of each filter is defined as:

[0103]

[0104] In the above formula, f l f is the lowest frequency in the filter's frequency range. h The highest frequency in the filter's frequency range is N, where N is the length of the audio after frame division, windowing, and Fourier transform, and f is... s Sampling rate,

[0105] F mel The function is defined as:

[0106]

[0107] Defined as:

[0108]

[0109] Finally, this application can be made through, for example Figure 4 The steps shown extract the third feature of each snoring audio sample. The third feature is the GFCC feature.

[0110] Step S206: Squaring each discrete power spectrum yields multiple energy spectra, and filtering each energy spectrum is performed using a pre-constructed second filter bank.

[0111] Step S207: Perform cepstral analysis on each energy spectrum after filtering to obtain the third feature corresponding to each snoring sample audio.

[0112] The second filter bank can be a Gammatone filter bank, which can be obtained according to the following method:

[0113] According to the critical frequency band table of the human ear and the maximum frequency f max Determine the center frequency f of the Gammatone filter c And the number of filters N;

[0114] Based on the center frequency f c Calculate the equivalent rectangular bandwidth ERB(f) c ),

[0115]

[0116] In the above formula, f c The center frequency;

[0117] The attenuation factor b of the filter is calculated based on the equivalent rectangular bandwidth. i b i =1.019ERB(f c );

[0118] Constructing a Gammatone filter bank g i (t),

[0119]

[0120] A is the filter gain, n is the filter order, and b i f is the attenuation factor of the filter. c Let U(t) be the center frequency of the filter, and U(t) be the step function. Let N be the phase and N be the number of filters.

[0121] Step S300: Based on the first audio features and sample labels of different snoring sample audios, obtain the positive sample set and confused sample set corresponding to each type of first audio feature respectively. The positive sample set includes first audio features whose sample labels are the same as other first audio features that meet the first condition of similarity with the target and have a similarity rate greater than or equal to a preset threshold. The confused sample set includes first audio features whose sample labels are the same as other first audio features that meet the first condition of similarity with the target and have a similarity rate less than a preset threshold.

[0122] After obtaining at least two types of audio features corresponding to each snoring sample audio according to the above method, a first feature set, a second feature set, and a third feature set, as mentioned above, are obtained. The similarity between each pair of features can be calculated according to the following formula:

[0123]

[0124] Among them, X i Y i It is the feature vector of the first audio feature of the same type for any two different snoring audio samples in each snoring audio sample.

[0125] For example, the first feature set includes: A1, A2, A3, A4, A5, A6, A7, A8, A9, A10...1000 frequency domain features.

[0126] First, the similarity between A1 and A2 is calculated using the above formula, then the similarity between A1 and A3, the similarity between A1 and A4, the similarity between A1 and A5, and so on, until the similarity between A1 and A1000 is calculated, resulting in 999 similarity scores.

[0127] Next, similarity scores greater than a preset threshold (such as 99%, 98%, 95%, 85%, 80%) can be selected from 999 similarity scores to determine frequency domain features with similarity scores greater than the preset threshold. For example, frequency domain features similar to A1 can be identified as A3, A5, and A8.

[0128] Then, obtain the labels of the snoring sample audios corresponding to A1, A3, A5 and A8: MFCC feature - neck, frequency domain feature - lung, frequency domain feature - tongue root, MFCC feature - lung; it can be seen that the label of the snoring sample audio corresponding to A1 is different from that of other snoring sample audios, so A1 is classified into the confused sample set.

[0129] It should be noted here that A1 can be classified into the positive sample set if the labels of the snoring sample audio corresponding to A1 are the same as those of other snoring sample audios, or if the similarity rate is greater than a preset threshold. The preset threshold can be 95%, 90%, 80%, 75%, etc., and is not limited here.

[0130] Following the above method, continue to divide A2, A3, A4, A5, A6, A7, A8, A9, A10...A1000, and finally divide the first feature set into the first positive sample set and the first confused sample set.

[0131] The second feature set includes: B1, B2, B3, B4, B5, B6, B7, B8, B9, B10...1000 MFCC features.

[0132] First, the similarity between B1 and B2, the similarity between B1 and B3, the similarity between B1 and B4, the similarity between B1 and B5, and so on, was calculated using the above formula, resulting in 999 similarity scores.

[0133] Next, similarity values ​​greater than a preset threshold can be selected from 999 similarities, and then frequency domain features with similarity values ​​greater than the preset threshold can be determined. For example, the frequency domain features most similar to B1 can be determined as B20, B50 and B800.

[0134] Furthermore, the labels for the snoring sample audios corresponding to B1, B20, B50, and B800 are obtained: MFCC feature - neck, frequency domain feature - lungs, frequency domain feature - tongue root, and MFCC feature - lungs. It can be seen that the labels for the snoring sample audios corresponding to B1 are different from the others, so B1 is classified into the confused sample set. If the labels for the snoring sample audios corresponding to B1 are the same as the others, then B1 is classified into the positive sample set.

[0135] Following the above method, continue to divide B2, B3, B4, B5, B6, B7, B8, B9, B10...B1000, and finally divide the second feature set into the second positive sample set and the second confused sample set.

[0136] The third feature set includes: C1, C2, C3, C4, C5, C6, C7, C8, C9, C10...1000 MFCC features.

[0137] First, the similarity between C1 and C2, the similarity between C1 and C3, the similarity between C1 and C4, the similarity between C1 and C5, ..., the similarity between C1 and C1000 was calculated using the above formula, resulting in 999 similarities.

[0138] Next, similarity scores greater than a preset threshold can be selected from 999 similarity scores to determine the frequency domain features most similar to C1. For example, the frequency domain features most similar to C1 can be C33, C45, and C700.

[0139] Then, the labels for the snoring sample audios corresponding to C1, C33, C45, and C700 are obtained: MFCC feature - palatal snoring, frequency domain feature - oropharyngeal snoring, frequency domain feature - tongue root snoring, and MFCC feature - epiglottic snoring. It can be seen that the label of the snoring sample audio corresponding to C1 is different from that of other snoring sample audios, or the similarity rate is less than a preset threshold. Therefore, C1 is classified into the confused sample set. It should be noted that C1 can be classified into the positive sample set if the label of the snoring sample audio corresponding to C1 is the same as that of other snoring sample audios, or the similarity rate is greater than a preset threshold.

[0140] Following the above method, continue to divide C2, C3, C4, C5, C6, C7, C8, C9, C10...C1000, and finally divide the third feature set into the third positive sample set and the third confused sample set.

[0141] In this way, we obtain three sets of positive samples and three sets of confused samples.

[0142] Step S400: Determine the target feature comparison order for snoring recognition based on the positive sample set, the confused sample set, and the validation set;

[0143] In one embodiment, it can be achieved through, for example Figure 5 The method steps shown are used to determine the target feature alignment order:

[0144] Step S401: Determine the first order. The first order is any one of the preset order sets. The order is used to indicate the extraction order of various features.

[0145] If there are two audio features, the order in the sequence set includes: A1 feature - A2 feature, A2 feature - A1 feature; if there are three audio features, the order in the sequence set includes: A1 feature - A2 feature - A3 feature, A1 feature - A2 feature - A2 feature, A2 feature - A1 feature - A3 feature, A2 feature - A3 feature - A1 feature, A3 feature - A1 feature - A2 feature, and A3 feature - A2 feature - A1 feature.

[0146] It should be noted here that the snoring recognition method provided in this application generally determines the excitation location based on three characteristics of the snoring audio frequency: frequency domain characteristics, MFCC characteristics, and GFCC characteristics. The first order can be any of the following: frequency domain characteristics - MFCC characteristics - GFCC characteristics, frequency domain characteristics - GFCC characteristics - MFCC characteristics, MFCC characteristics - frequency domain characteristics - GFCC characteristics, MFCC characteristics - GFCC characteristics - frequency domain characteristics, GFCC characteristics - frequency domain characteristics - MFCC characteristics, or GFCC characteristics - MFCC characteristics - frequency domain characteristics.

[0147] Step S402: Extract the target features to be extracted from the snoring verification audio in the verification set according to the first order;

[0148] The target feature to be extracted is related to the first sequence and can be any one of the first, second, and third features. That is, it can be any one of the frequency domain features, MFCC features, and GFCC features. The target feature can be obtained by extracting the corresponding features using the aforementioned feature extraction methods.

[0149] Step S403: Calculate the similarity between the target feature to be extracted and the first audio feature of the corresponding snoring sample audio in the training set. The target feature to be extracted and the first audio feature are of the same type.

[0150] If the target feature to be extracted is the first feature, then the first feature of the snoring verification audio can be compared with the first feature of the snoring sample audio. The comparison method is to calculate the similarity between the first feature of the snoring verification audio and the first feature of the snoring sample audio in sequence using the above formula.

[0151] For example, if there are 100 snoring verification audio samples, and the first order is frequency domain features - MFCC features - GFCC features, then 100 frequency domain features can be extracted first. Then, the similarity of one of these frequency domain features is calculated with all the frequency domain features of the snoring sample audio samples. A preset number of snoring sample audio samples with a similarity greater than a preset threshold are selected. Based on the similarity calculation and selection method, all target features to be extracted in the snoring verification audio samples that are similar to the snoring sample audio samples are obtained.

[0152] For example: the first feature of the snoring verification audio is A1, A2, A3, A4, ... A100; the first feature of the snoring sample audio is B1, B2, B3, B4, ... B1000. The similarity between A1 and B1, B2, B3, B4, ... B1000 is calculated, and the following are selected: B1, B4, B7 are similar to A1; B2, B10, B500 are similar to A2; B5, B20, B50 are similar to A3; ...

[0153] Step S404: Determine whether the proportion of the first audio feature with a similarity greater than the target feature to be extracted in the positive sample set is greater than a predetermined threshold.

[0154] The preset threshold can be any one of 100%, 95%, 90%, 88%, 80%, etc. Based on the features similar to the target features to be extracted from the snoring verification audio determined above, the corresponding positive sample set and confused sample set are obtained by dividing the audio features of the snoring sample audio as described above. It is then determined whether the first audio feature similar to the target feature to be extracted is located in the positive sample set. The proportion of the first audio feature similar to the target feature to be extracted located in the positive sample set is determined and compared with the preset threshold.

[0155] Step S405: If the value is greater than the target feature to be extracted, determine the order in which the target feature to be extracted is used for snoring recognition based on its position in the first sequence.

[0156] If the above proportion is greater than a preset threshold, the target features to be extracted in the first order will be used as the features to determine the excitation location of the audio that triggers snoring.

[0157] Step S406: If it is not greater than, extract the next feature to be extracted from the snoring verification audio in the verification set according to the first order, take the next feature to be extracted as the target feature to be extracted, and continue to determine the order of the re-determined target features to be extracted for snoring recognition until the next feature to be extracted is the last feature in the first order.

[0158] If the proportion of the first audio features similar to the target feature in the positive sample set is not greater than a preset threshold, the next feature to be extracted in the first sequence is then extracted. The next feature to be extracted is processed through the same steps as the target feature, and it is determined whether the next feature is the audio feature that determines the excitation site of the corresponding snoring verification audio. If it is determined that the feature is not the excitation site, the audio features following the next feature in the first sequence are extracted, and similarity calculation, filtering, and comparison are performed using the same method. If it is determined that the audio feature preceding the last audio feature in the first sequence is also not the excitation site, the last feature in the first sequence is used as the feature that determines the excitation site of the snoring verification audio. This process continues until all features that trigger the excitation sites of the snoring verification audio are determined.

[0159] Finally, the other sequences in the sequence set are calculated, filtered, and compared in the same way as described above to obtain the characteristics of the excitation sites of the snoring verification audio corresponding to the other sequences.

[0160] For example, based on the above method, the features corresponding to the first order in the sequence set that determine the excitation sites of each snoring verification audio are A feature, B feature, C feature, A feature, B feature, C feature, A feature, B feature, C feature, ..., C feature; the features corresponding to the second order that determine the excitation sites of each snoring verification audio are B feature, B feature, A feature, A feature, C feature, C feature, C feature, B feature, B feature, ..., B feature; the features corresponding to the third order that determine the excitation sites of each snoring verification audio are... Feature A, Feature B, Feature C, Feature C, Feature C, Feature A, Feature C, Feature B, Feature A, ..., Feature A; the fourth sequence corresponds to the features that determine the excitation location of each snoring verification audio, which are: Feature C, Feature C, Feature B, Feature B, Feature A, Feature A, Feature C, Feature B, Feature A, ..., Feature A; the fifth sequence corresponds to the features that determine the excitation location of each snoring verification audio, which are: Feature A, Feature B, Feature B, Feature A, Feature C, Feature C, Feature C, Feature B, Feature A, ..., Feature B; ...

[0161] Step S407: Compare the accuracy of the feature matching order corresponding to each order in the sequence set. Based on the accuracy, determine the target feature matching order in the sequence set. The target feature matching order is the feature matching order with the highest accuracy.

[0162] The process involves determining the features corresponding to each sequence based on the aforementioned method, then acquiring labels for each snoring verification audio. Since the labels for each snoring verification audio include audio features that identify the excitation sites that trigger each snoring verification audio, the audio features in the labels of each snoring verification audio are carefully compared with the audio features of each sequence. The accuracy rate of audio with the same label in each sequence is calculated, and finally, the sequence with the highest accuracy rate is determined as the target feature comparison sequence. For example, if the sequence set includes six sequences: the first sequence has an accuracy rate of 80%, the second sequence has an accuracy rate of 85%, the third sequence has an accuracy rate of 70%, the fourth sequence has an accuracy rate of 65%, the fifth sequence has an accuracy rate of 90%, and the sixth sequence has an accuracy rate of 75%, then the fifth sequence is determined as the target feature comparison sequence.

[0163] Step S500: Based on the target feature comparison order, compare the audio features of the snoring audio of the patient to be tested, and determine the triggering site of the snoring of the patient to be tested through the comparison results.

[0164] In one embodiment, it may be based on, for example Figure 6 The method shown obtains the audio characteristics of the snoring audio of the patient being tested:

[0165] Step S501: Determine the target features to be compared based on the target feature comparison order;

[0166] In this process, the target feature comparison order is determined based on the above method. For example, if the target feature comparison order is MFCC feature-frequency domain feature-GFCC feature, then after obtaining the snoring audio of the patient to be tested, the target feature of the snoring audio is determined to be MFCC feature.

[0167] Step S502: Extract the target features of the snoring audio of the patient to be tested, and compare the target features with the first audio features of each snoring sample audio. The target features and the first audio features are of the same type.

[0168] Then, the MFCC features of the snoring audio are extracted using the method described above, and the similarity between the MFCC features of the snoring audio and the MFCC features of each snoring sample audio is calculated using the method described above.

[0169] Step S503: If the comparison result does not meet the preset conditions, the target feature to be compared is determined again based on the target feature comparison order, and the comparison is performed based on the newly determined target feature until the preset conditions are met.

[0170] The preset conditions include: the similarity between the target feature and the first audio feature of a preset number of snoring sample audios is greater than a preset similarity threshold, and the proportion of the first audio feature of a preset number of snoring sample audios in the positive sample set is greater than a predetermined threshold.

[0171] Based on the above comparison process, a predetermined number of snoring sample audios with MFCC features greater than a threshold similarity to the snoring audio are obtained. Finally, it is determined whether the proportion of the MFCC features of the predetermined number of snoring sample audios in the positive sample set is greater than a predetermined threshold. If it is, the MFCC feature can be used as the feature to determine the excitation site of the snoring audio in the test patient. If it is not, frequency domain features are extracted, and the similarity between the frequency domain features of the snoring audio and the frequency domain features of each snoring sample audio is calculated using the above similarity calculation method. A predetermined number of snoring sample audios with frequency domain features greater than a threshold similarity to the snoring audio are obtained. Finally, it is determined whether the proportion of the frequency domain features of the predetermined number of snoring sample audios in the positive sample set is greater than a predetermined threshold. If it is, the frequency domain feature can be used as the feature to determine the excitation site of the snoring audio in the test patient. If it is not, the MFCC feature is directly used as the feature to determine the excitation site of the snoring audio in the test patient. This application allows for the identification of the triggering site for snoring audio in different patients by using different audio characteristics, enabling more targeted identification of the patient and thus improving the accuracy of the triggering site identification.

[0172] After determining the target features of the patient to be tested, the excitation sites that trigger the snoring audio of the patient to be tested can be identified from the corresponding features of snoring audio samples with a similarity greater than a preset threshold. For example, if the target feature of the patient to be tested is a frequency domain feature, and the features of the snoring audio samples with a similarity greater than a threshold are A5, A20, and A100, where A5 determines the excitation site as the palate, A20 determines the excitation site as the palate, and A100 determines the excitation site as the root of the tongue, then the excitation site that triggers the snoring audio of the patient to be tested is determined to be the palate. Other methods can also be used to determine the excitation site that triggers the snoring audio of the patient to be tested, and this application does not limit this method.

[0173] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0174] Based on the same inventive concept, this application also provides a snoring recognition device based on multi-feature fusion for implementing the aforementioned method. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more data backup device embodiments provided below can be found in the limitations of the snoring recognition method based on multi-feature fusion described above, and will not be repeated here.

[0175] In one embodiment, such as Figure 7 As shown, a snoring recognition device 700 based on multi-feature fusion is provided. The snoring recognition device 700 includes: a first acquisition module 701, an extraction module 702, a second acquisition module 703, a determination module 704, and a comparison and determination module 705.

[0176] The first acquisition module 701 is used to acquire a training set and a validation set. The training set includes: multiple snoring sample audios, and sample labels corresponding to each snoring sample audio. The sample labels are used to indicate the first excitation part of the snoring sample audio, and to indicate the first target feature for determining the first excitation part. The validation set includes: multiple snoring validation audios, and validation labels corresponding to each snoring validation audio. The validation labels include the second excitation part of the snoring validation audio, and the second target feature for determining the second excitation part.

[0177] Extraction module 702 is used to extract at least two types of first audio features corresponding to each snoring sample audio.

[0178] The second acquisition module 703 is used to acquire positive sample sets and confused sample sets corresponding to each type of first audio feature based on the first audio features and sample labels of different snoring sample audios. The positive sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is greater than or equal to a preset threshold. The confused sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is less than a preset threshold.

[0179] The determination module 704 is used to determine the target feature comparison order when performing snoring recognition based on the positive sample set, the confused sample set, and the validation set;

[0180] The comparison and determination module 705 is used to compare the audio features of the snoring audio of the test patient based on the target feature comparison order, and determine the triggering site of the snoring of the test patient through the comparison results.

[0181] In one embodiment, the extraction module 702 is specifically used to extract the frequency domain characteristics of each snoring sample audio using the fractional Fourier transform method to obtain the first feature corresponding to each snoring sample audio.

[0182] In one embodiment, the extraction module 702 is specifically used to perform frame-segmentation processing on each snoring sample audio to obtain a first audio corresponding to each snoring sample audio.

[0183] By applying a Hamming window to each of the first audio frequencies, multiple second audio frequencies can be obtained.

[0184] Perform a Fast Fourier Transform on each of the second audio frequencies to obtain multiple discrete power spectra;

[0185] The discrete power spectra are squared to obtain multiple energy spectra, and each energy spectrum is filtered by a pre-constructed first filter group.

[0186] Cepstral analysis was performed on the filtered energy spectra to obtain the second feature corresponding to the audio of each snoring sample.

[0187] In one embodiment, the extraction module 702 is specifically used to square each discrete power spectrum to obtain multiple energy spectra, and to filter each energy spectrum through a pre-constructed second filter group.

[0188] Cepstral analysis was performed on the filtered energy spectra to obtain the third feature corresponding to the audio of each snoring sample.

[0189] In one embodiment, the determining module 704 is specifically used to determine a first order, which is any one of the preset order sets, and the order is used to indicate the extraction order of various features;

[0190] Extract the target features to be extracted from the snoring verification audio in the verification set according to the first order;

[0191] Calculate the similarity between the target feature to be extracted and the first audio feature of the corresponding snoring sample audio in the training set, wherein the target feature to be extracted and the first audio feature are of the same type;

[0192] Determine whether the proportion of a first audio feature with a similarity greater than a similarity threshold to the target feature to be extracted in the positive sample set is greater than a predetermined threshold;

[0193] If the value is greater than the target feature to be extracted, the order in which the target feature to be extracted is used for snoring recognition is determined based on the position of the target feature to be extracted in the first order.

[0194] If it is not greater than, extract the next feature to be extracted from the snoring verification audio in the verification set according to the first order, take the next feature to be extracted as the target feature to be extracted, and continue to determine the order of the re-determined target features to be extracted for snoring recognition until the next feature to be extracted is the last feature in the first order.

[0195] The accuracy of each sequence in the sequence set corresponding to the feature matching sequence is compared. Based on the accuracy, the target feature matching sequence in the sequence set is determined. The target feature matching sequence is the feature matching sequence with the highest accuracy.

[0196] In one embodiment, the comparison determination module 705 is specifically used to determine the target features to be compared based on the target feature comparison order;

[0197] The target features of the snoring audio of the patient to be tested are extracted and compared with the first audio features of each snoring audio sample. The target features and the first audio features are of the same type.

[0198] If the comparison result does not meet the preset conditions, the target features to be compared are determined again based on the target feature comparison order, and the comparison is performed based on the newly determined target features until the preset conditions are met.

[0199] The modules in the aforementioned snoring recognition device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0200] In one embodiment, a computer device is provided, which may be a server, and the internal structure diagram of the server may be as follows: Figure 8 As shown, the server includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The server's database stores individual snoring audio recordings and their characteristics. The server's network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a snoring recognition method based on multi-feature fusion.

[0201] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the solution of this application and does not constitute a limitation on the server to which the solution of this application is applied. A specific server may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0202] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the steps of the above-described snoring recognition method based on multi-feature fusion.

[0203] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements any of the steps of the above-described snoring recognition method based on multi-feature fusion.

[0204] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the steps of the above-described snoring recognition method based on multi-feature fusion.

[0205] It is readily understood that, based on the several embodiments provided in this application, those skilled in the art can combine, split, or reorganize the embodiments of this application to obtain other embodiments, none of which exceed the protection scope of this application.

[0206] The above detailed embodiments further illustrate the purpose, technical solution, and beneficial effects of the embodiments of this application. It should be understood that the above are merely specific embodiments of the embodiments of this application and are not intended to limit the protection scope of the embodiments of this application. Any modifications, equivalent substitutions, improvements, etc., made on the basis of the technical solutions of the embodiments of this application should be included within the protection scope of the embodiments of this application.

Claims

1. A snoring recognition method based on multi-feature fusion, characterized in that, include: A training set and a validation set are obtained. The training set includes: multiple snoring sample audios, and sample labels corresponding to each of the snoring sample audios. The sample labels are used to indicate the first excitation part of the snoring sample audios and to indicate the first target feature for determining the first excitation part. The validation set includes: multiple snoring verification audios, and validation labels corresponding to each of the snoring verification audios. The validation labels include the second excitation part of the snoring verification audios and the second target feature for determining the second excitation part. Extract at least two types of first audio features corresponding to each of the snoring sample audios; Based on the first audio features of different snoring sample audios and the sample labels, obtain positive sample sets and confused sample sets corresponding to each type of first audio feature. The positive sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is greater than or equal to a preset threshold. The confused sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is less than a preset threshold. Based on the positive sample set, the confused sample set, and the validation set, the target feature comparison order for snoring recognition is determined. Specifically, a first order is determined, which is any order from a preset order set, indicating the extraction order of various features. Target features to be extracted from the snoring verification audio in the validation set are extracted according to the first order. The similarity between the target features to be extracted and the first audio features of the corresponding snoring sample audio in the training set is calculated, wherein the target features to be extracted and the first audio features are of the same type. It is determined whether the proportion of first audio features with a similarity greater than a similarity threshold in the positive sample set is greater than a predetermined threshold. The accuracy rates of the feature comparison orders corresponding to each order in the order set are compared, and based on the accuracy rates, the target feature comparison order in the order set is determined, wherein the target feature comparison order is the feature comparison order with the highest accuracy. Based on the target feature comparison order, the audio features of the snoring audio of the test patient are compared, and the triggering site of the snoring of the test patient is determined by the comparison results.

2. The snoring recognition method according to claim 1, characterized in that, The first audio feature includes a first feature, and at least two types of first audio features are extracted for each of the snoring sample audios, including: The frequency domain characteristics of each snoring audio sample are extracted using the fractional Fourier transform method to obtain the first feature corresponding to each snoring audio sample.

3. The snoring recognition method according to claim 1, characterized in that, The first audio feature includes a second feature, and at least two types of first audio features are extracted for each of the snoring sample audios, including: The audio samples of each snoring sound are segmented into frames to obtain a first audio corresponding to each of the audio samples of each snoring sound. Each of the first audio samples is subjected to Hamming windowing to obtain multiple second audio samples; Perform a Fast Fourier Transform on each of the second audio frequencies to obtain multiple discrete power spectra; The discrete power spectra are squared to obtain multiple energy spectra, and each energy spectrum is filtered by a pre-constructed first filter group. Cepstral analysis is performed on the filtered energy spectra to obtain the second feature corresponding to the audio of each snoring sample.

4. The snoring recognition method according to claim 3, characterized in that, The first audio feature includes a third feature, and at least two types of first audio features are extracted for each of the snoring sample audios, including: The discrete power spectra are squared to obtain multiple energy spectra, and each energy spectrum is filtered by a pre-constructed second filter group. Cepstral analysis is performed on the filtered energy spectra to obtain the third feature corresponding to the audio of each snoring sample.

5. The snoring recognition method according to claim 1, characterized in that, The target similarity is determined by the following formula: Among them, X i Y i It is the feature vector of the first audio feature of the same type for any two different snoring audio samples in each of the snoring audio samples.

6. The snoring recognition method according to claim 1, characterized in that, The step of determining whether the proportion of a first audio feature with a similarity greater than a similarity threshold to the target feature to be extracted in the positive sample set is greater than a predetermined threshold includes: If the value is greater than the target feature to be extracted, the order in which the target feature to be extracted is used for snoring recognition is determined based on the position of the target feature to be extracted in the first order. If the value is not greater than the target feature, the next feature to be extracted from the snoring verification audio in the verification set is extracted according to the first order. The next feature to be extracted is taken as the target feature to be extracted, and the order of the re-determined target features to be extracted for snoring recognition is determined until the next feature to be extracted is the last feature in the first order.

7. The snoring recognition method according to claim 1 or 6, characterized in that, The comparison of audio features of the snoring audio of the patient under test based on the target feature comparison order includes: Based on the target feature comparison order, the target features to be compared are determined; Extract the target features of the snoring audio of the patient to be tested, and compare the target features with the first audio features of each of the snoring audio samples, wherein the target features and the first audio features are of the same type; If the comparison result does not meet the preset conditions, the target feature to be compared is determined again based on the target feature comparison order, and the comparison is performed based on the newly determined target feature until the preset conditions are met.

8. The method according to claim 7, characterized in that, The preset conditions include: the similarity between the target feature and the first audio feature of a preset number of snoring sample audios is greater than a preset similarity threshold, and the proportion of the first audio feature of the preset number of snoring sample audios in the positive sample set is greater than a predetermined threshold.

9. A snoring recognition device based on multi-feature fusion, characterized in that, The device includes: The first acquisition module is used to acquire a training set and a validation set. The training set includes: multiple snoring sample audios, and sample labels corresponding to each of the snoring sample audios. The sample labels are used to indicate the first excitation part of the snoring sample audios and to indicate the first target feature for determining the first excitation part. The validation set includes: multiple snoring verification audios, and verification labels corresponding to each of the snoring verification audios. The verification labels include the second excitation part of the snoring verification audios and the second target feature for determining the second excitation part. The extraction module is used to extract at least two types of first audio features corresponding to each of the snoring sample audios; The second acquisition module is used to acquire, based on the first audio features of different snoring sample audios and the sample labels, a positive sample set and a confused sample set corresponding to each type of first audio feature, respectively. The positive sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is greater than or equal to a preset threshold. The confused sample set includes first audio features whose label similarity to the target satisfies the first condition and whose label similarity is less than a preset threshold. A determining module is used to determine the target feature comparison order for snoring recognition based on the positive sample set, the confused sample set, and the validation set. Specifically, it determines a first order, which is any order from a preset order set, indicating the extraction order of various features. Based on the first order, it extracts target features to be extracted from the snoring validation audio in the validation set. It calculates the similarity between the target features to be extracted and the first audio features of the corresponding snoring sample audio in the training set, wherein the target features to be extracted and the first audio features are of the same type. It determines whether the proportion of first audio features with a similarity greater than a similarity threshold in the positive sample set is greater than a predetermined threshold. It compares the accuracy of each order in the order set corresponding to the feature comparison order, and based on the accuracy, determines the target feature comparison order in the order set, wherein the target feature comparison order is the feature comparison order with the highest accuracy. The comparison and determination module is used to compare the audio features of the snoring audio of the patient under test with the target feature comparison order, and determine the triggering site of the snoring of the patient under test through the comparison results.

10. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores executable program code, and the processor is used to execute the executable program code to implement the steps of the method described in any one of claims 1-8.