Nighttime teeth grinding real-time diagnosis method and system based on non-verbal audio feature recognition

By using a non-verbal audio feature recognition method and employing audio acquisition and machine learning models, a non-contact, real-time, and accurate diagnosis of bruxism was achieved. This solves the problems of low specificity and inability to diagnose bruxism in existing technologies, and provides a basis for early prevention and treatment.

CN116312647BActive Publication Date: 2026-04-07FUDAN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-09
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing diagnostic methods for bruxism have low specificity and cannot achieve non-contact real-time diagnosis, failing to meet the needs of healthy bruxism patients. Furthermore, existing portable devices cause discomfort during sleep.

Method used

This method employs non-verbal audio feature recognition, which uses audio acquisition, framing and downsampling, feature extraction and machine learning models to diagnose bruxism and other sleep states in real time. It uses a microphone to collect audio data and a machine learning model to identify bruxism sounds and provide diagnosis.

Benefits of technology

It enables contactless, real-time, and accurate diagnosis of bruxism, reduces interference with patients' sleep, provides a basis for early prevention and treatment of bruxism, and can distinguish bruxism sounds from other sleep sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312647B_ABST
    Figure CN116312647B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of computer-aided diagnostic technology, specifically a real-time diagnostic method and system for bruxism based on non-verbal audio feature recognition. The invention includes: acquiring audio in a sleep environment; converting non-stationary signals into short-time stationary signals through frame segmentation; extracting multiple time-domain and frequency-domain features from the audio samples, and summarizing the variance and mean distribution feature groups by combining multiple feature arrays from each different frame; and optimizing the best predictive model for bruxism diagnosis using a machine learning model. This invention can accurately identify the occurrence of bruxism and other sleep states in complex sleep environments without requiring additional computational resources for preprocessing such as noise reduction and filtering of the audio signals. It enables non-contact monitoring and diagnosis of bruxism with a simple process, avoiding patient discomfort while providing practical reference and theoretical analysis for subsequent interventions for bruxism.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer-aided diagnosis, and particularly relates to a real-time diagnosis method and system for nocturnal bruxism based on non-verbal audio feature recognition. BACKGROUND

[0002] According to the international consensus definition for evaluating bruxism in 2018, nocturnal bruxism is the rhythmic (phasic) or non-rhythmic (tonic) activity of masticatory muscles during the sleep process of the human body, which is not a movement disorder and sleep disorder in healthy individuals, but increases the risk of other diseases, such as tooth wear, masticatory muscle hypertrophy, temporomandibular joint dysfunction, facial muscle soreness, and even causes headache, etc. According to reliable literature research, the population with nocturnal bruxism is widely distributed, and the proportion of adolescents accounts for 8%. In order to reduce the negative impact of nocturnal bruxism on others and themselves, it is of great significance to realize correct diagnosis and treatment of nocturnal bruxism and to the life and health.

[0003] At present, the diagnosis methods for bruxism mainly include questionnaire survey, clinical evaluation, and portable device diagnosis. The questionnaire survey and clinical evaluation have too low specificity for the diagnosis of bruxism, and they are used to evaluate the situation after the damage caused by bruxism, and cannot prevent the occurrence of bruxism. The portable device based on surface electromyography (SEMG) is widely studied, and the real-time diagnosis of bruxism and the detection of the degree of bruxism based on electromyography have made progress in methods and technology within a certain range, but the algorithms for diagnosing bruxism based on surface electromyography signals at home and abroad excessively rely on the equipment for collecting muscle signals, that is, the electromyography report record of bruxism lacks authority and uniformity, and the algorithm for identifying bruxism is only applicable to the equipment for collecting muscle signals developed by the author of the algorithm, and is not applicable to the nocturnal bruxism signals collected by other electromyography devices. Therefore, there is almost no nocturnal bruxism data and mature portable device for diagnosing bruxism based on surface electromyography signals. Most importantly, the current bruxism population belongs to healthy individuals (i.e. the bruxism behavior does not cause pathological harm to the patients), and most healthy bruxism patients do not have serious pathological behavior, and only have the trouble of disturbing the partner or rest caused by the grinding sound. In this case, the healthy bruxism patients will feel uncomfortable in sleep when wearing contact devices such as electromyography sensors and pressure sensors.

[0004] Therefore, a mature non-contact diagnosis method is needed to detect the occurrence of nocturnal bruxism, and to realize real-time diagnosis of the occurrence of nocturnal bruxism, so as to make a substantial contribution to reducing the occurrence of bruxism from the source, thereby meeting the needs of the majority of healthy bruxism patients, and also monitoring other sleep states such as sneezing, yawning, sighing, groaning, and clearing the throat. SUMMARY

[0005] In order to improve the existing diagnosis method of bruxism, overcome the low specificity of current clinical diagnosis of bruxism and the deficiency that it cannot realize non-contact real-time diagnosis, the application provides a portable, non-contact real-time diagnosis method and system for nocturnal bruxism (including related sleep states, the same below) based on non-verbal audio feature recognition.

[0006] The real-time diagnosis method and system for nocturnal bruxism based on non-verbal audio feature recognition provided by the application can not only quickly identify original lossless audio, but also has stable robustness and precision. The application can monitor the occurrence of bruxism in real time, add a new applicable algorithm to the portable bruxism diagnosis device, and provide an important basis for the subsequent research on the early prevention and treatment of nocturnal bruxism.

[0007] The real-time diagnosis method for nocturnal bruxism based on non-verbal audio feature recognition provided by the application comprises the following specific steps:

[0008] Step S1, audio acquisition: using an audio acquisition module to acquire audio data in a sleep environment;

[0009] Step S2, frame division and down-sampling: the acquired audio data is converted into a short-time stationary audio signal through a frame division method in the time domain, and the sampling frequency is reduced by half;

[0010] Step S3, feature extraction: extracting time domain features, frequency domain features and time-frequency domain features from the audio samples after frame division;

[0011] Step S4, feature group summary: according to the multiple feature arrays of each different frame, multiple time domain variance distribution and mean value distribution feature groups, and multiple frequency domain variance distribution and mean value distribution feature groups are summarized;

[0012] Step S5, machine learning model optimization: based on the above-mentioned time domain, frequency domain and time-frequency domain feature groups, multiple machine learning models are input, and the optimal model for predicting nocturnal bruxism is selected, which is used for diagnosis and evaluation of nocturnal bruxism (and other sleep states).

[0013] Further,

[0014] In step S1, the audio acquisition module is used to acquire audio data in a sleep environment, which means that the sound in the sleep environment is collected through a microphone in a natural sleep environment. The microphone is fixed in the central area of the bed, 30-50 cm above the head of the nocturnal bruxism patient, the sampling frequency of the acquired audio data is 44100HZ, the sampling bit number is 16, the audio format is lossless WAV format, and the length of the audio data sample captured at a time is 1s.

[0015] In step S2, the collected audio data is converted into short-time stationary audio signals by a frame method in time domain, and the sampling frequency is reduced by half; specifically including: reducing the sampling of the collected audio data; converting the non-stationary audio signal into a short-time stationary audio signal by a frame method in time domain, and the amount of one frame of data sampled is 1024 points, and the moving step of each frame is 512 points.

[0016] In step S3, the time domain features and frequency domain features of the audio samples after frame are extracted, and the time-frequency domain features are extracted, including the following sub-steps:

[0017] In step S3-1, the time domain features are extracted, and the time domain features are AE(Amplitude envelope), EMS(Root-mean square energy) and ZCR(Zero crossing rate), and the specific calculation method is as follows:

[0018] The AE calculation formula is:

[0019]

[0020] The EMS calculation formula is:

[0021]

[0022] The ZCR calculation formula is:

[0023]

[0024] In the above formula, s(k) represents a speech signal, k is the coordinate of the number of speech signal sampling points, K represents the number of sampling points of a frame of speech signal, and t represents the speech signal of the tth frame; that is, tk represents the position of the first sampling point of the tth frame, (t+1)*k-1 represents the position of the last sampling point of the tth frame, and sgn is a sign function.

[0025] In step S3-2, the frequency domain features are extracted, and the frequency domain features are BER(Band Energy Ratio), SC(Spectral centroid), BW(Bandwidth) and SR(Spectral Roll-off), and the calculation method is as follows:

[0026] The BER calculation formula is:

[0027]

[0028] The SC calculation formula is:

[0029]

[0030] BW is calculated as:

[0031]

[0032] SR is calculated as:

[0033]

[0034] In the above equations, t represents the t-th frame of speech signal, m t (n) represents the t-th frame of signal, n is the n-th sampling point, F represents the selected separation frequency in the frequency band ratio, N represents the number of sampling points per frame, f c represents the spectral point satisfying the inequality, m i represents the energy.

[0035] In the frequency domain feature, the complementary harmonic and the perceptual (Harmonics and Perceptrual) feature, the calculation formula is:

[0036] H(t) = A * sin(2 * pi * f * t + phi), (8)

[0037] S = K * (10 L / 10 ), (9)

[0038] H(t) is the amplitude function of the harmonic, which represents the amplitude of the harmonic at time t. A is the amplitude of the harmonic, which represents the maximum amplitude of the harmonic. f is the frequency of the harmonic, with the unit of hertz. t is the time, with the unit of second. phi is the phase of the harmonic, which represents the phase difference of the harmonic relative to the fundamental frequency. S is the sound pressure level, L is the sound intensity level, and K is a constant.

[0039] Step S3-3, extract the time-frequency domain feature, i.e. Mel Frequency Cepstral Coefficients (MFCC). Perform FFT transform on each frame in step S2, obtain the frequency spectrum, and then obtain the amplitude spectrum. Apply the Mel filter bank to the amplitude spectrum, take the logarithm of the filter bank output, and finally perform Discrete Cosine Transform (DCT) to obtain MFCC. The expression is as follows:

[0040] L is the number of filters, (10)

[0041] where i is the i-th MFCC coefficient, N is the number of sampling points per frame, l is the l-th filter, m(l) represents the output of each filter, and L is the total number of filters.

[0042] In step S4, the feature groups of the time domain variance distribution and mean distribution and the feature groups of the frequency domain variance distribution and mean distribution are summarized according to the feature arrays of each different frame. Specifically, the features of each frame are calculated according to the feature calculation formula in step S3 to obtain the feature arrays. For example, for the feature EMS, the feature arrays are EMS-1, EMS-2, EMS-3, EMS-4, …, EMS-i, where i is the frame number. The feature groups of the variance distribution and mean distribution, such as EMS_Mean and MES_Var, are obtained through the feature arrays. Similarly, other feature groups are obtained. Through the feature arrays, the feature groups of the variance distribution and mean distribution (such as EMS_Mean and MES_Var) are obtained, and the following total number of features can be obtained, which is 56 feature values in total, specifically as follows:

[0043] 'envelope_mean', 'envelope_var', 'RMS_mean', 'RMS_var', 'zero_mean', 'zero_var',

[0044] 'centroids_mean', 'centroids_var', 'bandwith_mean', 'bandwidth_var', 'rolloff_mean',

[0045] 'rolloff_var', 'harmonics_mean', 'harmonics_var', 'perceptrual_mean', 'perceptrual_var',

[0046] 'mfcc1_mean','mfcc2_mean','mfcc3_mean','mfcc4_mean','mfcc5_mean','mfcc6_mean',

[0047] 'mfcc7_mean','mfcc8_mean','mfcc9_mean','mfcc10_mean','mfcc11_mean','mfcc12_mean',

[0048] 'mfcc13_mean','mfcc14_mean','mfcc15_mean','mfcc16_mean','mfcc17_mean','mfcc18_mean','mfcc19_mean','mfcc20_mean',

[0049] 'mfcc1_var','mfcc2_var','mfcc3_var','mfcc4_var','mfcc5_var','mfcc6_var','mfcc7_var',

[0050] 'mfcc8_var','mfcc9_var','mfcc10_var','mfcc11_var','mfcc12_var','mfcc13_var',

[0051] 'mfcc14_var','mfcc15_var','mfcc16_var','mfcc17_var','mfcc18_var','mfcc18_var','mfcc20_var';

[0052] In step S5, the method of machine learning constructs a prediction optimal model for bruxism and other sleep states, for the diagnosis of bruxism and the evaluation of other sleep states, specifically:

[0053] After the correlation analysis of the characteristic values, the characteristic values in step S4 are input into a typical machine learning network model, which includes but is not limited to: Bayesian model (Bayes), K-Nearest Neighbor (KNN), Stochastic Gradient Descent (Stochastic Gradient Descent), Decision Tree (Decision), Random Forest (Random Forest), Support Vector Machine (Support Vector Machine), Logistic Regression (Logistic Regression), Neural Network (Neural Nets), Catboost, Cross Gradient Booster, Cross Gradient Booster (RandomForest); among the 11 models, the best model is selected for the application of non-verbal audio feature recognition in real-time diagnosis of bruxism and other sleep states. The standard for selecting the optimal model from different models is the value of F-score, which is based on the comprehensive trade-off value of precision (Precision) and recall (Recall). The introduction of F-Score as a comprehensive index is to balance the influence of accuracy and recall, and to evaluate a classifier more comprehensively. F-score is the harmonic mean of precision and recall.

[0054] The calculation formula of F-score (F1) is:

[0055] The calculation formula of F-score (F1) is:

[0056]

[0057]

[0058] Wherein, TP, TN, FP and FN are the number of true positive, true negative, false positive and false negative respectively. Wherein, true positive (TP) is a positive sample predicted as a positive class by the model, true negative (TN) is a negative sample predicted as a negative class by the model, false positive (FP) is a negative sample predicted as a positive class by the model, and false negative (FN) is a positive sample predicted as a negative class by the model.

[0059] In the application, the workflow of the machine learning network model training is as follows: data augmentation is performed on the audio sample to obtain a data set, and the data set is divided into a training set and a test set; the training set data set is down-sampled and framed to facilitate subsequent algorithm; after each data set is framed, the characteristic array of each sample is calculated according to the above characteristic value formula; after the characteristic value is analyzed, it is input into different network models for training and analysis, and the test set is used to verify the advantages and disadvantages of different machine learning network models; the optimal network model is selected according to the above F1 index, and the prediction results of the bruxism and other sleep states are analyzed.

[0060] The workflow of the application is as follows: in a sleep environment, the audio acquisition module close to the patient acquires audio sound during sleep; the audio acquisition module acquires 1s of audio data at a time; the MCU performs down-sampling and frame analysis on the audio data; the signal processing unit performs feature analysis on the processed audio data; the characteristic data is input into the trained optimal model, and the model outputs the results in real time, thereby realizing the diagnosis and prediction of the nocturnal bruxism and other sleep states.

[0061] The application also provides a nocturnal bruxism real-time diagnosis system based on the above nocturnal bruxism real-time diagnosis method. The nocturnal bruxism real-time diagnosis system comprises the following modules: an audio acquisition module, a frame and down-sampling module, a feature extraction module, a feature group summarization module and a machine learning model optimization module; the five modules sequentially perform the operations of the five steps of the nocturnal bruxism real-time diagnosis method. That is, the audio acquisition module performs the audio acquisition of step S1; the frame and down-sampling module performs the frame and down-sampling of step S2; the feature extraction module performs the feature extraction of step S3; the feature group summarization module performs the feature group summarization of step S4; and the machine learning model optimization module performs the machine learning model optimization of step S5.

[0062] The nocturnal bruxism real-time diagnosis method and system based on non-verbal audio feature recognition provided by the application have the following advantages:

[0063] 1. This invention diagnoses bruxism and other sleep states in real time by quickly identifying audio signals. The audio acquisition system does not have actual contact with the patient. Compared with other various adhesive portable devices and oral implantable devices, it will not cause any discomfort to the patient's sleep and meets the needs of a large number of bruxism patients without pathological conditions.

[0064] 2. This invention diagnoses bruxism in real time by quickly recognizing audio signals. It can make a diagnosis quickly when bruxism occurs, avoiding the previous diagnosis methods that summarized data from the whole night, but did not actually reduce the harm to the human body caused by nocturnal bruxism.

[0065] 3. This invention uses standard lossless WAV format audio as the input signal, ensuring data accuracy, reliability, and experimental reproducibility;

[0066] 4. The dataset of this invention comes from an open-source database, with abundant experimental data and broad applicability;

[0067] 5. The algorithm part of this invention fully explores the characteristics of machine learning, adopts a variety of commonly used machine learning models, and conducts comparative analysis;

[0068] 6. The algorithm can exhibit strong specificity and robustness. It can distinguish teeth grinding sounds from various other sleep state-related noises in a natural sleep environment with high accuracy, including sneezing, screaming, groaning, crying, yawning, tongue clicking, laughter, lip pursing, throat clearing, lip smacking, nose blowing, coughing, sighing, teeth grinding and chattering, and panting sounds, totaling 16 sleep states.

[0069] 7. This invention can not only monitor the occurrence of bruxism in real time, but also combine with the system to count key information such as the number of bruxisms and energy ratio, which is of great significance for bruxism diagnosis reports and subsequent treatment guidance. It also has the function of monitoring other sleep states. Attached Figure Description

[0070] Figure 1 This is a block diagram of the architecture for real-time diagnosis of nocturnal bruxism based on non-verbal audio feature recognition in this invention.

[0071] Figure 2 A comparative analysis diagram showing the selection of the network model for this invention.

[0072] Figure 3 A confusion matrix diagram for optimal network recognition of teeth grinding sounds (including other sleep states).

[0073] Figure 4 A flowchart of signals for identifying molars and other sleep states.

[0074] Figure 5 Hardware schematics for identifying real-world examples of teeth grinding and other sleep states. Detailed Implementation

[0075] The technical solutions of the embodiments of the present invention will be described in complete and clear form below with reference to the accompanying drawings. Obviously, the described embodiments are only representative embodiments of the present invention, and not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the protection scope of the present invention.

[0076] The embodiments of this invention comprise two parts: a network model training part and an implementation example part. The network model training process includes the following steps:

[0077] Step S1: Use the audio acquisition module to acquire audio data in a sleep environment;

[0078] Step S2: The acquired audio data is converted from a non-stationary audio signal to a short-time stationary audio signal by a time-domain frame-segmentation method, and the sampling frequency is reduced by half.

[0079] Step S3: Extract time-domain and frequency-domain features, as well as time-frequency domain features, from the audio samples after frame segmentation.

[0080] Step S4: Combine the multiple feature arrays of each different frame to summarize the feature groups of multiple time-domain variance distributions and mean distributions, and the feature groups of multiple frequency-domain variance distributions and mean distributions.

[0081] Step S5: Based on the feature groups in the time domain, frequency domain, and time-frequency domain mentioned above, construct the optimal prediction model for bruxism and other sleep states using machine learning methods, which is used for the diagnosis of bruxism and the evaluation of other sleep states.

[0082] In step S1, an audio acquisition module is used to acquire audio data in a sleep environment. Sounds in the sleep environment are acquired through a microphone in a natural sleep environment. The microphone is fixed in the center of the bed, 30-50cm vertically above the head of the patient with bruxism. The audio data is obtained as single-channel data with a sampling frequency of 44100HZ, a sampling bit depth of 16 bits, and a lossless audio format of WAV. The length of the audio data sample captured in a single session is 1 second.

[0083] S2 includes the following two sub-steps: In step S2-1, the acquired audio data is downsampled; In step S2-1, the non-stationary audio signal is converted into a short-time stationary audio signal by a time-domain framing method. The amount of data sampled in one frame is 1024 points, and the movement of each frame is 512 points.

[0084] The feature extraction part of step S3 includes the following sub-steps:

[0085] Step S3-1 extracts time-domain features based on step S2. The time-domain features AE, EMS, ZCR, and their calculation methods are as follows:

[0086] The calculation method for the corresponding AE (Amplitude envelope) is as follows:

[0087]

[0088] The EMS (Root-mean square energy) calculation method is as follows:

[0089]

[0090] The ZCR (Zero Crossing Rate) is calculated as follows:

[0091]

[0092] In the above formula, s(k) represents the speech signal, k is the coordinate of the number of sampling points of the speech signal, K represents the number of sampling points of a frame of speech signal, t represents the speech signal of the t-th frame; that is, tk represents the position of the first sampling point of the t-th frame, (t+1)*k-1 represents the position of the last sampling point of the t-th frame, and sgn is the sign function.

[0093] Step S3-2 extracts frequency domain features BER, SC, BW, and SR based on step S2. The calculation method is as follows:

[0094] The BER (Band Energy Ratio) calculation method is as follows:

[0095]

[0096] The SC (Spectral centroid) calculation method is as follows:

[0097]

[0098] The BW (Bandwidth) is calculated as follows:

[0099]

[0100] The SR (Spectral Roll-off) calculation method is as follows:

[0101]

[0102] In the above formula, t represents the speech signal of the t-th frame, and m t(n) represents the signal of the t-th frame, n is the nth sampling point, F represents the separation frequency selected in the bandwidth ratio, N represents the number of sampling points in each frame, f c m represents the spectral points that satisfy the inequality. i It represents energy.

[0103] In terms of frequency domain characteristics, the harmonics and perception features are supplemented, and their calculation formula is as follows:

[0104] H(t) = A*sin(2*pi*f*t+phi),

[0105] S=K*(10 L / 10 ),

[0106] H(t) is the amplitude function of the harmonic, representing the amplitude of the harmonic at time t. A is the amplitude of the harmonic, representing the maximum amplitude of the harmonic. f is the frequency of the harmonic, in Hertz (Hz). t is the time, in seconds. phi is the phase of the harmonic, representing the phase difference between the harmonic and the fundamental frequency. S is the sound pressure level, L is the sound intensity level, and K is a constant.

[0107] Step S3-3: Extract time-frequency domain features, i.e., Mel-frequency cepstral coefficients (MFCCs). Perform an FFT transform on each frame from step S2 to obtain the spectrum, and then obtain the amplitude spectrum. Add a Mel filter bank to the amplitude spectrum, perform a logarithmic operation on the filter bank output, and finally perform a Discrete Cosine Transform (DCT) to obtain the MFCCs, as shown in the following expression:

[0108] L is the number of filters.

[0109] Where i is the i-th MFCC coefficient, N is the number of sampling points in each frame, l is the l-th filter, m(l) represents the output of each filter, and L is the total number of filters.

[0110] In step S4, based on the feature calculation formula in claim 4, a feature array (e.g., [EMS-1, EMS-2, EMS-3, EMS-4...EMS-i], where i is the frame number) is obtained by calculating the features of each frame (e.g., feature EMS_Mean, MES_Var). The total number of features is 56.

[0111] 'envelope_mean','envelope_var','RMS_mean','RMS_var','zero_mean','zero_var',

[0112] 'centroids_mean','centroids_var','bandwith_mean','bandwidth_var','rolloff_mean',

[0113] 'rolloff_var','harmonics_mean','harmonics_var','perceptrual_mean','perceptrual_var',

[0114] 'mfcc1_mean','mfcc2_mean','mfcc3_mean','mfcc4_mean','mfcc5_mean','mfcc6_mean',

[0115] 'mfcc7_mean','mfcc8_mean','mfcc9_mean','mfcc10_mean','mfcc11_mean','mfcc12_mean',

[0116] 'mfcc13_mean','mfcc14_mean','mfcc15_mean','mfcc16_mean','mfcc17_mean','mfcc18_mean','mfcc19_mean','mfcc20_mean',

[0117] 'mfcc1_var','mfcc2_var','mfcc3_var','mfcc4_var','mfcc5_var','mfcc6_var','mfcc7_var',

[0118] 'mfcc8_var','mfcc9_var','mfcc10_var','mfcc11_var','mfcc12_var','mfcc13_var',

[0119] 'mfcc14_var','mfcc15_var','mfcc16_var','mfcc17_var','mfcc18_var','mfcc18_var','mfcc20_var';

[0120] The summary feature table is as follows:

[0121] Feature class envelope_ RMS zero Centroids Bandwidth Number 2 2 2 2 2 Feature class Roll-off harmonics_ Perceptrual MFCC Total Number 2 2 2 40 56 .

[0122] In step S5, after the correlation analysis of the feature values, the feature values ​​from step S4 are input into a typical machine learning network model, including but not limited to: Bayesian models ( The study compared 11 models, including Bayes, K-Nearest Neighbors (KNN), Stochastic Gradient Descent, Decision Tree, Random Forest, Support Vector Machine, Logistic Regression, Neural Nets, Catboost, Cross Gradient Booster, and Cross Gradient Booster (RandomForest). The analysis determined the optimal model for real-time diagnosis of bruxism and other sleep states using non-verbal audio feature recognition.

[0123] In step S5, the optimal model is selected from different models for real-time diagnosis of bruxism and other sleep states using non-verbal audio feature recognition. The selection criterion for the optimal model is based on the F-score, which is a comprehensive trade-off between precision and recall. Introducing the F-score as a comprehensive indicator is to balance the impact of precision and recall, and to evaluate a classifier more comprehensively. The F-score is the harmonic mean of precision and recall.

[0124] The calculation method for F1 (i.e., F-score) is as follows:

[0125]

[0126]

[0127] In the formula, TP, TN, FP, and FN represent the number of true positives, true negatives, false positives, and false negatives, respectively. TP represents the positive samples predicted as positive by the model, TN represents the negative samples predicted as negative by the model, FP represents the negative samples predicted as positive by the model, and FN represents the positive samples predicted as negative by the model.

[0128] In step S5, after the correlation analysis of the feature values, the feature values ​​from step S4 are input into a typical machine learning network model, including but not limited to: Bayesian models ( The study compared 11 models, including Bayes, K-Nearest Neighbors (KNN), Stochastic Gradient Descent, Decision Tree, Random Forest, Support Vector Machine, Logistic Regression, Neural Nets, Catboost, Cross Gradient Booster, and Cross Gradient Booster (RandomForest). The analysis determined the optimal model for real-time diagnosis of bruxism and other sleep states using non-verbal audio feature recognition.

[0129] Among methods for selecting the optimal model from different models for real-time diagnosis of bruxism and other sleep states using non-verbal audio feature recognition, the selection criterion is based on the F-score. The F-score is a comprehensive trade-off between precision and recall. Introducing the F-score as a comprehensive indicator is to balance the influence of precision and recall, and to evaluate a classifier more comprehensively. The F-score is the harmonic mean of precision and recall.

[0130] The calculation method for F1 (i.e., F-score) is as follows:

[0131] Precision : Recall rate :

[0132]

[0133] In the formula, TP, TN, FP, and FN represent the number of true positives, true negatives, false positives, and false negatives, respectively. TP represents the positive samples predicted as positive by the model, TN represents the negative samples predicted as negative by the model, FP represents the negative samples predicted as positive by the model, and FN represents the positive samples predicted as negative by the model.

[0134] The above network models, combined with the selection of the F1-score criterion, yield the following results:

[0135] Model Score 0 Cat-boost 0.88827 1 Cross Gradient Booster 0.84358 2 Random Forest 0.83240 3 Support Vector Machine 0.76536 4 Logistic Regression 0.74302 5 Cross Gradient Booster(Random Forest) 0.72905 6 Neural Nets 0.71229 7 KNN 0.62011 8 Naive Bayes 0.60335 9 Stochastic Gradient Descent 0.58380 10 Decision trees 0.58380

[0136] Among the 11 commonly used machine learning network models, Catboost's scoring criteria are the most prominent. The following table summarizes its precision and recall on the test set for predicting molars (including teething-chattering and teeth-grinding) and 14 sleep states:

[0137] Precision Recall F1-score coughing 0.85 0.96 0.90 crying 1.00 0.84 0.91 laughing 0.77 0.59 0.67 lip-popping 0.90 0.96 0.93 lip-smacking 0.79 0.83 0.81 moaning 0.85 0.88 0.87 nose-blowing 0.88 0.84 0.86 panting 0.92 0.92 0.92 screaming 1.00 0.97 0.98 sighing 1.00 0.80 0.89 sneezing 0.93 0.81 0.87 teeth-chattering 0.94 0.91 0.92 teeth-grinding 0.83 0.86 0.84 throat-clearing 0.71 0.95 0.82 tongue-clicking 0.96 0.96 0.96 yawning 0.94 0.89 0.91 .

[0138] The embodiments of this invention comprise two parts: a network model training part and an implementation example part. The network model training part is as follows... Figure 1 As shown, the implementation process can be as follows: Figure 5 The following simple system can be built in this way:

[0139] This invention relates to a method for real-time diagnosis of bruxism and other sleep states based on non-verbal audio feature recognition. It employs an audio acquisition module, an MCU microcontroller, a power supply module, a signal processing unit, a storage module, and an OLCD display. The MCU drives the audio acquisition module, acquiring a 1-second unit audio signal in lossless WAV format at a sampling frequency of 44100Hz. After acquiring the unit audio signal, the signal processing unit calculates the feature set of the unit audio signal and stores the data in memory. The feature set is then input into a pre-trained neural network, which outputs and retains specific indicator parameters corresponding to a specific sleep bruxism state. The OLCD display module implements the bruxism recognition command. The MCU... Figure 4 The algorithm identifies the occurrence of teeth grinding and displays the results on the screen. Under the control of the MCU, OLCD can make decisions and display the results in real time for each input signal. The storage module stores the total count of teeth grinding occurrences throughout the night, as well as characteristic parameters. This provides reliable data for doctors' evaluation and subsequent treatment.

Claims

1. A real-time diagnostic method for nocturnal bruxism based on non-verbal audio feature recognition, characterized in that, The specific steps are as follows: Step S1, Audio Acquisition: Use the audio acquisition module to acquire audio data in a sleep environment; Step S2, framing and downsampling: The acquired audio data is framed in the time domain to convert the non-stationary audio signal into a short-time stationary audio signal and reduce the sampling frequency by half. Step S3, Feature Extraction: Extract time-domain features, frequency-domain features, and time-frequency-domain features from the audio samples after frame segmentation; Includes the following sub-steps: Step S3-1: Extract time-domain features, which are AE, EMS, and ZCR. The specific calculation method is as follows: The formula for calculating AE is: The EMS calculation formula is: The formula for calculating ZCR is: In the formula, s(k) represents the speech signal, k is the coordinate of the number of sampling points of the speech signal, K represents the number of sampling points of a frame of speech signal, t represents the speech signal of the t-th frame; that is, t*k represents the position of the first sampling point of the t-th frame, (t+1)*k-1 represents the position of the last sampling point of the t-th frame, and sgn is the sign function. Step S3-2: Extract frequency domain features, which are BER, SC, BW, and SR, calculated as follows: The BER calculation formula is: The formula for calculating SC is: The formula for calculating BW is: The formula for calculating SR is: In the above formula, t represents the speech signal of the t-th frame, and m t (n) represents the signal of the t-th frame, n is the nth sampling point, F represents the separation frequency selected in the bandwidth ratio, N represents the number of sampling points in each frame, f c m represents the spectral points that satisfy the inequality. i Represents energy; In terms of frequency domain characteristics, harmonics and perception features are added, and their calculation formula is as follows: H(t)=A*sin(2*pi*f*t+phi), (8) S=K*(10 L / 10 ), (9) Where H(t) is the amplitude function of the harmonic, representing the amplitude of the harmonic at time t; A is the amplitude of the harmonic, representing the maximum amplitude of the harmonic; f is the frequency of the harmonic, in Hertz; t is time, in seconds; phi is the phase of the harmonic, representing the phase difference of the harmonic relative to the fundamental frequency; S is the sound pressure level; L is the sound intensity level; and K is a constant. Step S3-3: Extract time-frequency domain features, i.e. Mel-frequency cepstral coefficients (MFCCs); perform FFT transform on each frame from step S2 to obtain the spectrum, and then obtain the amplitude spectrum. Add a Mel filter bank to the amplitude spectrum, perform logarithmic operation on the output of the filter bank, and finally perform Discrete Cosine Transform (DCT) to obtain the MFCCs, as shown in the following expression: Where i is the i-th MFCC coefficient, N is the number of sampling points in each frame, l is the l-th filter, m(l) represents the output of each filter, and L is the total number of filters; Step S4, Feature group summarization: Based on the multiple feature arrays of each different frame, summarize the feature groups of multiple temporal variance distributions and mean distributions, as well as the feature groups of multiple frequency domain variance distributions and mean distributions. Step S5, Machine Learning Model Optimization: Based on the above feature groups in the time domain, frequency domain, and time-frequency domain, input multiple machine learning models and select the optimal model for bruxism prediction for diagnosis and evaluation.

2. The real-time diagnostic method for bruxism according to claim 1, characterized in that, The audio acquisition module described in step S1, which collects frequency data in a sleep environment, refers to the process of collecting sound in a natural sleep environment using a microphone. The microphone is fixed in the center of the bed, 30-50cm vertically above the head of the patient with bruxism. The audio data is obtained as single-channel data with a sampling frequency of 44100Hz, a sampling bit depth of 16 bits, and a lossless audio format of WAV. The length of the audio data sample captured in a single instance is 1 second.

3. The real-time diagnostic method for bruxism according to claim 2, characterized in that, In step S2, the non-stationary audio signal is converted into a short-time stationary audio signal by a time-domain framing method, and the sampling frequency is reduced by half. Specifically, this includes: downsampling the acquired audio data; converting the non-stationary audio signal into a short-time stationary audio signal by a time-domain framing method, with one frame of data containing 1024 points and each frame shift being a step size of 512 points.

4. The real-time diagnostic method for bruxism according to claim 1, characterized in that, In step S4, based on the multiple feature arrays of each different frame, feature groups of multiple temporal variance distributions and mean distributions, as well as feature groups of multiple frequency domain variance distributions and mean distributions, are summarized. Specifically, the features of each frame are calculated according to the feature calculation formula in step S3 to obtain a feature array; through the feature array, feature groups of variance distributions and mean distributions are obtained; through the feature array, feature groups of variance distributions and mean distributions are obtained, thus obtaining the following total number of features, a total of 56 feature values.

5. The real-time diagnostic method for bruxism according to claim 4, characterized in that, Step S5 describes constructing an optimal prediction model for bruxism and other sleep states using machine learning methods, for the diagnosis of bruxism and the assessment of other sleep states. Specifically: The feature values ​​from step S4 are input into multiple machine learning network models; these models are compared, and the best model for real-time diagnosis of bruxism is selected; the criterion for selecting the optimal model from different models is the F-score; the F-score is a comprehensive trade-off between precision and recall, and is the harmonic mean of precision and recall. Among them, TP, TN, FP, and FN represent the number of true positives, true negatives, false positives, and false negatives, respectively; F1 is the F-score.

6. The real-time diagnostic method for bruxism according to claim 5, characterized in that, The machine learning network models include Bayesian models, K-nearest neighbors, stochastic gradient descent, decision trees, random forests, support vector machines, logistic regression, neural networks, Catboost, Cross Gradient Booster, and Cross Gradient Booster.

7. A real-time bruxism diagnosis system based on non-verbal audio feature recognition, based on the real-time bruxism diagnosis method according to any one of claims 1-6, characterized in that, It includes the following modules: audio acquisition module, frame segmentation and downsampling module, feature extraction module, feature group summarization module, and machine learning model optimization module; these 5 modules sequentially perform the 5 steps of the real-time bruxism diagnosis method.

Citation Information

Patent Citations

  • Cough and sneeze recognition method for real-time voice streams

    CN111524537A