Voice state intelligent classification method based on voice spectrum characteristics and reinforcement learning optimization mechanism
By combining multi-scale convolution and attention mechanisms with long short-term memory networks, the problem of insufficient feature expression in the staging diagnosis of laryngeal tumors was solved, enabling refined auxiliary diagnosis of laryngeal tumor staging and improving the accuracy and interpretability of early screening.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DALIAN UNIV OF TECH
- Filing Date
- 2026-03-31
- Publication Date
- 2026-05-01
AI Technical Summary
Existing methods lack sufficient feature expression capabilities in the staging diagnosis of laryngeal tumors. The model structure fails to fully integrate spatial and temporal features, the diagnostic results are limited to binary classification, and there is a lack of dynamic attention mechanisms, making it difficult to achieve refined auxiliary diagnosis.
We employ a combination of multi-scale convolution and attention mechanisms with a long short-term memory network to perform intelligent classification of laryngeal tumor staging through voice spectrogram feature extraction and adaptive weighting. We utilize Mel spectrograms for feature extraction and introduce a reinforcement learning optimization mechanism for adaptive model adjustment.
It improves the accuracy and reliability of early screening for laryngeal tumors, enables refined auxiliary diagnosis of laryngeal tumor staging, has good interpretability and scalability, and is applicable to the identification of laryngeal tumors and other voice disorders.
Smart Images

Figure CN121963799A_ABST
Abstract
Description
A Smart Voice State Classification Method Based on Speech Spectrum Features and Reinforcement Learning Optimization Mechanism Technical Field
[0001] This invention belongs to the field of artificial intelligence and medical speech analysis technology, and relates to an intelligent classification method for voice state based on speech spectrogram features and reinforcement learning optimization mechanism. Background Technology
[0002] Laryngeal tumors are among the most common malignant diseases of the head and neck region, and early diagnosis is crucial for improving cure rates and prognosis. Studies have shown that laryngeal lesions directly alter the vibration state of the vocal cords, resulting in characteristics such as periodic perturbations, frequency fluctuations, and abnormal timbre in the voice signal. These abnormalities often appear in the early stages of the disease; therefore, non-invasive detection using voice signals has potential clinical application value.
[0003] Currently, speech-based pathological detection methods have mainly gone through three stages of development. The earliest methods mainly relied on traditional acoustic parameters, such as fundamental frequency jitter, amplitude jitter, harmonic noise ratio (HNR), and formant frequencies. These features can reflect some physiological abnormalities in vocal cord vibration, but due to the limited dimensionality of the parameters and significant individual differences, their diagnostic accuracy is insufficient, and they are mostly used for simple binary classification tasks of pathology and normality, making it difficult to assist in the determination of tumor staging.
[0004] With advancements in speech signal processing technology, researchers have begun to utilize more complex acoustic features such as Mel-frequency cepstral coefficients (MFCC) and linear predictive cepstral coefficients (LPCC), combining them with traditional machine learning models like support vector machines, K-nearest neighbors, and random forests for laryngeal disease identification. Compared to traditional acoustic parameters, these methods offer improved feature representation and classification performance. However, they rely on manually designed features, have limited robustness, struggle to capture the dynamic changes of speech signals over time, and lack the ability to learn and generalize from large-scale data.
[0005] In recent years, the rise of deep learning has provided new solutions for pathological speech detection. Convolutional Neural Networks (CNNs) can extract local spatial features from spectrograms, and Recurrent Neural Networks (RNNs), especially Long Short-Term Memory (LSTM) networks, can model the long-term dependencies of speech sequences. The CRNN architecture, combining the two, has shown good results in various pathological speech detection tasks. However, these methods still have shortcomings: First, most CNNs only use convolutional kernels of fixed size, making it difficult to capture speech details at different scales simultaneously, and easily missing lesion information; second, existing studies mostly focus on binary classification tasks of "whether or not there is disease," lacking the ability to identify different stages of laryngeal tumors; third, features at different scales are often treated equally in the model, lacking a dynamic weight adjustment mechanism, resulting in insufficient expression of key lesion features, affecting diagnostic accuracy and interpretability.
[0006] A search of existing published patents reveals that related technologies are mostly concentrated on medical image recognition or diagnosis based on traditional acoustic features. For example, patent application CN116564522A mainly utilizes a multivariate linear regression model, vital signs, and dynamic parameters for comprehensive evaluation, but it does not involve artificial intelligence-based staging diagnosis of laryngeal tumors. Therefore, there is still a lack of effective methods for non-invasive intelligent detection of laryngeal tumor staging.
[0007] In summary, existing methods generally suffer from insufficient feature representation capabilities, inadequate integration of spatial and temporal features in model structure, limited diagnostic results to binary classification, and a lack of dynamic attention mechanisms. Therefore, it is necessary to propose an intelligent hierarchical diagnostic method for laryngeal tumors based on speech spectrogram features to improve the accuracy and reliability of early screening and achieve refined auxiliary diagnosis of tumor staging. Summary of the Invention
[0008] To address the aforementioned technical challenges, a method for intelligent voice state classification based on speech spectrogram features and reinforcement learning optimization is proposed. The method takes the spectrogram (preferably Mel spectrogram) obtained by time-frequency transformation of the voice as input, uses multi-scale convolution for spatial feature extraction, introduces an attention mechanism to achieve adaptive weighting, combines a long short-term memory network for temporal modeling, and finally outputs the classification through Softmax to automatically determine whether a person has a laryngeal tumor and the severity of its stage.
[0009] The technical solution of this invention:
[0010] A method for intelligent voice state classification based on speech spectrogram features and reinforcement learning optimization mechanism, comprising the following steps:
[0011] S1, Target Data Acquisition;
[0012] Under uniform acquisition conditions, test voice samples are obtained by collecting test voice signals to form an original voice sample set;
[0013] S11. Determine the data collection conditions: Clarify the data collection environment and technical conditions, including:
[0014] The sampling environment had a background noise sound pressure level below 40 dB(A);
[0015] A recording device with a sampling rate of 16kHz and a quantization precision of 16bit was used.
[0016] The vocal content includes the vowels / a / , / i / , and / u / , as well as the pre-set short text readings, which are standard reading sentences used for collecting the vocal signals of the test subjects;
[0017] Voice status labels are divided into five categories according to a preset grading standard; voice status labels are generated based on pre-established reference labeling rules, or are labeled by no fewer than two labelers with relevant experience in accordance with the principle of consistency.
[0018] S12. Data Acquisition Equipment Setup and System Calibration: The data acquisition device includes a condenser microphone, an audio acquisition terminal, and a computer processing system; the condenser microphone is kept in a stable orientation by a pop filter and a fixed bracket; the distance between the condenser microphone and the subject's lips is maintained at 9-11 cm; the audio acquisition terminal is used to complete analog or digital conversion and synchronously record acquisition time information and equipment operating status information; before acquisition, the acquisition link is calibrated for consistency using an acoustic calibrator or standard white noise signal, and a short trial recording is performed before each test to confirm that the background noise meets the preset threshold; only voice samples that meet the preset threshold are retained;
[0019] S13. Voice Sample Collection and Recording: Collect no less than 10 voice samples from each subject; during the collection process, simultaneously record: voice waveform data, voice status labels, collection time information, and equipment operation status information to form a complete original voice sample set; use a subjective rating scale of 0-10 to evaluate the voice status, with a scoring interval of 1 point, where 0 points indicate no abnormality and 10 points indicate extremely significant abnormality; map the scores to five categories of voice status labels according to preset intervals: [0–2] is the first category of voice status label, (2–4] is the second category of voice status label, (4–6] is the third category of voice status label, (6–8] is the fourth category of voice status label, and (8–10] is the fifth category of voice status label; score each subject's voice sample in a timely manner after collection and record the score and the mapped voice status label.
[0020] S2, Data Processing and Feature Optimization;
[0021] The raw voice sample set obtained in step S1 is subjected to data preprocessing and quality control, acoustic feature extraction, and reinforcement learning optimization mechanism is introduced on this basis to achieve adaptive adjustment of feature subset selection and feature combination strategy, and output an optimized feature dataset for training deep learning classification model.
[0022] S21. Data Preprocessing and Quality Control: The test voice samples are used as input to the acoustic feature extraction module in the deep learning classification model. A unified data normalization process is executed, including: pre-emphasis processing to enhance high-frequency components; frame segmentation based on fixed frame length and frame shift; speech endpoint detection to distinguish between effective speech segments and silence segments; estimation of background noise level based on silence segments and noise reduction processing; amplitude normalization and standardization processing of the test voice signals; consistency verification of multiple test voice samples from the same subject is performed by combining acquisition time information and equipment operating status information; when there are voice status label conflicts or acquisition anomalies, only test voice samples that meet the preset threshold acquisition conditions and have complete time records are retained. The voice status label conflict refers to the same test voice sample corresponding to multiple different categories of voice status labels, and the acquisition anomaly refers to the test voice signal having amplitude saturation, signal loss, or abnormal equipment operating status recording; abnormal voice samples that deviate from the overall distribution by more than a preset threshold are removed through an anomaly detection mechanism based on the statistical distribution of acoustic features, forming a set of effective voice samples that meet the quality requirements.
[0023] S22. Acoustic Feature Extraction: Acoustic features are extracted from the effective voice sample set output in step S21 to form a set of voice acoustic features, including Mel frequency cepstral coefficient features, fundamental frequency correlation features, formant distribution features, energy correlation statistical features, and comprehensive statistical features in the time and frequency domains; the voice acoustic features are used to characterize the differences in voice state in terms of spectral structure, energy distribution, and temporal evolution;
[0024] To ensure consistency across different feature scales, the acoustic features of the voice are uniformly normalized, and a multidimensional feature vector is constructed to represent each test voice sample.
[0025] S23. Feature Selection and Combination Optimization Driven by Reinforcement Learning Mechanism: After acoustic feature extraction, a reinforcement learning optimization mechanism is introduced to adaptively optimize the feature subset selection and feature combination strategy in the voice acoustic feature set. The selection process of some features in the voice acoustic feature set is modeled as a sequential decision process, and a reinforcement learning feature selection strategy is constructed. Each sequential round of decision corresponds to selecting or removing a feature from the voice acoustic feature set. 20% of the test voice samples are randomly divided from the original voice sample set to form a validation set, and the performance index of the deep learning classification model on the validation set is used as the reward signal. The reinforcement learning feature selection strategy is iteratively updated according to the reward signal. After multiple rounds of reinforcement learning feature selection strategy optimization, a feature subset whose contribution to the classification of the test voice state meets the preset threshold is selected, and redundant or low-contribution features are removed to obtain the optimized feature dataset. The reinforcement learning optimization mechanism is executed during the offline training phase.
[0026] S3. Adaptive optimization modeling that combines deep learning classification models with reinforcement learning optimization mechanisms;
[0027] Deep feature learning and classification modeling are performed on the optimized feature dataset output in step S2, and a reinforcement learning optimization mechanism is introduced for adaptive optimization. The structure and training hyperparameters of the deep learning classification model are dynamically adjusted, and the predicted probability distribution of each voice state label and the corresponding classification results are output.
[0028] S31. Construction of Spectral Input for Deep Learning Classification Model: Based on the spectral information in the test voice samples and their optimized feature dataset output in step S2, construct a two-dimensional time-frequency representation for the input of the deep learning classification model; map the preprocessed test voice samples into a two-dimensional Mel spectrogram, and perform size unification, amplitude normalization and tensor quantization on the two-dimensional Mel spectrogram to form a standardized input tensor; wherein, the two-dimensional Mel spectrogram is used to characterize the joint distribution features of the test voice signal in the time dimension and frequency dimension, and serves as the input carrier for subsequent multi-scale convolutional feature extraction and temporal dependency modeling;
[0029] S32. Multi-scale Convolutional Feature Extraction and Fusion: A multi-scale feature extraction structure with multiple parallel convolutional branches is constructed. Multiple parallel convolutional branches are used to extract features from the standardized input tensor. Each parallel convolutional branch includes at least a first, second, and third convolutional branch. The first convolutional branch uses a 3×3 kernel, the second uses a 5×5 kernel, and the third uses a 7×7 kernel. Each convolutional branch is used to extract local spectral texture features, mesoscale spectral variation features, and large-scale spectral structure features from the two-dimensional Mel spectrogram of the test voice signal. After convolution and pooling processing, the output features corresponding to each convolutional branch are obtained. The features output by each convolutional branch are concatenated or combined along the feature channel dimension to achieve multi-scale feature fusion and form a fused feature.
[0030] S33. Temporal Dependency Modeling and Feature Enhancement: After completing the convolutional feature fusion, a bidirectional temporal feature learning structure is introduced to extract the temporal change features between continuous time segments in the fused features. The bidirectional temporal feature learning structure is used to learn the correlation between the test subject's voice signal in different time segments to enhance the ability to express the trend of voice state changes. The deep features after temporal feature learning are mapped to the space of voice state labels through a fully connected layer, and the predicted probability distribution of each voice state label is output.
[0031] S34. Reinforcement learning optimization mechanism-driven adaptive optimization of structure and parameters:
[0032] In the training phase of the deep learning classification model, a reinforcement learning optimization mechanism is introduced to adaptively adjust the structural configuration and training hyperparameters of the deep learning classification model. Specifically, this includes defining the structural configuration and training hyperparameters of the deep learning classification model as decision variables, where the structural configuration includes convolutional kernel size, the number of convolutional branches, and the network layer depth, and the training hyperparameters include learning rate and regularization strength. The classification accuracy on the validation set, the stability index of the deep learning classification model output, and the training convergence efficiency are used as reward signals. A policy update mechanism iteratively optimizes the selection of convolutional kernel size, the configuration of the number of convolutional branches, the network layer depth, the learning rate setting, and the regularization strength. The reinforcement learning optimization mechanism dynamically adjusts the structural configuration of the deep learning classification model based on training feedback, converging after multiple iterations to a combination of deep learning classification models that meets preset performance indicators. The iterative optimization process is executed during the offline training phase, ultimately outputting a deep learning classification model with a determined structure and stable performance for practical deployment.
[0033] S4. Classification result output, confidence calibration, interpretability generation and adaptive update;
[0034] The predicted voice state labels and predicted probability distribution output in step S3 are calibrated for confidence and uncertainty, and an interpretable result is generated. When the preset safety conditions are met, the optimization strategy of the deep learning classification model is updated in a controlled adaptive manner based on the reinforcement learning optimization mechanism, and the final voice state judgment result, calibrated confidence and acoustic region of interest are output.
[0035] S41. Classification Result Output and Recording: Based on the predicted probability distribution of each voice state label output in step S3, determine the predicted voice state category according to the maximum probability principle, and output the corresponding confidence level simultaneously; associate and record the predicted voice state category, original probability distribution, collection time information and device status information of each test voice sample for subsequent quality traceability and deep learning classification model maintenance.
[0036] S42. Confidence calibration and uncertainty determination: Perform confidence calibration on the predicted probability distribution output by the deep learning classification model; at the same time, set uncertainty determination rules: when the predicted confidence of the test voice sample is lower than the preset threshold or the predicted probability distribution shows significant confusion features, mark the test voice sample as an uncertain sample and output a prompt message so that it can enter the manual review or secondary collection process.
[0037] S43. Interpretability Enhancement and Region of Interest Output: Correlation analysis is performed on the two-dimensional Mel spectrogram input and the multi-scale convolutional feature response and temporal feature response in the deep learning classification model to generate interpretable results, which are used to indicate the acoustic regions of interest that the deep learning classification model focuses on when making classification decisions. Among them, the multi-scale convolutional feature response is used to characterize the spectral texture features, spectral variation features, and spectral structure features of the test voice signal under different receptive fields of convolutional kernels, and the temporal feature response is used to characterize the changing trends and correlation features of the test voice signal between consecutive time segments. The interpretable results include the time segment range, frequency band range, and acoustic regions of interest presented in the form of heatmaps, which are used to help determine the main sources of difference between different test voice states.
[0038] S44. Controlled Adaptive Update and Closed-Loop Optimization of Reinforcement Learning Optimization Mechanism: During the deployment phase of the deep learning classification model, output stability indicators are continuously monitored. For the test voice samples, the prediction consistency of consecutive samples, the proportion of low-confidence samples, and the proportion of uncertain samples are continuously monitored. When the proportion of uncertain samples in M consecutive test voice samples is higher than a preset threshold, or when K consecutive low-confidence samples appear, a deep learning classification model update evaluation process is triggered. In the update evaluation process, the feature selection strategy and preset threshold are adjusted first. If the evaluation shows improved performance, the update constraints are gradually relaxed in a graded manner by increasing the range of optional feature subsets or increasing the number of structure search steps. If the evaluation shows no improvement in performance or an increase in the proportion of uncertain samples, the update constraints are tightened and the model is rolled back to the most recently validated stable version. The relaxation and tightening rules are set as follows: when the proportion of uncertain samples decreases in two consecutive evaluations and the classification accuracy is not lower than the preset baseline, the update constraints are relaxed. When the classification accuracy is lower than the preset baseline or the proportion of uncertain samples increases by more than a preset amount in any evaluation, the constraints are immediately tightened and the model is rolled back.
[0039] The beneficial effects of this invention are:
[0040] First, this invention introduces a multi-size convolutional kernel structure in the feature extraction stage, which can simultaneously capture local minute perturbations and cross-frequency band global patterns, thereby improving the sensitivity and expressiveness of speech features from lesions of different granularities. Compared with traditional single-scale convolutional methods, this design is more suitable for characterizing multi-level acoustic abnormalities caused by laryngeal tumors. Second, this invention uses spectrograms as input features, preferably Mel spectrograms to simulate the characteristics of human hearing, and uniformly adjusts them to image sizes adapted for deep learning, improving the stability and convergence efficiency of model training while ensuring the integrity of acoustic information. In addition to Mel spectrograms, this invention can also be extended to constant Q-transform spectrograms and cochlear maps, achieving diversity and scalability of feature representation. Third, this invention integrates recurrent neural networks and attention mechanisms into the model structure, dynamically weights the multi-scale features output by the convolutional layers, and models the temporal dependence of speech signals, thereby significantly improving the ability to recognize complex acoustic evolution patterns and enhancing the sensitivity and interpretability of the model in early lesion identification. Furthermore, this invention is geared towards clinical applications, featuring a five-category output format covering laryngeal tumors and stage I-IV diagnoses, providing graded early warning and auxiliary judgment. Compared to most existing studies that only achieve binary classification, this invention achieves higher-resolution clinical decision support, meeting the practical needs of multi-level early warning.
[0041] Finally, the method of the present invention has good scalability and versatility. In the future, it can not only be applied to the screening and follow-up monitoring of laryngeal tumors, but also extended to the identification of other voice diseases such as vocal cord nodules and vocal cord polyps. Furthermore, it can be combined with more modalities (such as medical images and medical history data) to further improve the comprehensiveness of diagnosis. Attached Figure Description
[0042] Figure 1 is a flowchart of the overall process of a voice state intelligent classification method based on speech spectrogram features and reinforcement learning optimization mechanism according to the present invention.
[0043] Figure 2 is a schematic diagram of the adaptive optimization modeling of an intelligent voice state classification method based on speech spectrogram features and reinforcement learning optimization mechanism according to the present invention.
[0044] Figure 3 is a schematic diagram of the classification result output, confidence calibration, interpretability generation, and controlled adaptive update of an intelligent voice state classification method based on speech spectrogram features and reinforcement learning optimization mechanism according to the present invention. Detailed Implementation
[0045] The specific embodiments of the present invention are further described below with reference to the accompanying drawings and technical solutions:
[0046] As shown in Figure 1, this embodiment provides an intelligent voice state classification method based on speech spectrogram features and reinforcement learning optimization mechanism, including the following steps:
[0047] 1) Target Data Acquisition: Under uniform acquisition conditions, test voice samples are obtained by collecting the test subject's voice signals to form a raw voice sample set. The acquisition environment is preferably a quiet environment with a background noise sound pressure level below 40 dB(A), using a recording device with a sampling rate of 16 kHz and a quantization accuracy of 16 bits. The vocal content includes vowels / a / , / i / , / u / , and pre-set short sentence text readings. The acquisition device includes a condenser measurement microphone, an audio acquisition terminal, and a computer processing system. The condenser measurement microphone is kept stable by a pop filter and a fixed bracket, and the distance between it and the subject's lips is maintained at 9-11 cm. Before acquisition, the acquisition link is calibrated for consistency using an acoustic calibrator or a standard white noise signal, and a short trial recording is performed before each test to confirm that the background noise meets the preset threshold. For each subject, no less than 10 test voice samples are collected, and voice waveform data, voice status labels, acquisition time information, and device status information are recorded simultaneously to form a complete raw voice sample set. Voice status labels are divided into five categories according to a preset grading standard. They can be generated based on pre-established reference labeling rules or labeled by no fewer than two labelers with relevant experience in accordance with the principle of consistency.
[0048] In one embodiment, a subjective rating scale of 0-10 points is used to evaluate voice state, and the rating results are mapped to five voice state labels according to preset intervals. To ensure acquisition quality, the signal-to-noise ratio is preferably used as the evaluation index, and the effective speech segment power is denoted as... The power of the background noise segment is Then the signal-to-noise ratio It can be represented as:
[0049]
[0050] Only the test voice samples that meet the preset threshold are retained for subsequent steps.
[0051] 2) Data Processing and Feature Optimization: As shown in Figure 1, after completing the target data acquisition, the original voice sample set undergoes data preprocessing and quality control, acoustic feature extraction, and a reinforcement learning optimization mechanism is introduced to achieve adaptive optimization of feature subset selection and feature combination strategies, resulting in an optimized feature dataset. Specifically, the test voice samples are first processed with pre-emphasis, framing, speech endpoint detection, noise reduction, amplitude normalization, and standardization. The pre-emphasized signal... It can be represented as:
[0052]
[0053] in, The original subject's voice signal, This is the pre-emphasis coefficient. After framing, the first... Short-time energy of frame signal It can be represented as:
[0054]
[0055] Its zero-crossing rate It can be represented as:
[0056]
[0057] in, The number of sampling points per frame. For the first The first frame signal The amplitude of each sampling point For the first The first frame signal The amplitude of each sampling point This is a sign function, taking the value 1 when the independent variable is greater than or equal to 0, and -1 when the independent variable is less than 0. Combining short-time energy and zero-crossing rate, effective speech segments and silent segments can be distinguished.
[0058] Combining short-time energy and zero-crossing rate can distinguish between valid speech segments and silent segments. Subsequently, noise reduction and standardization are performed on the samples, and the standardized features... It can be represented as:
[0059]
[0060] in, These are the original eigenvalues. The mean, The standard deviation is denoted as .
[0061] After preprocessing, acoustic features are extracted from the effective voice sample set, forming a set of voice acoustic features including Mel frequency cepstral coefficient features, fundamental frequency correlation features, formant distribution features, energy correlation statistical features, and comprehensive time-domain and frequency-domain statistical features. A multi-dimensional feature vector is then constructed to represent each test voice sample. Based on this, a reinforcement learning optimization mechanism is introduced to adaptively optimize the feature subset selection and feature combination strategy. The feature selection process is modeled as a sequential decision-making process, denoted as state . The action is The reward is Its objective function is:
[0062]
[0063] in, Feature selection strategy, Indicating in strategy The expected return is as follows For the decision-making step, For the sequence decision length, As a discount factor, For the first The immediate reward corresponding to each step of the decision. Indicates the first The state of the step, Indicates the first The steps taken are as follows. The reinforcement learning optimization mechanism iteratively updates the feature selection strategy based on the performance feedback on the validation set, selects a subset of features whose contribution to the classification of the subject's voice state meets the preset threshold, and removes redundant or low-contribution features to obtain the optimized feature dataset.
[0064] 3) Adaptive optimization modeling combining deep learning classification models and reinforcement learning optimization mechanisms: As shown in Figure 2, the preprocessed test voice samples are mapped to two-dimensional Mel spectrograms. Size unification, amplitude normalization, and tensor quantization are then performed on the two-dimensional Mel spectrograms to form standardized input tensors. Based on this, a multi-scale feature extraction structure containing first, second, and third convolutional branches is used to extract spectral features under different receptive field ranges. After feature fusion and temporal dependency modeling, the classification results are output. Simultaneously, a reinforcement learning optimization mechanism is constructed based on training feedback to perform closed-loop optimization of the deep learning classification model's structural configuration and training hyperparameters. Two-dimensional Mel spectrogram features. It can be represented as:
[0065]
[0066] in, Indicates at time First Eigenvalues of the two-dimensional Mel spectrum corresponding to each Mel filter. For the subject's voice signal at time The short-time Fourier transform results, For frequency index, This represents the total number of frequency sampling points. For the first A Mel filter at frequency index The response at the location, This refers to the Mel filter number. To prevent extremely small constants from having zero values in logarithmic operations.
[0067] Then, a multi-scale feature extraction structure containing multiple parallel convolutional branches is constructed. These multiple parallel convolutional branches include at least a first convolutional branch, a second convolutional branch, and a third convolutional branch, which are respectively employed... , and Convolutional kernels. The first convolutional branch extracts local spectral texture features, the second convolutional branch extracts mesoscale spectral variation features, and the third convolutional branch extracts large-scale spectral structure features. For the third... Each convolutional branch has a convolutional output. It can be represented as:
[0068]
[0069] in, To standardize the input tensor, For convolution kernel parameters, For bias terms, This is a non-linear activation function. The outputs of each convolutional branch are pooled and then concatenated or combined along the feature channel dimension to form a fused feature. :
[0070]
[0071] in, As a feature of fusion, , , These represent the output features of the first, second, and third convolutional branches after convolution and pooling, respectively. This represents a splicing operation along the feature channel dimension.
[0072] After completing the convolutional feature fusion, a bidirectional temporal feature learning structure is introduced to extract the temporal variation features between consecutive time segments in the fused features, thereby enhancing the ability to express the changing trends of voice state. Let the first... The input features at each time step are Then the bidirectional time series feature representation can be written as:
[0073]
[0074] in, Indicates the first Input features at each time step, The forward temporal feature learning structure represents the first... The hidden state at each time step The backward temporal feature learning structure represents the first... The hidden state at each time step This represents the bidirectional temporal feature representation obtained by concatenating the forward hidden state and the backward hidden state. This represents a forward temporal feature learning unit. Represents the backward temporal feature learning unit. This indicates a splicing operation.
[0075] The deep features learned from the temporal features are mapped to the voice state label space through a fully connected layer, and the predicted probability distribution of each voice state label is output through Softmax:
[0076]
[0077] in, Indicates the first The predicted probability of voice-like state labels. Indicates the first The output score corresponding to the voice state label. Indicates the first The output score corresponding to the voice state label. For the summation index, The number of voice status label categories. Represented by natural constant An exponential function with base 0.
[0078] In the training phase of the deep learning classification model, a reinforcement learning optimization mechanism is introduced to adaptively adjust the model's structural configuration and training hyperparameters. The structural configuration includes kernel size, number of convolutional branches, and network depth; the training hyperparameters include learning rate and regularization strength. The classification accuracy on the validation set, the stability index of the deep learning classification model output, and training convergence efficiency are used as reward signals. A policy update mechanism iteratively optimizes the selection of kernel size, the configuration of the number of convolutional branches, the network depth, the learning rate setting, and the regularization strength. Reward function. It can be represented as:
[0079]
[0080] in, Indicates classification accuracy. This indicates the output stability index. This indicates the training convergence efficiency. , and These represent the weight coefficients corresponding to classification accuracy, output stability index, and training convergence efficiency, respectively.
[0081] 4) Classification Result Output, Confidence Calibration, Interpretability Generation, and Adaptive Update: As shown in Figure 3, the predicted probability distribution and classification results of each voice state label output by the deep learning classification model are post-processed. This involves sequentially recording results, calibrating confidence, and determining uncertainty, generating interpretability results, and monitoring output stability. When the update evaluation conditions are met, the controlled adaptive update and closed-loop optimization of the reinforcement learning optimization mechanism are triggered. Specifically, the predicted voice state category is determined according to the maximum probability principle.
[0082]
[0083] The predicted voice state category, original probability distribution, acquisition time information, and device status information are then linked and recorded. Subsequently, confidence calibration processing is performed on the predicted probability distribution, preferably using a temperature scaling method. It can be represented as:
[0084]
[0085] in, The original output score. The temperature parameter is used. When the prediction confidence is lower than a preset threshold, or when the prediction probability distribution exhibits significant confusion characteristics, the test subject's voice sample is marked as an uncertain sample.
[0086] Simultaneously, correlation analysis is performed on the Mel spectrogram input and the multi-scale convolutional feature responses and time-series feature responses in the deep learning classification model to generate interpretable results. These interpretable results include time segment ranges, frequency band ranges, and acoustic regions of interest presented in heatmap form, used to help determine the main sources of difference in the voice states of different test subjects. In one embodiment, the interpretable response map... It can be represented as:
[0087]
[0088] in, This represents an interpretability response diagram. Indicates the first Feature response maps corresponding to each convolutional branch Represents the time-series characteristic response diagram. For convolution branch index, The total number of convolution branches. and These are the corresponding weighting coefficients.
[0089] During the deployment phase of the deep learning classification model, output stability metrics are continuously monitored, and statistics are performed on the prediction consistency of continuous samples, the proportion of low-confidence samples, and the proportion of uncertain samples. If the proportion of uncertain samples in the tested voice samples is higher than a preset threshold, or if uncertain samples appear consecutively... When a low-confidence sample is encountered, a deep learning classification model update and evaluation process is triggered. The proportion of uncertain samples in continuous samples. It can be represented as:
[0090]
[0091] in, For an uncertain sample size, To continuously monitor the total number of samples, the feature selection strategy and preset threshold are first adjusted during the update evaluation process. If the evaluation shows improved performance, the update constraints are gradually relaxed in a tiered manner, such as increasing the range of selectable feature subsets or increasing the number of structure search steps. If the evaluation shows no improvement in performance or an increase in the proportion of uncertain samples, the update constraints are tightened and the system reverts to the most recently validated stable version, thereby achieving controlled adaptive updates and closed-loop optimization of the reinforcement learning optimization mechanism.
[0092] The present invention has at least the following beneficial effects:
[0093] 1. The multi-scale feature extraction structure, consisting of the first, second, and third convolutional branches, can extract local spectral texture features, mesoscale spectral variation features, and large-scale spectral structure features, thereby improving the ability to represent differences in different voice states.
[0094] 2. By introducing a reinforcement learning optimization mechanism, the feature selection and combination strategy, the structural configuration of the deep learning classification model, and the training hyperparameters are adaptively optimized, which helps to improve the accuracy, stability, and training efficiency of intelligent voice state classification.
[0095] 3. By modeling the variation features between consecutive time segments in the test voice samples through a bidirectional temporal feature learning structure, the ability to characterize the temporal evolution of voice state can be enhanced, thereby improving the stability of classification results.
[0096] 4. Through confidence calibration, uncertainty determination, interpretability result generation, and controlled adaptive update and closed-loop optimization, the final voice state determination result, calibrated confidence, and acoustic region of interest can be output simultaneously, thereby improving the interpretability, reliability, and practical application value of the system.
[0097] This invention is not limited to the specific embodiments described above. Without departing from the spirit and scope of this invention, those skilled in the art can make various equivalent changes or substitutions, all of which should fall within the scope of protection of this invention.
Claims
1. A method for intelligent classification of voice states based on speech spectrogram features and reinforcement learning optimization mechanism, characterized in that, The steps are as follows: S1, Target data acquisition; Under uniform acquisition conditions, acquire test voice samples by collecting test voice signals to form an original voice sample set; S2, Data processing and feature optimization; The raw voice sample set obtained in step S1 is preprocessed and quality controlled, and acoustic features are extracted. Based on this, a reinforcement learning optimization mechanism is introduced to achieve adaptive adjustment of feature subset selection and feature combination strategy, and output an optimized feature dataset for training the deep learning classification model. S3, Adaptive optimization modeling combining deep learning classification model and reinforcement learning optimization mechanism: Deep feature learning and classification modeling are performed on the optimized feature dataset output in step S2, and a reinforcement learning optimization mechanism is introduced for adaptive optimization. The structure and training hyperparameters of the deep learning classification model are dynamically adjusted, and the predicted probability distribution of each voice state label and the corresponding classification results are output. S4. Classification result output, confidence calibration, interpretability generation and adaptive update; perform confidence calibration and uncertainty determination on the predicted voice state labels and predicted probability distribution output in step S3, and generate interpretable results; when the preset safety conditions are met, perform controlled adaptive update on the optimization strategy of the deep learning classification model based on the reinforcement learning optimization mechanism, and output the final voice state determination result, calibrated confidence and acoustic region of interest.
2. The intelligent voice state classification method based on speech spectrogram features and reinforcement learning optimization mechanism according to claim 1, characterized in that, The specific implementation process of step S1 is as follows: S11. Determine the acquisition conditions: Clarify the acquisition environment and technical conditions, including: the acquisition environment has a background noise sound pressure level of less than 40 dB(A); a recording device with a sampling rate of 16 kHz and a quantization accuracy of 16 bits is used; the vocal content includes the vowels / a / , / i / , / u / and the pre-set short sentence text readings, the short sentence text being a standard reading sentence used for the acquisition of the subject's voice signal; the voice status labels are divided into five categories according to the pre-set grading standards; the voice status labels are generated according to the pre-established reference labeling rules, or are labeled by no less than two labelers with relevant experience according to the consistency principle; S12. Set up the acquisition equipment and calibrate the system: the acquisition device includes a condenser measurement microphone, an audio acquisition terminal, and a computer processing system; the condenser measurement microphone is kept stable in orientation by a pop filter and a fixed bracket; the distance between the condenser measurement microphone and the subject's lips is kept at 9-11 cm; the audio acquisition terminal is used to complete analog or digital conversion and synchronously record the acquisition time information and equipment operating status information. Information; Before collection, use an acoustic calibrator or standard white noise signal to perform consistency calibration on the acquisition link, and perform a short trial recording before each test to confirm that the background noise meets the preset threshold; only retain voice samples that meet the preset threshold; S13, Voice Sample Collection and Recording: Collect no less than 10 voice samples for each subject; during the collection process, synchronously record: voice waveform data, voice status labels, collection time information, and equipment operating status information to form a complete original voice sample set; use a subjective rating of 0-10 points. The scale evaluates vocal condition with a 1-point interval, where 0 points indicate no abnormality and 10 points indicate extremely significant abnormality. The scores are mapped to five vocal condition labels according to preset intervals: [0–2] for the first vocal condition label, (2–4] for the second vocal condition label, (4–6] for the third vocal condition label, (6–8] for the fourth vocal condition label, and (8–10] for the fifth vocal condition label. Each subject's vocal sample is scored and recorded immediately after collection, along with the mapped vocal condition label.
3. The intelligent voice state classification method based on speech spectrogram features and reinforcement learning optimization mechanism according to claim 2, characterized in that, The specific implementation process of step S2 is as follows: S21, Data Preprocessing and Quality Control: The test voice samples are used as input to the acoustic feature extraction module in the deep learning classification model, and a unified data normalization process is executed, including: pre-emphasis processing to enhance high-frequency components; frame segmentation processing based on fixed frame length and frame shift; speech endpoint detection to distinguish between effective speech segments and silence segments; estimation of background noise level based on silence segments and noise reduction processing; amplitude normalization and standardization processing of the test voice signals; and consistency verification of multiple test voice samples from the same subject by combining the acquisition time information and equipment operating status information; when there are voice status label conflicts or When an acquisition anomaly occurs, only test voice samples that meet the acquisition conditions of a preset threshold and have complete time recordings are retained. Here, voice state label conflict refers to the same test voice sample corresponding to multiple different categories of voice state labels. Acquisition anomaly refers to situations where the test voice signal exhibits amplitude saturation, signal loss, or abnormal equipment operation status recording. An anomaly detection mechanism based on acoustic feature statistical distribution removes abnormal voice samples that deviate from the overall distribution by more than a preset threshold, forming a valid voice sample set that meets quality requirements. S22, Acoustic Feature Extraction: Acoustic feature extraction is performed on the valid voice sample set output in step S21, forming features including Mel-frequency cepstral coefficients and fundamental frequency phase... A set of voice acoustic features, including key features, formant distribution features, energy correlation statistical features, and comprehensive statistical features in the time and frequency domains; these voice acoustic features are used to characterize the differences in voice state in terms of spectral structure, energy distribution, and temporal evolution; to ensure consistency across different feature scales, the voice acoustic features are uniformly normalized, and a multi-dimensional feature vector is constructed to represent each test voice sample; S23, Feature selection and combination optimization driven by reinforcement learning optimization mechanism: after completing the acoustic feature extraction, a reinforcement learning optimization mechanism is introduced to adaptively optimize the feature subset selection and feature combination strategy in the voice acoustic feature set; some features in the voice acoustic feature set are optimized... The feature selection process is modeled as a sequential decision-making process, and a reinforcement learning feature selection strategy is constructed. Each sequential round of decision corresponds to selecting or removing a feature from the set of voice acoustic features. 20% of the test voice samples are randomly divided from the original voice sample set to form a validation set, and the performance index of the deep learning classification model on the validation set is used as the reward signal. The reinforcement learning feature selection strategy is iteratively updated according to the reward signal. After multiple rounds of optimization of the reinforcement learning feature selection strategy, a subset of features whose contribution to the classification of the test voice state meets the preset threshold is selected, and redundant or low-contribution features are removed to obtain the optimized feature dataset. The reinforcement learning optimization mechanism is executed during the offline training phase.
4. The intelligent voice state classification method based on speech spectrogram features and reinforcement learning optimization mechanism according to claim 3, characterized in that, The specific implementation process of step S3 is as follows: S31, Spectral input construction for deep learning classification model: Based on the spectral information in the test voice samples and their optimized feature dataset output in step S2, construct a two-dimensional time-frequency representation for the input of the deep learning classification model; map the preprocessed test voice samples into a two-dimensional Mel spectrogram, and perform size unification, amplitude normalization, and tensor quantization on the two-dimensional Mel spectrogram to form a standardized input tensor; wherein, the two-dimensional Mel spectrogram is used to characterize the joint distribution features of the test voice signal in the time and frequency dimensions, serving as the input carrier for subsequent multi-scale convolutional feature extraction and temporal dependency modeling; S32, Multi-scale convolutional feature extraction and fusion: construct a multi-scale convolutional feature extraction and fusion model containing multiple parallel convolutional branches. The scale feature extraction structure employs multiple parallel convolutional branches to extract features from the standardized input tensor. Each parallel convolutional branch includes at least a first, second, and third convolutional branch. The first branch uses a 3×3 kernel, the second a 5×5 kernel, and the third a 7×7 kernel. Each convolutional branch is used to extract local spectral texture features, mesoscale spectral variation features, and large-scale spectral structure features from the two-dimensional Mel spectrogram of the test voice signal, respectively. After convolution and pooling processing, the output features corresponding to each convolutional branch are obtained. The features output from each convolutional branch are concatenated or combined along the feature channel dimension to achieve multi-scale feature fusion and form a fused feature. (S3) 3. Temporal Dependency Modeling and Feature Enhancement: After completing convolutional feature fusion, a bidirectional temporal feature learning structure is introduced to extract temporal variation features between consecutive time segments in the fused features. This structure is used to learn the correlation between different time segments of the subject's voice signal, enhancing its ability to express the trend of voice state changes. The deep features learned through temporal feature learning are mapped to the voice state label space through a fully connected layer, outputting the predicted probability distribution of each voice state label. S34. Reinforcement Learning Optimization Mechanism-Driven Adaptive Optimization of Structure and Parameters: During the training phase of the deep learning classification model, a reinforcement learning optimization mechanism is introduced to adaptively adjust the structure configuration and training hyperparameters of the deep learning classification model. The process involves: defining the structural configuration and training hyperparameter settings of the deep learning classification model as decision variables, where the structural configuration includes kernel size, number of convolutional branches, and network layer depth, and the training hyperparameters include learning rate and regularization strength; using the classification accuracy of the deep learning classification model on the validation set, the output stability index of the deep learning classification model, and the training convergence efficiency as reward signals; iteratively optimizing the selection of kernel size, configuration of the number of convolutional branches, network layer depth, learning rate setting, and regularization strength through a policy update mechanism; and using a reinforcement learning optimization mechanism to dynamically adjust the structural configuration of the deep learning classification model based on training feedback, converging to a combination of deep learning classification models that meets the preset performance indicators after multiple iterations.The iterative optimization process is performed during the offline training phase, ultimately outputting a deep learning classification model with a defined structure and stable performance for practical deployment.
5. The intelligent voice state classification method based on speech spectrogram features and reinforcement learning optimization mechanism according to claim 4, characterized in that, The specific implementation process of step S4 is as follows: S41, classification result output and recording: based on the predicted probability distribution of each voice state label output in step S3, determine the category of the predicted voice state according to the principle of maximum probability, and output the corresponding confidence level simultaneously. The predicted voice state category, original probability distribution, collection time information, and device status information of each test voice sample are associated and recorded for subsequent quality traceability and deep learning classification model maintenance; S42, Confidence calibration and uncertainty judgment: Confidence calibration processing is performed on the predicted probability distribution output by the deep learning classification model; At the same time, uncertainty judgment rules are set: When the predicted confidence of the test voice sample is lower than the preset threshold or the predicted probability distribution shows significant confusion characteristics, the test voice sample is marked as an uncertain sample and a prompt message is output so that it can enter the manual review or secondary collection process; S43. Interpretability Enhancement and Region of Interest Output: Correlation analysis is performed on the 2D Mel spectrogram input and the multi-scale convolutional feature responses and temporal feature responses in the deep learning classification model to generate interpretable results. These results indicate the acoustic regions of interest that the deep learning classification model should focus on when making classification decisions. The multi-scale convolutional feature responses characterize the spectral texture, spectral variation, and spectral structure features of the tested voice signal under different convolutional kernel receptive fields. The temporal feature responses characterize the changing trends and correlation features of the tested voice signal between consecutive time segments. The interpretable results include the time segment range, frequency band range, and acoustic regions of interest presented in heatmap form, used to help determine the main sources of difference between different tested voice states. S44. Controlled Adaptive Update and Closed-Loop Optimization of Reinforcement Learning Optimization Mechanism: Output stability indicators are continuously monitored during the deployment phase of the deep learning classification model. For the tested voice... The system continuously monitors the prediction consistency, low-confidence sample ratio, and uncertain sample ratio of consecutive samples. When the uncertain sample ratio in M consecutive test voice samples exceeds a preset threshold, or K consecutive low-confidence samples appear, a deep learning classification model update evaluation process is triggered. In the update evaluation process, the feature selection strategy and preset threshold are adjusted first. If the evaluation shows improved performance, the update constraints are gradually relaxed in a tiered manner by increasing the range of selectable feature subsets or increasing the number of structure search steps. If the evaluation shows no improvement in performance or an increase in the uncertain sample ratio, the update constraints are tightened and the system reverts to the most recently validated stable version. The relaxation and tightening rules are set as follows: if the uncertain sample ratio decreases in two consecutive evaluations and the classification accuracy is not lower than the preset baseline, the update constraints are relaxed; if the classification accuracy is lower than the preset baseline or the uncertain sample ratio increases by more than a preset amount in any evaluation, the constraints are immediately tightened and the system reverts.
Citation Information
Patent Citations
Intelligent evaluation method for laryngocarcinoma screening
CN116564522A
Detection method and system for pathological voice
CN103730130A
Passive underwater acoustic signal identification method and system based on MFCC characteristics
CN117009781A
Intention recognition method for voice question answering based on large-model multi-agent
CN119831043A
Method, system, apparatus, medium and program product for vocal cord disease classification based on ResNet model of MFCC
CN121306201A