Voice training scheme generation method and system

Through data cleaning and feature extraction, combined with the improved S4 neural network model and multi-objective loss function, a personalized voice training solution is dynamically generated, which solves the shortcomings of the existing voice training solution, improves the accuracy and adaptability of voice training, and is suitable for voice health care and voice training in professional populations.

CN120472933APending Publication Date: 2025-08-12BEIJING FENGSHANGTIANCHENG CULTURAL DEV CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510561093.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing voice training schemes have shortcomings in data cleaning, feature extraction, training optimization, etc., resulting in unsatisfactory voice training results and the inability to personalize customization and dynamic adjustment, which affects the improvement of voice quality and the recovery of voice function.

Method used

Voice data and physiological characteristic data were collected, data cleaning was performed through mutual power spectral density and Shapiro-Wilk test, time and frequency domain features were extracted, and the improved S4 neural network model and multi-objective loss function were used for training, and a personalized voice training scheme was dynamically generated.

Benefits of technology

It improves the accuracy and robustness of the voice training program, enhances the accuracy of voice evaluation and the classification stability of the model, and is suitable for voice health and rehabilitation in the general and professional populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472933A_ABST
    Figure CN120472933A_ABST
Patent Text Reader

Abstract

The invention provides a voice training scheme generation method and system. The method is applied to the technical field of voice training, and comprises the following steps: acquiring voice data, voice signals and physiological feature data related to voice generation, and performing preprocessing; extracting time domain and frequency domain features from the preprocessed data, wherein the time domain features comprise fundamental frequency, peak amplitude, signal energy and the like; the frequency domain characteristics comprise a harmonic noise ratio, a fast Fourier transform median, energy distribution of a voice signal in each frequency band, a frequency spectrum centroid, a width of the frequency spectrum centroid and the like; screening high-discrimination features from the initially extracted time domain features and frequency domain features through backward / forward sequence feature selection to construct an optimized feature set, and inputting the optimized feature set into an improved S4 neural network model for training; and analyzing the voice type based on the trained model, dynamically generating a personalized voice training scheme, and performing dynamic adjustment. According to the invention, the accuracy, adaptability and robustness of the voice training scheme are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice training, and in particular to a method and system for generating a voice training program. Background Art

[0002] With the continuous development of speech signal processing technology, voice training has been widely used in fields such as speech rehabilitation, speech correction, and speech enhancement. For example, in the field of voice medicine, for patients with voice disorders, scientific and reasonable voice training can improve the vibration pattern of the vocal cords and enhance voice quality. In application scenarios such as singing training and speech skills training, refined voice assessment and training programs can effectively improve voice control ability. However, existing voice training programs still face many challenges in data cleaning, feature extraction, training optimization, etc.:

[0003] Traditional methods for processing and analyzing voice data have numerous limitations. During voice data collection, the acquisition equipment is susceptible to factors such as environmental noise and inherent device errors. This results in the collected data containing significant amounts of noise, weak signals, and irrelevant leakage signals, severely impacting data quality and interfering with subsequent analysis and model training. Previous methods for extracting data features are insufficiently comprehensive and precise. Time-domain feature extraction struggles to fully capture the complex temporal characteristics of speech signals, while frequency-domain feature extraction also suffers from limitations in analyzing frequency distribution and energy concentration. For example, traditional methods are unable to accurately capture and analyze key information such as the ratio of periodic and aperiodic components in a signal, and the center and spread of the frequency distribution. During model construction and training, traditional models have limited ability to learn long-range dependencies in voice data, making them ineffective in processing complex time series data. Furthermore, traditional loss function designs are relatively simplistic and fail to address multiple optimization objectives, resulting in poor model training results. The accuracy of voice defect detection and classification stability require improvement. Furthermore, existing voice assessment methods lack scientific and comprehensive evaluation metrics, making it difficult to accurately identify various voice abnormalities. In terms of training program generation, most of them are general training programs that cannot be personalized and dynamically adjusted according to each user's specific voice defects and training progress, resulting in unsatisfactory training results and unable to effectively help users improve voice quality and restore normal vocal function. Summary of the Invention

[0004] The present invention provides a method and system for generating a voice training program, aiming to address the deficiencies of existing voice training programs in terms of data cleaning, feature extraction, and training optimization, so as to improve the accuracy, adaptability, and robustness of the voice training program.

[0005] To achieve the above objectives, the following technical solutions are adopted:

[0006] According to a first aspect of the present invention, there is provided a method for generating a voice training program, comprising the following steps:

[0007] Collecting voice data, voice signals, and physiological characteristic data related to voice production and preprocessing them; wherein the physiological characteristic data related to voice production includes static and dynamic information of breathing, body posture information, and muscle state information;

[0008] Extract time domain and frequency domain features from the preprocessed data: the time domain features include any two or more of fundamental frequency, peak amplitude, signal energy, waveform features including skewness and autocorrelation measurements, nonlinear chaotic features, and entropy features; the frequency domain features include any two or more of harmonic noise ratio, fast Fourier transform median, energy distribution of speech signals in each frequency band, spectral centroid and its width, Mel spectrum, short-time Fourier transform phase estimation optimized by the improved Griffin-Lim algorithm, modulation spectrum, linear predictive coding coefficients, and Cepstral coefficients;

[0009] An optimized feature set is constructed by screening high-discriminative features from the initially extracted time-domain and frequency-domain features through backward / forward sequential feature selection. This feature set is then fed into an improved S4 neural network model for training. The S4 neural network model uses a diagonalized state transfer matrix and implicit Z-transform discretization to enhance long-term time series modeling capabilities.

[0010] Construct a weighted loss function that includes sample similarity contrast loss, asynchronous attention loss, and speech representation contrast loss to optimize model classification stability;

[0011] Analyze voice type based on the trained model, dynamically generate personalized voice training plans and make dynamic adjustments.

[0012] According to a second aspect of the present invention, a voice training program generation system is further provided, for implementing the voice training program generation method according to the first aspect, comprising:

[0013] A data acquisition and preprocessing module is used to configure multiple sensors to collect voice signals and physiological characteristic data related to voice production and perform data preprocessing;

[0014] The time-frequency feature extraction module is used to extract time domain and frequency domain features from the preprocessed data;

[0015] A model building and training module is used to integrate the improved S4 neural network model and the multi-objective loss function, optimize the feature set through backward / forward sequential feature selection, and then input the feature set into the improved S4 neural network model for training;

[0016] The voice assessment module includes a square envelope spectrum analyzer and a personalized program generator, which is used to analyze the voice type based on the trained model, dynamically generate personalized training programs and make dynamic adjustments.

[0017] Compared with the prior art, the present invention achieves the following beneficial effects:

[0018] 1. This invention cleans collected data using cross-power spectral density (CPSD). By calculating the cross-spectrum of data collected by different sensors, low-quality signals are filtered out, improving data purity. The frequency domain characteristics of acoustic signals are acquired through fast Fourier transform (FFT), and the CPD value is calculated to remove noise and irrelevant signals, improving data quality.

[0019] 2. The present invention performs a Shapiro-Wilk test on the collected data to determine the data distribution. The Shapiro-Wilk test is used to assess whether the voice feature data conforms to a normal distribution and to determine whether nonparametric processing or data transformation (such as Box-Cox transformation or logarithmic transformation) is necessary. This method ensures the rationality of data preprocessing and improves the stability of subsequent model training.

[0020] 3. Spectral centroid width, a new metric for time-domain feature extraction, measures the energy distribution of the voice signal near the center of the spectrum by calculating the spectral centroid width, enhancing feature differentiation. This more accurately reflects the timbre characteristics of the speech signal, improving the accuracy of voice assessment. A modified Griffin-Lim algorithm is used to estimate the phase of the short-time Fourier transform results, optimizing the reconstruction quality of the time-frequency features. The introduction of a momentum mechanism accelerates phase convergence, improves signal reconstruction accuracy, and reduces the impact of phase distortion on voice analysis.

[0021] 4. The present invention adopts the backward sequential feature selection method and the forward sequential feature selection method to optimize the feature set. Combining the two methods effectively reduces the model complexity and improves the classification accuracy and generalization ability of voice recognition.

[0022] 5. This invention uses an improved S4 model for training to enhance time series modeling capabilities. By optimizing calculations through diagonalization transformations, state update calculations are made more efficient, improving the model's ability to handle long-term dependencies and making it suitable for long-term time series modeling of voice signals.

[0023] 6. This invention forms a loss function based on a weighted combination of sample similarity contrast loss, asynchronous attention loss, and speech representation contrast loss. This combination of multiple optimization objectives effectively improves the model's classification stability, prevents overfitting, and enhances the accuracy of voice defect detection.

[0024] 7. The present invention is not limited to the health care and rehabilitation of the voice of the general population, but can even slow down the decline of the respiratory and swallowing functions of middle-aged and elderly people, and serve to improve the pronunciation ability of professional groups such as speakers, singers, and drama performers.

[0025] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:

[0027] Figure 1 1 is a schematic diagram of specific steps of a method for generating a voice training program according to an embodiment of the present invention;

[0028] Figure 2 1 is a flow chart of a method for generating a voice training program according to an embodiment of the present invention;

[0029] Figure 3 1 is a module diagram of a voice training program generation system according to an embodiment of the present invention;

[0030] Figure 4 Comparison of the performance of the algorithm proposed in this invention with other models;

[0031] Figure 5 Confusion matrix of the proposed algorithm. DETAILED DESCRIPTION

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0033] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0034] Figure 1A schematic diagram showing the specific steps of a method for generating a voice training program is shown. Figure 2 FIG. 1 shows a flow chart of a method for generating a voice training program. Figure 1 and Figure 2 As shown, a method 100 for generating a voice training program includes the following steps:

[0035] S110: collecting voice data, voice signals, and physiological characteristic data related to voice production and preprocessing them;

[0036] Among them, data collection: using a variety of sensors to collect voice data, such as microphones to collect sound signals, and using physiological sensors to obtain physiological characteristic data related to voice production, such as: static and dynamic information of breathing; body posture information, such as shrugging shoulders, hunching back, pelvic tilt, etc.; body muscle status information, because vocalization is a whole-body movement, the body's muscle status has a great influence on voice quality, such as (1) tension in the throat muscles, affecting the movement of the vocal cords and vocal muscles, (2) adhesion of the chest or back muscles affecting lung capacity, (3) weakness in the pelvic position and surrounding muscles, weak voice or even air leakage, etc.

[0037] The collected voice data includes multi-dimensional information such as pitch, tone, timbre, sound quality, and volume that reflects the voice data signal. Fundamental frequency and overtones are the physical basis of acoustic characteristics. The former corresponds to pitch perception, and the latter constructs timbre characteristics. Among them, the fundamental frequency determines the pitch, the overtone structure shapes the timbre and sound quality, and the overtone series affects the sound quality evaluation (such as the spectral characteristics of breathy and hoarse voices). The overtones of singers can extend to above 8000Hz (for example, Pavarotti's high-frequency overtones reach 4kHz).

[0038] The preprocessing of collected data includes:

[0039] The collected voice and physiological data are de-noised and filtered, the data format is standardized, and outliers are corrected or removed to ensure data quality. Data augmentation techniques are used to generate expanded samples, including performing audio cutting, audio padding, pitch shifting, fading, time masking, and time shifting on the original audio to generate expanded data samples.

[0040] Among them, the cross power spectrum density is used to clean the collected data. The cross power spectrum density is calculated by fast Fourier transforming the acoustic wave data collected by the sensor, converting it to the frequency domain, and then taking the average or absolute value to obtain the cross power spectrum density value CPSD. xy (f), Cross-power spectrum density calculation process:

[0041]

[0042] Among them, CPSD xy (f) is the cross power spectral density, Xn (f) and Y n (f) Fast Fourier transform result of the sound signal collected by the microphone or sensor; It's X n The complex conjugate of (f); N is the number of signal data points, that is, the number of samples.

[0043] When the calculated cross power spectrum density value CPSD xy (f)>1500 (through experimental testing of the impact of different thresholds on classification performance, it was found that when CPSD xy When (f)>1500, the signal-to-noise ratio is improved by ≥12dB and the classification accuracy is improved by 21.6%, so this threshold is preferred) and the corresponding data is retained to filter out weak signals, noise or irrelevant leakage signals and improve data quality.

[0044] S120: Extracting time domain and frequency domain features from the preprocessed data: the time domain features include any two or more of fundamental frequency, peak amplitude, signal energy, waveform features including skewness and autocorrelation measurement, nonlinear chaotic features, and entropy features; the frequency domain features include any two or more of harmonic noise ratio, fast Fourier transform median, energy distribution of speech signals in each frequency band, spectrum centroid and its width, Mel spectrum, short-time Fourier transform phase estimation optimized by improved Griffin-Lim algorithm, modulation spectrum, linear predictive coding coefficient, and Cepstral coefficient;

[0045] In step S120, the following steps are specifically included:

[0046] S121, perform a Shapiro-Wilk normality test on the preprocessed data to determine whether the features conform to a normal distribution. Normality testing is crucial for subsequent feature processing (such as standardization, dimensionality reduction, and model selection) because many statistical and machine learning methods assume that the data is normally distributed. If the features do not conform to a normal distribution, nonparametric methods or data transformations (such as Box-Cox transformation or logarithmic transformation) should be considered.

[0047] The Shapiro-Wilk test measures whether the sample is close to a normal distribution by calculating the statistic W, which is defined as follows:

[0048]

[0049] Among them, x (i) Represents the i-th data point after the sample data is arranged in ascending order; a i are precomputed weights (related to the expected value of the standard normal distribution); is the sample mean; the numerator Used to measure the degree of fit between data and normal distribution; the denominator is the population variance of the data.

[0050] If the data strictly follows a normal distribution, the W value is close to 1; if the data deviates from a normal distribution, the W value is small. If the W value is low, the normality assumption is rejected, indicating that most features do not follow a normal distribution. At this time, adopt at least one of the following strategies: 1) Data transformation: Logarithmic transformation, Box-Cox transformation, or Z-score standardization of the data to make it closer to a normal distribution. 2) Use non-parametric methods: such as the Mann-Whitney U test, Kruskal-Wallis test, or kernel density estimation. 3) Use robust methods to replace parametric statistics, including replacing the mean with the median, M estimation, or quantile regression.

[0051] Then, feature extraction of voice data is performed, which is divided into two dimensions: time domain feature extraction and frequency domain feature extraction.

[0052] S122, time domain feature extraction

[0053] Time domain features are mainly used to analyze the changes of speech signals over time and capture information such as their energy distribution, amplitude changes, and waveform shape.

[0054] 1) The fundamental frequency represents the dominant pitch of the voice and is the primary frequency of vocal cord vibration. This can be calculated by calculating the logarithmic spectrum of the speech signal, performing a Fourier transform on the spectrum, extracting the position of the highest peak, and converting it to the fundamental frequency.

[0055] 2) Peak amplitude refers to the maximum amplitude value of the speech signal within a certain time window.

[0056] 3) Signal energy is used to measure the strength of the speech signal.

[0057] 4) Waveform features are used to describe the shape and changing trend of speech signals. Common features include skewness and autocorrelation measurement. Among them, skewness is used to measure the symmetry of the waveform:

[0058]

[0059] SS: Skewness, describes the asymmetry of the data distribution and reflects the degree of skewness of the data relative to the mean. SS>0 means that the data is right-skewed (positive skewness), that is, the data is distributed more to the right of the mean. SS<0 means that the data is left-skewed (negative skewness), that is, the data is distributed more to the left of the mean. SS=0 means that the data distribution is symmetrical (close to normal distribution). N: The number of signal data points, that is, the number of samples. x(n): The value of the nth data point. μ: Mean, that is, the average value of the data set:

[0060]

[0061] σ: standard deviation, which measures the degree of dispersion of data:

[0062]

[0063] Autocorrelation measurements are used to assess the periodicity of a signal:

[0064]

[0065] R(τ): Autocorrelation function, a measure of the similarity of signals after a time lag (τ), is often used to evaluate the periodicity of a signal. If R(τ) peaks at certain intervals τ, it indicates that the signal has obvious periodicity at these time intervals. If R(τ) decays rapidly, it indicates that the signal is random or non-periodic. N: The total number of signal data points. x(n): The original signal data, that is, the signal value at time point n. x(n+τ): The signal data lagged by τ time steps, that is, the signal value after a delay of τ. τ: The time delay step in the autocorrelation calculation, usually taking different integer values (τ = 0, 1, 2, ...) to analyze the similarity of signals over time.

[0066] 5) Nonlinear chaotic characteristics, used to characterize the chaotic, fractal, and non-stationary characteristics of audio signals, are particularly suitable for analyzing pathological speech and emotional speech. The phase space is constructed by the delay embedding method, and the maximum Lyapunov exponent is calculated to measure the system's sensitivity to initial conditions and reflect the degree of chaos in the signal. The Lyapunov exponent calculation formula is:

[0067]

[0068] Where λ is the maximum Lyapunov exponent (unit: bits / s), which is used to quantify the chaotic characteristics of the speech signal. The larger the value, the more sensitive the system is to the initial conditions. max : Maximum evolution time (unit: ms), take 3 times of the fundamental frequency period, for example, when the fundamental frequency is 100Hz, t max =30ms. l (t): The Euclidean distance between the lth pair of adjacent points in the phase space at time t (unit: Pa·s). The phase space is constructed using Takens' theorem. l (0): The distance between neighboring points at the initial moment (unit: Pa·s). L: The number of valid neighboring point pairs, which must satisfy L ≥ 10M (M is the embedding dimension, usually 3-5) to ensure the stability of phase space reconstruction.

[0069] 6) Entropy features are used to measure the uncertainty and randomness of signals, including: Shannon entropy (measures the randomness of signals), spectral entropy (measures the flatness of spectrum distribution and distinguishes noisy signals from structured signals), and multiscale entropy (calculates entropy at different time scales and analyzes the dynamic characteristics of signals).

[0070] S123, frequency domain feature extraction

[0071] Frequency domain features are used to analyze the frequency distribution, energy concentration, and noise ratio of speech signals.

[0072] 1) The harmonic-to-noise ratio measures the ratio of a signal's periodic components (harmonics) to its nonperiodic components (noise). The fundamental frequency and its multiples are found using the autocorrelation function. The power of the periodic component and the power of the nonperiodic noise are calculated, and the ratio is then calculated.

[0073] 2) The Fast Fourier Transform (FFT) median represents the median frequency of the spectrum after the FFT calculation, meaning that half of the energy is distributed on either side of this frequency. The frequency at which the cumulative energy of the spectrum reaches 50% is the FFT median. The FFT spectra are calculated over different time windows, and the median frequency of the spectrum for each time period is recorded to show how the FFT median changes over time. This feature provides an overall picture of the frequency distribution of the audio, which is helpful for timbre analysis and classification tasks.

[0074] 3) Energy distribution of speech signals in different frequency bands (low frequency, medium frequency, high frequency). By setting the frequency bands, such as: low frequency band (0-500Hz), medium frequency band (500-2000Hz), high frequency band (above 2000Hz), calculate the total energy of each frequency band. Among them, overtones are usually located at integer multiples of the fundamental frequency, so they will contribute energy in the corresponding frequency band. By dividing the frequency bands more finely and calculating the total energy in each frequency band, the energy distribution of the overtone structure can be more accurately reflected. By extracting the energy distribution characteristics at integer multiples of the fundamental frequency, the contribution of the overtone series to the timbre characteristics is quantified. Specifically including: a. Overtone energy ratio: Calculate the ratio of the sum of the energy of the first N overtones (2-5 times the fundamental frequency) to the energy of the fundamental frequency. The formula is: Where f0 is the fundamental frequency and E(f) represents the energy at frequency f. b. Overtone attenuation slope: By linearly fitting the logarithmic attenuation curve of the first D overtone amplitudes, the slope value αα is calculated, which reflects the attenuation rate of high-frequency overtones. The formula is: α=slope(log(|A(c·f0)|)) where c=1,2,…,D; A(f) is the amplitude at frequency f. Overtone consistency index: Calculate the variance of the energy difference between adjacent overtone frequencies to evaluate the stability of the overtone series. The formula is: This indicator can effectively identify irregular vocal cord vibration (such as overtone breaks caused by vocal cord nodules).

[0075] 4) The spectrum centroid represents the center of the frequency distribution and reflects the brightness of the voice data.

[0076] Core: Spectral centroid, measures the center frequency of the spectrum, and is calculated as:

[0077]

[0078] The spectral centroid represents the "center of gravity" of the spectrum and is usually associated with the pitch or timbre characteristics of the sound. A higher spectral centroid indicates that the audio has stronger high-frequency components, while a lower spectral centroid indicates that the audio has stronger low-frequency components.

[0079] In this formula, K is the total number of frequency components, that is, the number of frequency points in the discrete spectrum. k : The kth frequency component, usually a discrete frequency point calculated in the discrete Fourier transform (DFT) or fast Fourier transform (FFT). S(f k ): The amplitude spectrum value of the kth frequency component, that is, at frequency f k The power or energy at a point (such as the amplitude calculated by FFT).

[0080] 5) Spectral centroid width is used to measure the extent of spectrum expansion and indicates the distribution range of energy:

[0081]

[0082] This formula is used to calculate the spectrum centroid width, which measures the dispersion of the spectrum relative to its center frequency (the spectrum centroid). A larger Band indicates a wider spectrum, and a smaller Band indicates a narrower spectrum.

[0083] Band refers to the width of the spectrum centroid, which indicates the extent of spectrum expansion relative to the center frequency, that is, the distribution of spectrum energy around the center frequency.

[0084] This formula calculates the square root of the mean square deviation of all frequency components relative to the spectrum's centroid, C, or the standard deviation of the frequencies. This describes the dispersion of the spectrum. If B is large, the frequency distribution of the spectrum is broad, meaning the signal contains many high-frequency components. If B is small, the spectrum is primarily concentrated within a narrow range, indicating a relatively concentrated frequency distribution.

[0085] 6) Mel Spectrogram: This converts the linear frequency scale to the Mel frequency scale (which matches the human ear's perception) and then calculates the power spectral density at the Mel frequency scale. This feature is widely used in speech processing and music analysis to capture the timbre and perceived frequency distribution of audio signals.

[0086] 7) Short-time Fourier transform (SFT) is used to analyze the time-frequency variation of a signal and provide a time-localized spectral representation. However, SFT usually only extracts amplitude information, while the lack of phase information affects the quality of signal reconstruction. Therefore, the present invention uses an improved Griffin-Lim algorithm to perform phase estimation on the SFT result:

[0087]

[0088] Among them, X (t′) is the spectrum estimate of the t′th iteration; |Y| is the target amplitude spectrum (calculated by short-time Fourier transform); stands for the inverse short-time Fourier transform operation; Represents element-wise multiplication (Hadamard product).

[0089] Phase recovery initialization starts from the target amplitude spectrum |Y| and randomly initializes the phase θ0 to form the initial spectrum By iteratively updating the phase, the inverse short-time Fourier transform is calculated to estimate the current spectrum X (t′) Transform back to the time domain. Perform a new short-time Fourier transform on the time domain signal to obtain a new spectrum estimate X (t′+1) The phase is adjusted according to the Griffin-Lim formula, gradually converging to the optimal value. If the mean square error of the phase change falls below a set threshold or the maximum number of iterations is reached, the iteration is terminated. By introducing a momentum mechanism during the Griffin-Lim iteration process, the phase estimation converges faster.

[0090] 8) Modulation spectrum is used to analyze how the amplitude of an audio signal changes over time. It is primarily used in tasks such as speech processing, music analysis, and pathological speech recognition. First, the short-time Fourier transform (STFT) is calculated to obtain the time-frequency representation of the audio. The short-time Fourier transform amplitude spectrum is low-pass filtered to obtain the modulation characteristics at different frequencies. The amplitude envelope signal on each frequency channel is Fourier transformed to obtain its frequency component, i.e., the modulation frequency. Modulation spectrum features include the modulation frequency distribution (the distribution of modulation amplitudes at different frequencies) and the modulation spectrum entropy (used to measure the complexity of the modulation spectrum).

[0091] 9) Linear Prediction Coding (LPC) is a feature extraction method used for spectral modeling and is commonly used in speech signal processing. The LPC coefficients estimate the signal to be linearly predicted from the past p samples. LPC coefficients can replace some spectral tilt parameters to more compactly represent audio features.

[0092] 10) Cepstral coefficients are used to represent the smooth shape of the spectrum. Common Cepstral features include Mel-frequency cepstral coefficients, first-order differences of Mel-frequency cepstral coefficients, and second-order differences of Mel-frequency cepstral coefficients.

[0093] S130: Filtering high-discriminative features from the initially extracted time-domain and frequency-domain features through backward / forward sequential feature selection to construct an optimized feature set, which is then input into an improved S4 neural network model for training. The S4 neural network model uses a diagonalized state transfer matrix and implicit Z-transform discretization to enhance long-term time series modeling capabilities.

[0094] The goals and objects of feature selection in step S130 include: Optimization object: From the initial 16-dimensional feature set consisting of time domain features (6 items) and frequency domain features (10 items), screen out the feature subset with the most discriminativeness for voice defect detection (such as incomplete vocal cord closure, vocal cord nodules, etc.). Optimization goal: Maximize the model classification performance (such as accuracy, F1 score, AUC) while reducing the model complexity (reducing the number of features). Preferably, an embodiment of the present invention can automatically screen out 5-8 core features through backward / forward sequential feature selection to avoid model overfitting. If extreme efficiency is pursued, the final optimized feature set of the embodiment includes no more than 10 feature subsets. Specifically, it includes:

[0095] S131: Feature Subset Screening

[0096] Sequential feature selection is a greedy algorithm that gradually builds feature subsets by adding or removing features one at a time. This embodiment of the present invention uses both backward and forward sequential feature selection methods to optimize the feature set based on the model's performance in detecting voice defects and determine the final feature subset.

[0097] Among them, backward sequential feature selection: starting from the initial set containing all time domain and frequency domain features, iteratively deletes the single feature that optimizes the model classification performance until the first termination condition is met (wherein the first termination condition is that deleting the feature causes the performance to drop by more than a preset threshold or the number of remaining features reaches a lower limit), thereby obtaining the first feature subset. Specifically: Backward sequential feature selection starts from the set containing all features. In each iteration, an attempt is made to delete a feature from the current feature subset. Then, an evaluation function (such as the accuracy of the classifier, F1 score, or mean squared error) is used to evaluate the performance of the feature subset after deleting the feature. All deletable features are traversed, and the feature that can optimize the evaluation function value (such as the highest accuracy, the lowest error, etc.) after deletion is selected for deletion. When further deleting features causes the evaluation function value to deteriorate, or when the preset stopping condition is met (such as the number of features is reduced to a certain threshold), the algorithm stops. The feature subset obtained at this time is the result of backward sequential feature selection. The following example is further given:

[0098] Initialization: Input: A complete set of all 16 time domain and frequency domain features F full ={f1, f2, ..., f 16}; Current feature subset: SS curren =F full , Model benchmark performance: Use the S4 neural network model to calculate the benchmark performance index P on the validation set base (such as accuracy).

[0099] Iteratively delete features: Traverse all features: from SS curren Temporarily remove a feature f from s Get subset SS temp =SS curren \{f s}; Evaluate performance: For each SS temp Train the S4 neural network model and calculate the validation set performance P s . Select the optimal deletion: find the s The highest f s , delete it permanently, update SS curren =SS temp Logical: Remove features that have the least impact on model performance (may be redundant or noisy features).

[0100] Termination condition: Performance degradation threshold: If deleting any feature causes P s If the value drops by more than δ (e.g. δ = 1%), the iteration is stopped. curren The process ends when the number of features in the function drops to a preset value (e.g. 5). Output: Feature subset SS filtered by backward feature selection backward .

[0101] Among them, forward sequential feature selection: taking the first feature subset as input, iteratively adding a single feature that optimizes the model classification performance until the second termination condition is met (wherein, the second termination condition is that the performance improvement after adding the feature is lower than the preset threshold or the number of features reaches the upper limit), and the final feature subset is obtained. Specifically: forward sequential feature selection starts with an empty feature set. In each iteration, consider adding the features that have not been selected in the original feature set to the current feature subset one by one, and then use the evaluation function to evaluate the performance of the added feature subset. Traverse all unselected features, select the feature that can make the evaluation function value optimal after addition, and add it to the feature subset. Termination condition: When continuing to add features can no longer further optimize the evaluation function value, or when the preset stopping condition is reached (such as the number of features reaches a certain upper limit), the algorithm stops and obtains the final feature subset. The following example is further given:

[0102] Initialization: Input: Empty feature set Remaining feature pool: F pool =F full .

[0103] Iteratively add features: Traverse the remaining features: Start from F pool Add a feature f s′ to SS curren , get the subset SS temp =SS curren ∪{f s′}. Evaluate performance: For each SS temp Train the S4 neural network model and calculate the validation set performance P s′ . Select the optimal addition: find the s′ The highest f s′ , add it permanently, update SS curren =SS temp and from F pool Remove f s′ .

[0104] Termination condition: Performance saturation threshold: If P s′ If the improvement is less than ∈ (e.g. ∈ = 0.5%), stop. Upper limit of feature quantity: When SS curren The process terminates when the number of features reaches a preset value (e.g. 10). Output: Feature subset SS selected by forward sequential feature selection forward .

[0105] Forward sequential feature selection, which starts with an empty set and gradually adds features, is more suitable for scenarios with a large number of features and limited understanding of their interactions. Backward sequential feature selection, which starts with the full set and removes features, can be more effective when it's known that some features may be less important or when you want to gradually reduce a large feature set. Both methods are greedy and cannot guarantee a globally optimal feature subset, but they often find good local optima in practice, have relatively low computational complexity, and are easy to implement.

[0106] Furthermore, in some embodiments of the present invention, a combination strategy of backward and forward serial selection is adopted. Specifically, backward sequential feature selection is first performed to remove redundant features from the 16 features to obtain SS backward (If there are 6 features left). backwardAs the input feature pool for forward sequential feature selection, important features that were mistakenly deleted by backward sequential feature selection are further added. Thus, by first coarsely screening to remove obvious redundant features and then finely supplementing key features, it is possible to avoid falling into local optimality in the early stage of forward selection. For example, in one embodiment of the present invention, the process of determining the final feature subset based on model performance includes: cross-validation: for each candidate feature subset, a 5-fold cross-validation is used to calculate the average accuracy / F1 score. Voice defect classification task: the model output is the type of voice abnormality (such as vocal cord paralysis, vocal cord nodule), and the confusion matrix and AUC are calculated. High-discrimination features: innovative features such as spectral centroid width (Band), improved Griffin-Lim phase estimation, and modulation spectrum are retained in most subsets. Low-contribution features: waveform skewness (SS): due to its low contribution to classification (<1% accuracy improvement), it is preferentially deleted in the Backward stage. Linear predictive coding coefficient (LPC): highly correlated with MFCC and discarded in the Forward stage. Finally, eight core features were screened out through the combination of backward and forward concatenation strategies, including: spectral centroid width (Band), nonlinear chaotic characteristics (Lyapunov exponent), improved Griffin-Lim phase estimation, modulation spectrum, harmonic-to-noise ratio, Mel spectrum diagram, autocorrelation measurement (R(τ)), and fast Fourier transform median.

[0107] This step uses performance-driven feature addition and subtraction to avoid relying on manual experience to select features. Combining the two approaches offsets their respective limitations (e.g., the potential for inadvertent deletion of weakly relevant features in the backward pass and the potential for falling into a local optimum in the forward pass). The resulting feature subset dimension is reduced (e.g., from 16 to 8), reducing S4 model training time while maintaining or improving classification performance.

[0108] S132: Neural Network Model Construction

[0109] A modified S4 neural network model is used as the selective state-space model component. Six sequential selective state-space model layers are used to learn long-range dependencies between features in the temporal dimension. Each selective state-space model layer processes multiple channels in parallel, extracting long-range dependency features from the sequence without changing the frequency channel. This complements the frequency-domain features extracted by the multi-criteria asynchronous neural network. The multi-criteria asynchronous neural network uses an asynchronous attention mechanism to enhance feature selection, eliminating the need for data clustering, reducing statistical irrelevance, and improving classification stability.

[0110] Among them, the basic idea of the improved S4 model is to convert sequence modeling into a continuous-time state equation:

[0111]

[0112] y(t)=Ch(t)+Dx(t)

[0113] Among them, h(t) is the hidden state; x(t) is the input signal; y(t) is the output signal; A, B, C, D are trainable parameters.

[0114] Since the neural network needs to be calculated at discrete time steps, implicit Z transform is used for discretization:

[0115] h[t+1]=A Δ h[t]+B Δ x[t]

[0116] y[t]=C Δ h[t]+D Δ x[t]

[0117] Among them, h[t]: state variable, used to store the state information of the system at time step t. It is similar to the hidden state in the recurrent neural network, but here it is based on the selective state space model; x[t]: input variable, representing the input data of the time series, such as speech signals, text embeddings, etc.; y[t]: output variable, that is, the output of the model at time step t, which can be used for prediction, classification or regression tasks; A Δ : The state transition matrix controls how the state variable h[t] evolves over time. After discretization, it is used to calculate the state update from the current state to the next moment; B Δ : The mapping matrix from input to state, which controls how the input x[t] affects the state variable h[t]; C Δ : The mapping matrix from state to output, which determines how the state variable h[t] affects the final output y[t]; D Δ : The direct mapping matrix from input to output, which determines whether the input x[t] directly affects the output without passing through the state variable h[t]; A Δ , B Δ , C Δ , D Δ The matrix is numerically approximated and optimized, which effectively improves the time series modeling capability and reduces the computational complexity.

[0118] The implicit Z transform is discretized using a bilinear transformation method with a step size of Δ=0.1.

[0119] The improved S4 model uses diagonal transformation to optimize the calculation:

[0120]

[0121] Where P is a diagonalization matrix used to transform the state matrix A into a diagonal form to simplify calculations. The diagonalized state transfer matrix is used to accelerate the state update calculation. Since the power calculation of the diagonalized matrix can be simplified to element-by-element exponential operation, the computational complexity is reduced. The transformed input mapping matrix and state-to-output mapping matrix. The diagonalized matrix P is iteratively optimized by QR decomposition, for example, once every 100 training steps.

[0122] After diagonal transformation, the recursive calculation formula becomes:

[0123]

[0124] The computational approach is more efficient, as diagonal matrix exponential operations can be performed element-by-element, significantly increasing computational speed. This enables the improved S4 model to more efficiently model long time series, parallelizing multiple time steps and improving computational efficiency, making it particularly suitable for environments with limited computing resources.

[0125] The improved S4 neural network model adopts a multi-task output head design, which specifically includes:

[0126] Classification branch: Softmax output layer, nodes corresponding to six abnormal types of vocal cords (good vocal cords, incomplete closure, excessive tension, vocal cord paralysis, vocal cord nodules, unstable pitch control) and normal categories, output probability distribution.

[0127] Regression branch: Outputs a multidimensional deviation vector ΔF∈R of key voice features and standard values 8 (Such as fundamental frequency jitter rate, harmonic noise ratio and other core characteristics).

[0128] Training parameters: The model was initialized with He, using the AdamW optimizer (learning rate 3e-4, weight decay 0.01), updating the diagonalized matrix P every 100 steps, with a batch size of 32 and 200 training epochs.

[0129] S140: Construct a weighted loss function that includes sample similarity contrast loss, asynchronous attention loss, and speech representation contrast loss to optimize model classification stability.

[0130] Using multiple optimization objectives, define the total loss function:

[0131]

[0132] in, is the sample similarity contrast loss (for classification tasks); is asynchronous attention loss (used to reduce statistical irrelevance); is the speech characterization contrast loss (used to prevent overfitting); λ1, λ2, λ3 are weight hyperparameters, preferably, for example, λ1 = 0.6 (contrastive loss), λ2 = 0.3, λ3 = 0.1.

[0133] Among them, the sample similarity comparison loss and speech representation contrast loss Use contrastive loss to learn robust representations for unlabeled data. Sample similarity contrastive loss A contrastive learning method is used, which aims to maximize the similarity of positive sample pairs while minimizing the similarity of negative sample pairs.

[0134] Among them, the sample similarity comparison loss The calculation formula is as follows:

[0135]

[0136] Among them, z i and It is the representation vector of the same data after different data enhancements, which serves as a positive sample pair; is a hyperparameter used to scale similarity; N is the mini-batch size, and each sample is matched with the enhanced version during calculation; sim(z i , z j ) is the cosine similarity, which measures their similarity in vector space; Similarity contribution of positive pairs (different augmented versions of the same data); Include all samples in the current mini-batch (excluding itself) as a set of negative sample pairs and calculate the similarity contribution of all samples.

[0137] Speech Representation Contrastive Loss For learning robust speech representations, the method is similar to But different sets of positive sample pairs are defined:

[0138]

[0139] P(i): This set contains all i The relevant positive sample index includes not only the data-enhanced version, but also similar speech segments. For example, different speech segments from the same speaker or segments with similar acoustic features; With z i The sum of similarity contributions of all relevant positive pairs, not just one enhanced version; and The same, including all samples (excluding themselves), as a set of negative sample pairs.

[0140] The characteristic of this method is that it can utilize the continuity of the sound signal and take adjacent time segments as positive samples to enhance the robustness to timing changes.

[0141] Asynchronous Attention Loss The core idea is to reduce the temporal correlation of attention distribution, preventing the model from over-relying on short-term patterns and ignoring long-term dependencies. This is achieved by making the attention weight matrix A as close to the identity matrix as possible to reduce its autocorrelation:

[0142]

[0143] Among them, A tt : Attention weight matrix, each row represents the attention distribution at different time steps; I: Identity matrix, ideally It should be as close to the identity matrix as possible to reduce the correlation between time steps; Frobenius norm, which measures the difference between matrices.

[0144] This loss encourages attention distributions to remain independent across different time steps and reduces redundant information.

[0145] S150: Analyze voice type based on the trained model, dynamically generate personalized voice training plan and make dynamic adjustments.

[0146] First, based on the S4 neural network model's output of the probability distribution of six abnormality types and normal categories, and the multidimensional deviation vector ΔF of key voice features from standard values, possible voice abnormalities, such as vocal cord insufficiency, vocal cord hypertonia, vocal cord paralysis, or vocal cord nodules, were identified.

[0147] Abnormality determination process: ① When the maximum probability of a classification branch exceeds the threshold θc (θc = 0.7), it is directly determined to be the corresponding abnormal category; ② If the probability does not reach θc, the square envelope spectrum index ENVSI is used to verify the effectiveness of the model in detecting voice defects. For example, if the highest confidence level is in the range of 50%-70%, the square envelope spectrum index ENVSI is used to verify the effectiveness of the model in detecting voice defects. The formula is:

[0148]

[0149] Among them, AIS 2 (m) is the square amplitude of the information signal, SES 2 (q) is the square envelope spectrum, M1 and M2 are the number of characteristic points of the signal and noise, respectively. The algorithm's ability to detect voice defects is evaluated by calculating metrics based on the square envelope spectrum.

[0150] The user's ENVSI value is calculated and then checked for consistency with the anomaly type output by the model. If the ENVSI is abnormal and the model does not detect it, a secondary analysis is triggered (such as resampling the signal or increasing the depth of the time series modeling).

[0151] Combined with a voice quality perception model or a standard voice sample library, the voice is classified based on changes in spectral characteristics and analyzed for potential physiological characteristics. For example, a decrease in high-frequency components and an increase in breathy components may indicate uneven vocal cord vibration or incomplete glottal closure. Excessive fundamental frequency fluctuations may indicate weak vocal cord muscle control, affecting pitch stability.

[0152] Among them, generating a personalized voice training program and making dynamic adjustments include:

[0153] Based on the voice defect classification results, targeted training movements are matched from the training movement library:

[0154] 1) Semi-closed voice training for vocal cord insufficiency, including bubble blowing training (using the LSVT method of blowing bubbles in a cup, maintaining a 10 cm bubble chain for 15 seconds), nasopharyngeal resonance training (soft palate elevation ≥45° when pronouncing / m / , glottal depression controlled at 8-12 cmH2O), and low-frequency humming training (fundamental frequency maintained at 120±5 Hz, sound intensity level 65-70 dB SPL), to enhance glottal closure ability and improve voice efficiency;

[0155] 2) Excessive vocal cord tension requires relaxation training, including soft vocalization or chewing phonation, to help relax the throat muscles and reduce unnecessary muscle tension;

[0156] 3) For unstable pitch control, corresponding scale exercises, including glissando training or crescendo and decrescendo training, can be used to improve vocal cord control, pitch stability and sound quality fluency.

[0157] In addition, through real-time feedback monitoring of training effects, combined with model output data, the user's training progress, physical condition and feedback, training parameters can be dynamically adjusted or training plans can be modified:

[0158] 1) Initial stage of training: To avoid vocal cord fatigue, set lower intensity training, such as reducing training time, reducing practice frequency or adjusting volume requirements.

[0159] 2) Adaptation stage: As the user's vocal control improves, the training difficulty is gradually increased, such as increasing the scale span, extending the vocalization time, or introducing more complex resonance training.

[0160] 3) Advanced stage: After users have mastered basic vocal skills, they can further optimize their personalized training, such as strengthening specific vocal ranges, improving speech clarity, or conducting vocal endurance training.

[0161] Based on physiological adaptability, progressive training is designed (initial fatigue prevention → adaptation period reinforcement → advanced period customization) to improve rehabilitation efficiency. The entire training process can be monitored through a neural network model with real-time feedback to prompt users to adjust their training methods, ensuring efficient and scientific training, ultimately helping users improve voice quality and restore normal vocal function. Real-time model monitoring (such as fundamental frequency jitter and spectral centroid changes) drives program iteration, which is different from a static training plan.

[0162] Furthermore, the improved S4 neural network model is deployed in a low computing resource environment, supports parallel state update calculations, and outputs voice abnormality classification results in real time.

[0163] Figure 3 FIG. 1 shows a module diagram of a voice training program generation system. Figure 3 As shown, a voice training program generation system 200 includes:

[0164] The data acquisition and preprocessing module 210 is used to configure multiple sensors to collect voice signals and physiological characteristic data related to voice production and perform data preprocessing;

[0165] The time-frequency feature extraction module 220 is used to extract time-domain and frequency-domain features from the preprocessed data;

[0166] A model training module 230 is used to integrate the improved S4 neural network model and the multi-objective loss function, optimize the feature set through backward / forward sequential feature selection, and then input the feature set into the improved S4 neural network model for training;

[0167] The voice evaluation module 240 includes a square envelope spectrum analyzer and a personalized program generator, which is used to analyze the voice type based on the trained model, dynamically generate a personalized training program and perform dynamic adjustments.

[0168] In the above-mentioned voice training program generation system, the data acquisition and preprocessing module 210 transmits the preprocessed data to the time-frequency feature extraction module 220. The feature set is input into the model training module 230 after feature selection optimization. The voice evaluation module 240 receives the model output in real time and generates a training program, forming a closed-loop feedback.

[0169] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0170] For example, user data: fundamental frequency jitter rate: 3.2% (normal <1%), spectrum center of mass offset: +15%, harmonic-to-noise ratio: 8dB (normal >12dB), generate the following plan: 1-3 days: bubble sound training (intensity 0.3, 5 minutes / time); 4-7 days: add glissando training (interval span ±3 semitones); after 8 days: add resonance training (enhance nasal resonance).

[0171] Experimental comparison:

[0172] Figure 4 The proposed algorithm is compared with Faster R-CNN, EfficientDet, WaveNet, SSD, and RetinaNet in terms of mean average precision during training. The proposed algorithm outperforms the comparison algorithms at all training rounds, and especially after the 10th round, the mean average precision is significantly higher than those of Faster R-CNN, EfficientDet, SSD, and other methods.

[0173] Figure 5 The confusion matrix performance of the proposed algorithm is demonstrated. This experiment classifies vocal cord disorders into categories such as "incomplete vocal cord closure," "excessive vocal cord tension," "unstable pitch control," "vocal cord paralysis," "vocal cord nodules," and "healthy vocal cords." By analyzing the confusion matrix, we can further explore the model's classification performance and identify potential optimization areas.

[0174] The diagonal elements of the confusion matrix represent the number of correct classifications for each category, with high values indicating that the model performs well in that category. Specifically, the category with good vocal cords had the highest number of correct classifications, at 45, indicating that the features of normal vocal cords are relatively clear and easy to identify. The correct classifications for vocal cord nodules were 42, indicating that this category has good discriminability in the data distribution and that the model can easily identify its features. The correct classifications for vocal cord paralysis and vocal cord insufficiency were 38 and 36, respectively, which performed relatively well, but there were still some misclassifications. The correct classifications for unstable pitch control were 34, indicating that this category has complex features and is easily confused with other categories. The correct classifications for vocal cord overstrain were 33, the worst performance of all categories, possibly because the features of this category are similar to those of multiple categories, resulting in a high number of misclassifications.

[0175] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. This is not limited herein.

[0176] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for generating a voice training program, characterized in that: The following steps are involved: Collecting voice data, voice signals, and physiological characteristic data related to voice production and preprocessing them; wherein the physiological characteristic data related to voice production includes static and dynamic information of breathing, body posture information, and muscle state information; Extract time domain and frequency domain features from the preprocessed data: the time domain features include any two or more of fundamental frequency, peak amplitude, signal energy, waveform features including skewness and autocorrelation measurements, nonlinear chaotic features, and entropy features; the frequency domain features include any two or more of harmonic noise ratio, fast Fourier transform median, energy distribution of speech signals in each frequency band, spectral centroid and its width, Mel spectrum, short-time Fourier transform phase estimation optimized by the improved Griffin-Lim algorithm, modulation spectrum, linear predictive coding coefficients, and Cepstral coefficients; An optimized feature set is constructed by screening high-discriminative features from the initially extracted time-domain and frequency-domain features through backward / forward sequential feature selection. This feature set is then fed into an improved S4 neural network model for training. The S4 neural network model uses a diagonalized state transfer matrix and implicit Z-transform discretization to enhance long-term time series modeling capabilities. Construct a weighted loss function that includes sample similarity contrast loss, asynchronous attention loss, and speech representation contrast loss to optimize model classification stability; Analyze voice type based on the trained model, dynamically generate personalized voice training plans and make dynamic adjustments.

2. The method for generating a voice training program according to claim 1, wherein: in, Preprocessing of the collected voice data includes: The cross-power spectrum density method is used to calculate the frequency domain cross-spectrum of the acoustic wave signals collected by multiple sensors, and low-quality signals with a cross-power spectrum density value of ≤1500 are filtered out; The data enhancement technology is used to generate the expanded samples, including performing at least one operation of audio cutting, audio filling, pitch conversion, fade-in and fade-out, time masking and time shifting on the original audio.

3. The method for generating a voice training program according to claim 2, wherein: in, Extract time domain and frequency domain features from the preprocessed data, including: The Shapiro-Wilk normality test is performed on the preprocessed data. The data distribution is determined by calculating the statistic W. If the W value is lower than the preset threshold, the data is considered to be non-normally distributed, and at least one of the following treatments is performed: 1) Data were subjected to Box-Cox transformation, logarithmic transformation, or Z-score standardization; 2) nonparametric methods were used, including the Mann-Whitney U test, Kruskal-Wallis test, or kernel density estimation; 3) robust methods were used to replace parametric statistics, including replacing the mean with the median, M estimation, or quantile regression.

4. The method for generating a voice training program according to claim 3, wherein: in, The method of selecting high-discrimination features from the initially extracted time domain features and frequency domain features through backward / forward sequential feature selection to construct an optimized feature set includes: Perform backward sequential feature selection: starting from the initial set containing all time domain and frequency domain features, iteratively delete the single feature that optimizes the model classification performance until the first termination condition is met, obtaining the first feature subset; Perform forward sequential feature selection: using the first feature subset as input, iteratively add a single feature that optimizes the model classification performance until the second termination condition is met to obtain the final feature subset; wherein, the model classification performance is quantified by an evaluation function, including classification accuracy, F1 score or mean square error; the first termination condition is that the performance degradation caused by deleting a feature exceeds a preset threshold or the number of remaining features reaches a lower limit; the second termination condition is that the performance improvement after adding a feature is lower than a preset threshold or the number of features reaches an upper limit.

5. The method for generating a voice training program according to claim 1, wherein: in, The construction of the improved S4 neural network model includes: The sequence modeling is converted into a continuous-time state equation; the state equation is discretized through implicit Z transformation; the state equation after diagonal transformation is obtained by optimization calculation using diagonalization transformation; multi-channel input is processed in parallel through six layers of selective state space model layers connected in series, the long-range dependency features of the time series are extracted, and the features are complemented with the frequency domain feature outputs of a multi-criteria asynchronous neural network, which uses an asynchronous attention mechanism to reduce statistical irrelevance.

6. The method for generating a voice training program according to claim 5, wherein: in, Construct a weighted loss function that includes sample similarity contrast loss, asynchronous attention loss, and speech representation contrast loss to optimize model classification stability, including: Use multiple optimization objectives to define the total loss function in, is the sample similarity comparison loss; for asynchronous attention loss; is the speech representation contrast loss; λ1, λ2, λ3 are weight hyperparameters.

7. A method for generating a voice training program according to claim 6, characterized in that: in, The voice type analysis based on the trained model includes: Based on the output of the S4 neural network model, identifying at least one abnormality type of vocal cord insufficiency, vocal cord hypertonia, vocal cord paralysis, or vocal cord nodule; The effectiveness of voice defect detection was verified by using the squared envelope spectrum index; Combined with the voice quality perception model or standard voice sample library, the voice is classified according to the changes in spectral characteristics and its possible corresponding physiological characteristics are analyzed.

8. The method for generating a voice training program according to claim 7, wherein: in, Generating a personalized voice training program and dynamically adjusting it includes: Based on the results of voice defect classification, targeted training exercises are matched from the training exercise library: 1) Incomplete vocal cord closure corresponds to semi-closed phonation training, including bubble blowing training, nasopharyngeal resonance training, or low-frequency humming training; 2) Excessive vocal cord tension corresponds to relaxation training, including soft pronunciation or chewing phonation; 3) Unstable pitch control corresponds to scale exercises, including glissando training or crescendo and decrescendo training; Dynamically adjust training parameters based on the user's training stage: 1) Reduce intensity, duration, frequency, or volume in the early stages of training; 2) Increase scale span, extend phonation time, or introduce complex resonance training during the adaptation stage; 3) Strengthen specific vocal ranges, optimize speech clarity, or conduct vocal endurance training during the advanced stage; Monitor the training effect through real-time feedback, combine it with model output data, and dynamically modify the training plan.

9. The method for generating a voice training program according to claim 8, wherein: in, The improved S4 neural network model is deployed in a low computing resource environment, supports parallel state update calculations, and outputs voice abnormality classification results in real time.

10. A voice training program generation system, used to implement a voice training program generation method according to any one of claims 1 to 9, characterized in that: include: A data acquisition and preprocessing module, configured to configure multiple sensors to collect voice signals and physiological characteristic data related to voice production and perform data preprocessing; The time-frequency feature extraction module is used to extract time domain and frequency domain features from the preprocessed data; A model building and training module is used to integrate the improved S4 neural network model and the multi-objective loss function, optimize the feature set through backward / forward sequential feature selection, and then input the feature set into the improved S4 neural network model for training; The voice assessment module includes a square envelope spectrum analyzer and a personalized program generator, which is used to analyze the voice type based on the trained model, dynamically generate personalized training programs and make dynamic adjustments.

Citation Information

Cited By

  • Voice processing method based on artificial intelligence

    CN120748414A

  • Training method and system of vocal cord tumor multi-classification analysis model based on voice

    CN121980362A