Sound standardization evaluation system for multi-dimensional acoustic parameter and emotional scene dynamic analysis

By constructing a sound standardized evaluation system for multi-dimensional acoustic parameters and dynamic analysis of emotional scenes, the shortcomings of multi-dimensional sound control and emotional adaptability in the existing technology are solved, and comprehensive and accurate evaluation and personalized feedback of speech expression are achieved.

CN120340540AActive Publication Date: 2025-07-18GUANGZHOU SENJI SOFTWARE TECH CO LTD

Patent Information

Application Number
CN202510822206.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The existing voice evaluation system has shortcomings in multi-dimensional sound control, emotional and deductive scene adaptability, pronunciation parameter system and system structure, and has failed to achieve comprehensive and dynamic evaluation and feedback.

Method used

A sound standardized evaluation system is built for multi-dimensional acoustic parameters and dynamic analysis of emotional scenes, including audio acquisition, processing, standard evaluation and integrated processing modules, multi-modal fusion analysis is carried out through machine learning models, and a structured evaluation report is generated.

Benefits of technology

Multi-dimensional evaluation of speech expression is realized, the accuracy of evaluation and personalized matching capabilities are improved, and a closed-loop evaluation and training system is formed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340540A_ABST
    Figure CN120340540A_ABST
Patent Text Reader

Abstract

The invention provides a sound standardization evaluation system for multi-dimensional acoustic parameter and emotional scene dynamic analysis, and aims to realize structured and objective evaluation of voice performance. The system comprises an audio acquisition module, an audio processing module, a standard evaluation module and an integration processing module. The audio processing module performs noise reduction, pre-emphasis, sampling rate adjustment, framing and slicing operation on the original voice data; the standard evaluation module executes character and pronunciation accuracy analysis, acoustic basic skill feature evaluation or emotion and scene adaptation degree analysis based on multi-modal fusion according to an evaluation type set by a user; parameters including fundamental frequency, formant, sound intensity, fundamental frequency perturbation, amplitude perturbation and the like are extracted in evaluation, and classification prediction is completed in combination with a machine learning model; and finally, the integration processing module generates an evaluation report including a scoring result, question prompts and personalized training suggestions, and provides closed-loop feedback to support sound training and expression optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice evaluation, and particularly relates to a voice standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios. Background Art

[0002] With the wide application of voice training and pronunciation evaluation technologies in fields such as education, entertainment, and medical treatment, the related technical systems have gradually shifted from manual scoring to automated evaluation. Most existing voice evaluation systems mainly focus on basic voice recognition and pronunciation accuracy judgment, such as the recognition of Chinese phonetic initials, finals, and tones, and the comparison of English phonetic symbol pronunciations. Some systems have initially introduced emotion recognition models to identify basic emotion types such as happiness, anger, sadness, and fear, and use voice scoring and feedback mechanisms to assist language learning. However, most systems still rely on static scoring mechanisms and are mainly single-dimensional.

[0003] The above systems generally have the following problems: First, in terms of multi-dimensional voice control, the existing technologies fail to achieve systematic quantitative evaluation of parameters such as breath strength, resonance cavities (such as oral cavity resonance, thoracic cavity resonance), and the conversion between real and virtual voices. Second, in terms of the adaptability of emotions and performance scenarios, they can only identify basic emotion types, lack the ability to adapt to fine-grained emotions such as smugness, pleading, and questioning, and also lack the ability to dynamically prompt and integrate scenarios such as role-playing and narration. Third, in terms of the pronunciation parameter system, the existing solutions mostly stay at the level of the correctness of pronunciation, and do not cover professional intonation features such as sandhi, rhythm control, pauses, and connections. Fourth, there are problems such as module fragmentation, single evaluation algorithm, and fragmented feedback mechanism in the system structure, and a closed-loop system integrating evaluation, analysis, and training has not yet been formed.

[0004] Therefore, there is an urgent need to provide a voice standardization evaluation system that integrates acoustic analysis, linguistic modeling, and multi-modal emotion recognition to meet the comprehensive evaluation requirements for multi-dimensional vocal performance and deductive ability. Summary of the Invention

[0005] The present application provides a voice standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios to achieve the integrated evaluation of multi-dimensional acoustic features and complex emotional scenarios in speech expression, thereby improving the accuracy of voice evaluation.

[0006] The present application provides a voice standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios, including: An audio acquisition module for acquiring the user's voice audio and generating original audio data; An audio processing module for preprocessing the original audio data to obtain preprocessed audio data; wherein, the preprocessing includes noise reduction, sampling rate adjustment, pre-emphasis, framing, and slicing operations; A standard evaluation module is used to determine the evaluation type according to user settings; receive the preprocessed audio data and perform an evaluation operation corresponding to the evaluation type, and the evaluation operation includes at least one of the following operations: Identify speech units in the preprocessed audio data, calculate the accuracy rates of initial consonants, final vowels, and phonetic sandhi linguistic features, and generate a pronunciation evaluation result; Extract acoustic features from the preprocessed audio data, including fundamental frequency, formant, sound intensity, fundamental frequency perturbation, and amplitude perturbation, and input the features into machine learning models separately trained for different evaluation items for classification prediction to generate a basic skills evaluation result; According to the user's emotion or scenario information, construct a prompt word template, and perform multimodal fusion analysis on the preprocessed audio data and scenario text, output emotional matching degree, scene restoration degree, and emotion coherence indicators, and generate an emotion evaluation result; An integration processing module is used to summarize the evaluation results generated by the standard evaluation module, generate a structured evaluation report, and the evaluation report includes dimension scores, corresponding problem prompts, and personalized training suggestions, and is transmitted back to the user terminal.

[0007] The beneficial effects of this application mainly include: (1) By constructing a complete audio preprocessing link, including noise reduction, sampling rate adjustment, pre-emphasis, framing, and slicing operations, the quality of audio signals and the stability of subsequent analysis are significantly improved, providing technical support for high-precision acoustic feature extraction. (2) By integrating linguistic features, acoustic parameters, and emotional context, it is possible to evaluate the user's speech from multiple dimensions, realizing a comprehensive evaluation from pronunciation accuracy to vocal skills and then to emotional expression, overcoming the problem of fragmented evaluation dimensions in the prior art. (3) Based on the evaluation task call and training model matching mechanism, it is possible to dynamically select the optimal algorithm model for classification prediction, ensuring that various evaluation items operate under the dual optimization of accuracy and adaptability, and improving the scientific nature and personalized matching ability of the evaluation. Description of the Drawings

[0008] Figure 1 is a schematic diagram of a sound standard evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios provided by the first embodiment of this application.

[0009] Figure 2 is a timing diagram of the user speech evaluation process involved in the first embodiment of this application.

[0010] Figure 3 is a heat map of the correlation matrix between acoustic features involved in the first embodiment of this application. Detailed Embodiment

[0011] In the following description, numerous specific details are set forth to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0012] The first embodiment of the present application provides a sound standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios. Please refer to Figure 1 , which is a schematic diagram of the first embodiment of the present application. The following will combine Figure 1 to detail a sound standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios provided by the first embodiment of the present application.

[0013] The sound standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios includes an audio acquisition module 101, an audio processing module 102, a standard evaluation module 103, and an integration processing module 104.

[0014] The audio acquisition module 101 is used to acquire the user's voice audio and generate raw audio data.

[0015] In the present invention, the function of the audio acquisition module 101 is to acquire the user's voice audio and generate raw audio data, specifically including but not limited to obtaining the voice signal from the user's terminal device (such as a microphone, mobile phone, tablet computer, or computer with voice input function), and converting the obtained analog audio signal into a digital signal form in real-time or non-real-time for subsequent data processing and analysis.

[0016] In a preferred embodiment, the audio acquisition module 101 includes an audio input interface, an analog-to-digital conversion unit, and a buffer unit. The audio input interface is used to receive the user's original voice, such as the user's reading of a demonstration text, an emotional recitation text, or a free expression voice input. The analog-to-digital conversion unit converts the analog voice signal into digital audio data, and the sampling rate is preferably set to 16 kHz or 44.1 kHz, and the quantization bit depth is not less than 16 bits to ensure clear enough sound quality for accurate extraction of subsequent acoustic features. The sampling duration is dynamically controlled according to specific application tasks. For example, the sampling duration for basic pronunciation evaluation is generally 3 to 10 seconds, and the audio duration for emotional interpretation or role-playing evaluation can be 10 seconds to 60 seconds.

[0017] During the acquisition process, the audio acquisition module 101 preferably integrates a voice activity detection mechanism to filter out invalid inputs. For example, when the user is not speaking or the background noise intensity exceeds a threshold, the acquisition is not started or the user is prompted to re-record. In addition, to ensure the recording effect, the system can monitor the signal-to-noise ratio (SNR) of the recording signal in real time. When it is detected that the SNR is lower than a set standard (e.g., lower than 30 dB), a prompt can be issued through the user interface or the acquisition can be automatically paused.

[0018] In a specific implementation, the audio acquisition module 101 supports multiple re-recordings of voice acquisition and a version control mechanism. The user can input multiple recording versions for the same task. The system identifies and numbers the audio data through a metadata management mechanism to ensure that each piece of original voice data has a unique identifier and is accompanied by information such as the acquisition time, acquisition device number, and acquisition user identifier, which is convenient for subsequent task invocation, training optimization, and evaluation comparison.

[0019] The audio acquisition module 101 may further include an echo cancellation and noise reduction function to reduce the echo interference caused by the hardware device itself and eliminate the ambient background noise in non-speech segments, such as air conditioner sounds and electric fan sounds, etc., to improve the cleanliness and effectiveness of the original audio data from the source.

[0020] Therefore, the audio acquisition module 101 not only undertakes the basic responsibility of digitizing the user's voice signal, but also provides original voice input data with reliable technical quality, clear structure, and strong traceability for the subsequent evaluation process through mechanisms such as analog-to-digital conversion, SNR monitoring, silent filtering, multi-version recording, and original data identification. This module constitutes the data entry point of the entire voice standardization evaluation system and plays a fundamental supporting role in the system performance.

[0021] The audio processing module 102 is used to preprocess the original audio data to obtain preprocessed audio data; wherein, the preprocessing includes noise reduction, sampling rate adjustment, pre-emphasis, framing, and slicing operations.

[0022] The audio processing module 102 is used to perform comprehensive preprocessing operations on the original audio data generated by the audio acquisition module 101, aiming to improve the audio quality, unify the data format, and provide a stable and standardized input basis for feature extraction and model analysis in subsequent evaluation links. This module specifically includes a series of sequentially executed processing steps, including noise reduction, sampling rate adjustment, pre-emphasis, framing, and slicing, and each step has a clear technical goal and implementation method.

[0023] In the actual operation of the system, the audio processing module first performs noise reduction on the input raw audio data. The noise reduction process can use classic signal processing methods such as spectral subtraction or wavelet denoising to estimate and filter the background noise. The system establishes a noise model through silent segment analysis, and then removes background interference from the entire audio signal, reducing the risk of offset caused by environmental noise to subsequent acoustic feature extraction.

[0024] After noise reduction, the system performs sampling rate adjustment on the audio data according to unified specifications. Preferably, if the sampling rate of the original audio is higher or lower than the standard value (for example, 16kHz), it will be upsampled or downsampled through a bandpass filter and a resampling algorithm to ensure that all input data has consistency in frequency domain characteristics. This processing not only helps improve the robustness of the system, but also keeps the time window processing that subsequent feature extraction relies on aligned.

[0025] After the sampling rate is adjusted, the audio data enters the pre-emphasis stage. This step mainly applies a linear filter function (such as ,in The value ranges from 0.95 to 0.97) to enhance the proportion of high-frequency components and suppress low-frequency noise interference, thereby improving the expression of high-frequency speech features such as voiceless consonants and plosives.

[0026] Subsequently, the system performs frame processing on the audio signal, dividing the continuous audio signal into short time frames of equal length. Preferably, each frame is 20 milliseconds long, and there is a 50% overlap between adjacent frames, that is, a 10 millisecond frame shift. This setting ensures that the speech signal has local stability within the time window, so that the extracted features such as fundamental frequency and formant have higher time accuracy and anti-jitter capability.

[0027] Finally, in order to meet the input requirements of various evaluation models, the audio processing module also slices the framed data. The purpose of slicing is to combine several consecutive frames into a set of time context segments, such as every 5 frames as an analysis unit, so as to support feature modeling and model input based on short-term semantic context. Slicing can also dynamically adjust the segment length according to the speech duration of different tasks, such as setting the short sentence pronunciation evaluation to 0.5 second segments and setting the long text interpretation evaluation to more than 1.5 seconds.

[0028] To achieve the above functions, the audio processing module can encapsulate the underlying processing logic based on the existing open source toolkit (such as PyDub), and gradually transmit the processing results through the data pipeline to ensure data consistency between each step. All output pre-processed audio data adopts a unified floating point PCM format, and is accompanied by timestamp and frame number information for feature alignment and model scheduling by the standard evaluation module 103 during the evaluation process.

[0029] The standard evaluation module 103 is used to determine the evaluation type according to the user's settings; receive the preprocessed audio data and perform evaluation operations corresponding to the evaluation type, and the evaluation operations include at least one of the following operations: Identify the speech units in the preprocessed audio data, calculate the accuracy rates of the initial consonants, final vowels, and connected speech sound change linguistic features, and generate a pronunciation evaluation result; Extract the acoustic features in the preprocessed audio data, including fundamental frequency, formant, sound intensity, fundamental frequency perturbation, and amplitude perturbation, and input the features into machine learning models separately trained for different evaluation items for classification prediction to generate a basic skills evaluation result; According to the user's emotion or scenario information, construct a prompt word template, and perform multi-modal fusion analysis on the preprocessed audio data and the scenario text, output the emotional matching degree, scene restoration degree, and emotional coherence index, and generate an emotion evaluation result.

[0030] The standard evaluation module 103 is the core analysis unit in this system that undertakes speech analysis, feature classification, and evaluation output. Functionally, it directly corresponds to the evaluation target set by the user. It is used to receive the preprocessed audio data output by the audio processing module 102, and execute different analysis paths according to the evaluation type, and finally generate a structured and operable evaluation result. This module supports three types of evaluation operations: pronunciation evaluation, sound basic skills evaluation, and emotion and scenario adaptability evaluation, and can realize multi-dimensional speech ability evaluation and personalized feedback.

[0031] In terms of pronunciation evaluation, the module performs speech unit segmentation and recognition on the input preprocessed audio data. By introducing a trained speech recognition model, it can automatically identify the initial consonants, final vowels, and whole syllable compositions in the user's speech, and further refine to the connected speech sound change phenomena within the syllable, such as neutral tone, rhotacization, and changes in modal particles. These models usually adopt a structure combining deep neural networks and time series modeling and can adapt to different pronunciation styles and speaking speeds. The system compares the recognized syllables with the standard pronunciations one by one, aligns the positions at the phoneme level, and determines whether the pronunciation is accurate. If a syllable deviates from the standard pronunciation, the system will locate the position of the pronunciation error and generate corresponding error prompts, such as "the tongue position is too high", "insufficient aspiration", or "unclear nasal sound at the end of the rhyme". This evaluation result not only includes the overall accuracy rate value, but also includes the error location, phoneme analysis, and oral suggestions for each segment of speech, which is convenient for users to make improvements item by item.

[0032] In terms of the evaluation of basic vocal skills, the module extracts core parameters reflecting vocal physiology and technical capabilities based on continuous frame-level audio feature extraction and in combination with acoustic signal processing algorithms. These parameters include, but are not limited to, fundamental frequency, formants, sound intensity, fundamental frequency perturbation, and amplitude perturbation. The system determines whether the vocalization is stable, whether the pitch is accurate, whether the speech rate is uniform, and whether the breath is coherent by dynamically tracking and analyzing the frequency changes of each frame and combining the information of the previous and subsequent frames. For example, if the amplitude fluctuations in a certain segment of speech are large and the periods are irregular, the system can identify it as unstable breath and prompt the user to strengthen breathing control exercises.

[0033] For each specific vocal dimension, the system has preset evaluation items and their corresponding high-correlation acoustic feature combinations. For example, the evaluation of voice brightness focuses on analyzing sound intensity and high-frequency energy ratio; oral resonance mainly relies on the spectral center frequency and formant shift; breath control is analyzed by combining sound intensity fluctuations and period stability. In the data preparation stage, the system constructed a sample set based on approximately 28,000 pieces of expert-annotated data, and through feature selection, cross-validation, and performance index testing, trained machine learning models with optimal performance for each evaluation dimension. These models include support vector machines, random forests, gradient boosting trees, multi-layer perceptrons, decision trees, and logistic regression, etc., and can all be dynamically scheduled through model configuration files. The system can intelligently select the most suitable model type for real-time classification prediction based on the discrimination ability of each model for positive and negative samples, the actual accuracy rate, and the operating efficiency, and output the score, scoring basis, and training suggestions for this dimension.

[0034] The execution mechanism of emotion and scenario evaluation operations is more complex. First, the system constructs a corresponding prompt word template from the emotion template library according to the target emotion selected by the user or the specified role and scenario information. The composition of the prompt word template not only includes emotion types (such as sadness, anger, surprise, pleading, etc.), but also can introduce role identity settings (such as age, gender, social identity) and scenario situations (such as formal meetings, family conversations, public speeches, etc.). For example, if the user selects "simulate an elderly person expressing apology", the system will generate a prompt word template containing features such as a low tone, a slow speech rate, and long tone words.

[0035] Next, the module inputs the prompt word template, preprocessed audio data, and reference emotion example audio into the multi-modal fusion analysis engine. This engine can synchronously process acoustic features (such as pitch change range, speech rate rhythm, tone intensity, pitch coherence) and semantic content (emotion keywords, tone structures, word choices, etc. in the text content transcribed by automatic speech recognition), and determine whether the actual expression matches the target emotion, role setting, and scenario requirements through joint modeling. The evaluation outputs include emotion matching degree (i.e., whether it has the specified emotion), scenario restoration degree (i.e., whether the expression fits the scenario setting), and emotion coherence (i.e., whether the emotion expression in the speech is consistent and natural).

[0036] The multimodal fusion analysis engine is used to jointly model and quantitatively evaluate the emotional matching degree, scene restoration degree, and emotional coherence of the user's voice samples.

[0037] The input of this engine consists of three parts, namely: the frame-level acoustic feature sequence of the preprocessed audio samples of the user, the time-aligned semantic coding stream converted from the structured prompt word template, and the frame-level acoustic comparison features of the reference demonstration audio. The first part of the input comes from the audio processing module, and its feature dimensions include MFCC (Mel Frequency Cepstral Coefficients), pitch, energy, formant positions, pitch curvature (such as delta-pitch), speech rate boundaries, and rhythm changes. All frame features are organized in the form of a two-dimensional tensor, and the tensor dimension is , where T is the number of audio frames and F is the number of feature dimensions per frame. The second part of the input embeds the prompt word template through a syntax-semantic embedding model (such as BERT) to obtain a set of structured semantic coding streams. Time position encoding is added to each semantic segment vector to keep it frame-level aligned with the audio frames, generating a semantic tensor with dimensions of T × E, where E is the embedding dimension. The third part is the reference demonstration audio tensor with the same structure as the preprocessed audio, which is used to provide a comparison benchmark for the target emotional expression.

[0038] The engine itself adopts a fusion structure based on a dual-channel attention mechanism. One channel of the input is the acoustic feature sequence, and the other channel is the semantic coding stream. In the core cross-attention layer, the system constructs multiple attention heads (the recommended number is 4), which respectively perform feature mapping and weight calculation on the acoustic frame features and semantic segments. Each attention head adopts the standard dot-product attention method to obtain the mutual attention weight matrix by calculating the similarity between the audio frames and the corresponding semantic vectors. To avoid relying on large models, this attention calculation process can be linearly transformed and normalized through preset projection matrices (such as W_Q, W_K, W_V), enabling the system to run on resource-constrained platforms.

[0039] The preset projection matrix is a linear transformation weight matrix used to map the input acoustic feature vector or semantic encoding vector to a unified attention space, representing the transformation matrices of Query, Key, and Value respectively. For example, W_Q is used to map the acoustic input vector to a query vector, and W_K and W_V are used to map the semantic input vector to a key vector and a value vector respectively. Each projection matrix is a two-dimensional floating-point array, and its dimensions are set according to the input and output dimensions (for example, the input dimension is F or E, and the output is the unified dimension D), and it is optimized and learned during the training phase after random initialization or import from a pre-trained model. To ensure feasibility under the minimum implementation conditions, these matrices can also be fixed as small-scale parameter matrices set manually for the preliminary implementation of the attention mapping ability.

[0040] After completing the cross-attention, the system forms a set of fused feature representations by concatenating the weight matrix and the original features, and then inputs them into a fully connected neural network with a two-layer structure. The number of nodes in each layer of this network is 128 and 64 respectively, using the ReLU activation function, and having basic non-linear expression ability. The final output is three independent scalars, corresponding to the emotional matching degree, scene restoration degree, and emotional coherence respectively. The value range of each scalar is from 0 to 1, which can represent the degree of consistency with the target setting and is used to drive the subsequent structured evaluation feedback.

[0041] Through the collaborative work of the above structure, the multi-modal fusion analysis engine can simultaneously capture the acoustic change features in the speech and the emotional expression directions in the prompt word semantics, extract the dynamic coupling relationship between them, so as to realize the in-depth modeling of expression consistency, semantic matching, and emotional flow.

[0042] When the evaluation task is "role-playing", the system will also perform voice feature matching analysis on the role feature parameters input by the user. For example, older characters usually have a lower pitch, a slower speaking speed, and fewer intonation fluctuations; while lively child-like characters have obvious intonation jumps and large fluctuations in sound intensity. The system will evaluate the matching degree between the relevant indicators in the user's speech and the target features, and generate a role adaptation score for judging whether the voice performance is successful.

[0043] The standard evaluation module 103 also has an intelligent adaptive mechanism, which can dynamically adjust the threshold parameters, prompt word generation logic, and feature weight strategy of each model during the execution process by combining the user's historical evaluation records, so as to adapt to individual vocal differences and optimize the stability and personalization of the evaluation results.

[0044] The evaluation results are output in a unified format, including the dimension scores, score explanations, error location, improvement suggestions, and priority training directions of each evaluation type, for the integration processing module 104 to perform integration, visual display, and personalized training plan push.

[0045] To further improve the accuracy and pertinence of the evaluation model, the standard evaluation module 103 introduces a systematic feature correlation analysis and algorithm matching mechanism in the feature processing and model training processes. Specifically, for the evaluation of basic voice skills, the system first conducts a correlation matrix analysis between acoustic features and each evaluation dimension based on a large amount of user data. Through this analysis, key feature groups under different dimensions can be clearly identified. For example, in the evaluation of voice brightness and dullness, sound intensity, fundamental frequency, harmonic noise ratio, and frequency band energy ratio are identified as core features with significant impacts. For the oral resonance dimension, the center frequency of the spectrum and energy distribution are found to be highly correlated with the evaluation results. Again, the evaluation of breathiness and breath stability highly depends on two parameters reflecting periodic perturbations, namely fundamental frequency perturbation and amplitude perturbation; thoracic resonance is jointly affected by the center frequency of the spectrum, sound intensity, amplitude perturbation, and fundamental frequency; dimensions such as forte, piano, and strong and weak breath control jointly depend on features such as sound intensity, fundamental frequency, and the position change of the first resonance peak.

[0046] The system uses the above analysis results to construct a feature engineering strategy, that is, selects the strongly correlated feature combinations for each evaluation item and constructs feature vectors accordingly. In the model training stage, the system conducts multiple rounds of iterative training and evaluation for each dimension based on more than 28,000 evaluation sample data annotated by professional voice teachers. During the training process, mainstream machine learning algorithm models such as support vector machine, random forest, gradient boosting tree, decision tree, logistic regression, and multi-layer perceptron are introduced respectively, and key performance indicators such as accuracy, recall rate, F1 score, and area under the curve (AUC) of each model are comprehensively evaluated through cross-validation. Finally, the system automatically selects the algorithm model with the best performance and applies it to each evaluation dimension.

[0047] The results show that different evaluation items are indeed suitable for different algorithm models, and there are significant performance differences. For example, evaluation items such as breathiness, strong breath control, weak breath control, and voice dullness perform best under the support vector machine model, with accuracies reaching 82.45%, 88.62%, 88.33%, and 91.13% respectively; the evaluation tasks of forte and thoracic resonance are more suitable for the random forest algorithm, with corresponding accuracies reaching 88.89% and 91% respectively; the evaluation accuracy of oral resonance reaches 83% in the gradient boosting algorithm model; piano is supported by the decision tree model with an accuracy of 82.64%; while the voice brightness performs most prominently in the logistic regression model with an accuracy of 94.35%; the accuracy of the evaluation of full voice using the multi-layer perceptron neural network model is 91.14%.

[0048] In the part of emotion and scenario evaluation, the system also summarizes the best implementation path through a large number of experiments. The emotion evaluation items are first divided into three categories: single emotion, scenario deduction, and role-playing. The single emotion evaluation covers specific emotion types such as emotional fullness, happiness, anger, anxiety, coldness, fierceness, sadness, fear, love, secret joy, ecstasy, interrogation, calmness, humility, and pleading. For the recognition of emotional fullness, the system directly uses acoustic features for modeling and adopts the logistic regression algorithm to achieve binary classification prediction, with an accuracy rate of 84.47%. In other single emotion evaluations, the system combines prompt word templates to guide expressions, uses a multi-modal fusion model to input reference audio and target speech, and after preprocessing and template optimization, the emotion matching accuracy rate is stable between 88% and 95%.

[0049] The core of the scenario deduction evaluation is to evaluate whether the speech expressed by the user matches the specified scenario and target emotion. The system automatically generates structured prompt content, such as "express excitement in a speech" and "simulate nervous reading in an exam room", through a dynamic prompt word generation mechanism, combined with the scenario label and target emotion input by the user, so as to guide the user to express. Subsequently, the system evaluates the performance of the speech in three dimensions: emotion matching degree, scenario restoration degree, and emotion coherence, and outputs a comprehensive score. After actual measurement, the comprehensive judgment accuracy rate of this module reaches 87%.

[0050] The role-playing module further introduces role characteristic parameters such as age and personality on the basis of scenario deduction. For example, for tasks such as "simulate a young child expressing happiness" or "act as a steady leader expressing worry", the system integrates the scenario characteristics and role characteristics in the prompt word template to generate a composite template to guide the user to express the corresponding speech characteristics. Subsequently, the system conducts a coupling matching analysis of the user's speech and the target role through indicators such as age characteristics such as speech rate, pitch, and vowel length changes, and personality characteristics such as tone intensity and intonation jump amplitude, and generates a role adaptation degree score. In multiple scenario tests, the comprehensive accuracy rate of this module reaches the range of 85% to 90%.

[0051] Through the above modeling, training, and verification strategies, the standard evaluation module 103 not only achieves high-accuracy and multi-dimensional speech ability evaluation, but also establishes a complete, closed-loop, and extensible evaluation system based on scientific feature selection, classification algorithms, and template scheduling strategies, laying a solid foundation for the engineering implementation of this system in the fields of speech teaching, expression training, and intelligent speech evaluation.

[0052] In the sound normalization evaluation system for multi-dimensional acoustic parameter and emotion scenario dynamic analysis according to the present invention, when the standard evaluation module performs the basic sound evaluation operation, in order to improve the evaluation accuracy and pertinence of each preset evaluation item (such as strong voice, breathy voice, chest resonance, etc.), a multi-stage and highly controllable acoustic feature selection mechanism is designed. This mechanism consists of three closely connected sub-steps, ensuring that the features finally input into the classification model have clear discriminative power and a low interference risk, thereby effectively improving the system's fine-grained modeling ability for different-dimensional speech skills.

[0053] First, for each preset evaluation item, the system extracts the user's historical voice samples from the evaluation database according to the scoring dimension corresponding to the evaluation item, and groups and pairs them with the teacher samples annotated by professional voice teaching staff to construct a control sample set covering each scoring level. For example, for the "breath stability" evaluation item, the system groups the samples that accurately express stable breath into one group, and the samples with manifestations such as breath jitter and weakness into another group. Subsequently, among these control group samples, the system extracts multiple acoustic parameters that are significantly related to vocal stability, spectral pattern changes, and energy distribution patterns through frame-level voice signal analysis methods. These parameters include fundamental frequency perturbation (used to measure the fluctuation degree of the vocal cord vibration period), amplitude perturbation (reflecting the instability of sound energy output), spectral density centroid position (representing the resonance focus position), and non-steady state energy fluctuation amplitude (describing the short-term speech intensity change). These features together constitute an acoustic parameter set, providing a data basis for subsequent differential modeling.

[0054] Next, the system performs between-group difference enhancement processing on the above acoustic parameter set. Instead of using a simple statistical difference method, this process introduces a shallow convolutional structure constructed by learnable weight kernels to model the parameter change trends between different level groups. Through training, this convolutional structure can adaptively capture and enhance the feature gradients related to the performance level, and finally form a multi-dimensional feature response map. In this map, the system can accurately mark the feature indicators that are the most sensitive to changes and have the clearest trends in high and low-level performances. These feature change patterns constitute the "feature-sensitive subspace" of the corresponding evaluation item. This subspace has a high directionality in the parameter dimension, ensuring that subsequent model training only focuses on the feature channels that truly have discriminative significance.

[0055] Finally, within the obtained feature-sensitive subspace, the system further executes a feature conflict suppression strategy to avoid interference or ambiguous determination caused by overlapping features used in multiple evaluation tasks. The system introduces an "overlap suppression factor" which judges the interference potential by calculating the occurrence frequency, sharing intensity, and feature direction consistency of the current candidate feature in other evaluation items. For features with a high overlap degree, the system reduces their feature weights; for features with uniqueness and outstanding performance, the system increases their proportion in the final feature combination. This strategy ensures that the finally selected target feature group not only has good discrimination but also has the least impact on the prediction process of non-target evaluation items, enhancing the specificity and anti-interference ability of the model.

[0056] The above final feature group is input into the trained classification model corresponding to the evaluation item. The model can be structured as a support vector machine, random forest, gradient boosting, or neural network, etc., and is used to predict and score newly input speech samples and compare them with the evaluation grade standard, thereby generating accurate basic vocal ability evaluation results. Through the above three-stage structured feature selection process, the present invention realizes a speech feature extraction and modeling mechanism deeply bound to the evaluation target, significantly improving the accuracy, stability, and adaptability of the system during parallel evaluation of multiple tasks.

[0057] During the implementation of the present invention, in order to extract acoustic parameters closely related to the evaluation of basic vocal skills, the system first pre-emphasizes each speech sample to enhance high-frequency components and suppress low-frequency noise interference. Subsequently, the speech signal is divided into multiple overlapping short-time frames, each frame being about twenty milliseconds long with an interval of ten milliseconds between adjacent frames to capture the minute changes of the speech signal over time.

[0058] For each frame of audio data, the system extracts multiple acoustic features, including key parameters such as fundamental frequency perturbation, amplitude perturbation, spectral density centroid position, and non-steady energy fluctuation amplitude. Fundamental frequency perturbation is used to measure the stability of the vocal cycle. The system first locates the main peak position of the fundamental frequency cycle using the autocorrelation function of the short-time frame, and then calculates the ratio of the standard deviation of the adjacent cycle durations to the average cycle to obtain the fundamental frequency fluctuation degree. Amplitude perturbation reflects the stability of energy output during vocalization. The system quantifies its jitter amplitude by calculating the relative change of the maximum amplitude of adjacent cycles. The spectral density centroid position is obtained by performing a fast Fourier transform on each frame to obtain a frequency distribution map, and then using a weighted average calculation method of the product of frequency and the corresponding energy amplitude to extract the spectral energy concentration position, thereby inferring the resonance focus during vocalization. The non-steady energy fluctuation amplitude evaluates the short-time energy fluctuation level by statistically analyzing the energy change degree between multiple consecutive frames, and outputs the fluctuation variance obtained after normalization as a parameter.

[0059] Through the above-mentioned feature extraction process, the system can obtain a set of parameters that highly represent the stability and resonance structure of the speech signal, providing a reliable input basis for subsequent differential analysis based on scoring levels, feature trajectory modeling, and multi-task interference control. The entire process can be automatically completed within the standard evaluation module, ensuring the stability, timeliness, and reproducibility of feature extraction, making the system have clear physical interpretability and engineering feasibility in the automated and personalized evaluation of voice skills.

[0060] In a preferred embodiment of the present invention, when the standard evaluation module performs the basic voice skill evaluation operation, for each preset evaluation item (such as strong voice, breathy voice, resonance-related indicators), the acoustic feature trajectory evolving with the change of the scoring level is extracted by using a scoring-level differential encoder. Instead of using traditional static mean or variance features as input, this module constructs a control sample set covering each scoring level, and based on the difference relationship between level embedding and sample features, guides a learnable differential neural network structure to generate a feature response map. Through this structure, a highly discriminative trajectory feature vector can be obtained, effectively mapping the performance trend of the user sample in the target voice skill dimension and supporting the accuracy requirements of the evaluation model in the micro-difference expression and grading determination scenarios.

[0061] Furthermore, to improve the robustness during the parallel execution of multiple evaluation tasks, the present invention designs a perturbation consistency resonance filtering mechanism. This mechanism evaluates the resonance degree triggered by the candidate features of the current evaluation item in other evaluation tasks by introducing a linear kernel structure based on a shared perturbation channel, and calculates the interference risk of the features. The system dynamically adjusts the weights of the features in the final sensitive subspace according to the perturbation common risk, and preferentially retains the feature dimensions that are unique, highly discriminative in this evaluation item and are not easily shared and interfered by other tasks. Compared with the existing technologies based on methods such as information gain and variance screening, this mechanism significantly enhances the practical effects of the present system in task isolation, multi-dimensional modeling, and personalized training recommendation. Please refer to the following implementation code: import numpy as np import torch import torch.nn as nn import torch.nn.functional as F # Build a scoring-level trajectory differential encoder class GradedContrastEncoder(nn.Module): # Scoring-level differential encoder: Simulate the "evolution trajectory" of features when the user progresses from a low score to a high score, different from traditional static feature extraction, emphasizing score-level differences rather than means""" def __init__(self, feature_dim): super(GradedContrastEncoder, self).__init__() self.delta_fc = nn.Linear(feature_dim, feature_dim) self.score_embed = nn.Parameter(torch.randn(5, feature_dim))# Assume 5 - level labels def forward(self, features, scores): #features: [N, D], input feature vectors #scores: [N], the scoring labels (0 - 4) corresponding to each sample # Get the embedding of the corresponding scoring level score_vecs = self.score_embed[scores] # Differential representation: feature - level vector delta = features - score_vecs out = F.relu(self.delta_fc(delta)) return out # As the feature evolution response # Perturbation Consistency Resonance Filter class PerturbationConsistencySuppressor(nn.Module): Multi - task resonance risk assessment: Avoid the "resonance" interference of selected features in other evaluation tasks, based on perturbation commonality weights rather than simple cosine similarity def __init__(self, feature_dim): super(PerturbationConsistencySuppressor, self).__init__() self.shared_filter = nn.Linear(feature_dim, 1, bias=False) def forward(self, target_feature, other_features): #target_feature: [D] #other_features: [T, D], where T is the feature center of other tasks # Perturbation intensity: Look at the response of each task to the shared perturbation channels shared_risks = self.shared_filter(other_features) # [T, 1] target_risk = self.shared_filter(target_feature.unsqueeze(0))# [1, 1] # Calculate the "resonance degree" with the target perturbation resonance_score = torch.mean(torch.sigmoid(shared_risks - target_risk)) # Suppression factor: The stronger the resonance → the greater the suppression suppression_weight = torch.clamp(1 - resonance_score, min=0.2) return target_feature suppression_weight # Main process function: Integrate multiple modules to construct the final feature vector def construct_inventive_feature_vector(feature_tensor, score_tensor, other_tasks_center_features): # feature_tensor: Basic features of all samples [N, D] #score_tensor: Rating level of each sample [N] #other_tasks_center_features: Central features of other evaluation tasks [T, D] # Stage 1: Rating level trajectory encoding encoder = GradedContrastEncoder(feature_dim=feature_tensor.shape[1]) trajectory_map = encoder(feature_tensor, score_tensor) # [N, D] # Stage 2: Aggregate feature responses of different levels to obtain the sensitive vector of the target evaluation item trajectory_mean = torch.mean(trajectory_map, dim=0) # [D] # Stage 3: The resonance filter suppresses the overlapping parts with other tasks suppressor = PerturbationConsistencySuppressor(feature_dim=feature_tensor.shape[1]) optimized_vector = suppressor(trajectory_mean, other_tasks_center_features) # [D] return optimized_vector # The final output can be fed into the exclusive classification model # Construct test data for simulation run N = 10 # Number of samples D = 8 # Feature dimension T = 3 # Number of other tasks torch.manual_seed(42) features = torch.randn(N, D) scores = torch.randint(0, 5, (N,)) # Simulated scoring labels other_task_feats = torch.randn(T, D) # Other task center features # Execute the process final_feat = construct_inventive_feature_vector(features, scores,other_task_feats) # Output the final feature vector final_feat Further, when performing the basic voice evaluation operation, the standard evaluation module includes an acoustic channel weighted adjustment strategy based on feature interference suppression and discrimination improvement. The acoustic channel weighted adjustment strategy includes the following steps: In the feature-sensitive subspace, for each dominant acoustic channel, construct a channel response scoring function , which is used to characterize its performance discrimination ability under the target evaluation item, and is defined according to the following formula 1: ; Wherein, represents the mean value of channel in the high-scoring sample group; high-scoring samples refer to voice samples with an expert scoring level higher than the preset threshold; represents the mean value of channel in the low-scoring sample group; low-scoring samples refer to voice samples with an expert scoring level lower than the preset threshold; , respectively represent the standard deviations of channel in the high and low scoring groups; is a very small positive number introduced to prevent the denominator from being zero; Based on the overlap situation among candidate channels in different evaluation items, introduce a channel exclusivity index , to control the channel interference degree, and its calculation method adopts the following formula 2: ; Wherein, represents whether channel appears in the target feature group of evaluation item (appears as 1, does not appear as 0); is the index of the current evaluation item, is the total number of evaluation items; the calculation formula expresses the sharing ratio of the current channel in all evaluation tasks except the target evaluation item, so as to deduce its exclusivity degree inversely. The closer the index is to 1, the fewer times the channel appears in other tasks and the smaller the interference possibility. In this system, the evaluation item refers to different ability dimensions used to evaluate the user's voice performance. For example, a common evaluation item can be breath stability, which is used to judge whether the user maintains uniform breathing when speaking, and whether the voice trembles or breaks. Another evaluation item may be chest resonance, which is used to evaluate whether the user's voice has depth and penetration. There may also be pitch control ability, which is used to check whether the high and low pitch changes are natural during speaking, and speech rate rhythm control ability, which measures whether the language rhythm is appropriate and whether there are inappropriate fast or slow changes. In addition, it may also include whether the emotional expression fits the specific situation, such as whether the voice is low and the speech rate slows down when expressing "sadness", or whether there are features such as volume increase and intonation rise when expressing "anger".

[0062] Define the channel weighting coefficient according to the following formula 3 as: ; wherein, and are hyperparameters used to adjust the weight balance between discriminability and exclusivity. The recommended value is , and this parameter ratio can be adjusted and optimized in the experiment to adapt to the vocal behavior characteristics of different user groups.

[0063] Through the above weighting strategy, the system can optimize the weight of the dominant acoustic channel during the training phase, making the input features of the final classification model more focused on the channels with high discriminability and low interference, thereby improving the model evaluation accuracy and task adaptability.

[0064] For example, in the task of evaluating the "chest resonance" ability, after analyzing multiple candidate acoustic channels, the system finds that the change in the spectral centroid position is significantly different between high-scoring and low-scoring users, and this channel is used less frequently in other evaluation tasks, with strong discriminability and exclusivity. At the same time, another channel such as the sound intensity parameter can also reflect the vocal difference, but it appears frequently in multiple tasks such as "loud voice" and "breath control", which is likely to introduce cross-interference.

[0065] Under the action of the weighting strategy, the system will assign a higher weight to the spectral centroid channel and a relatively lower weight to the sound intensity channel. Subsequently, during the training process, the classification model will pay more attention to the information change of the former and give priority to using its pattern for discrimination and determination during modeling. Finally, when the model conducts an automatic evaluation of chest resonance, it shows stronger specificity and recognition accuracy, significantly reducing the risk of misjudgment caused by interference from other task features.

[0066] In the implementation of the present invention, in order to achieve the dynamic binding of the optimal classification model for each preset sound basic skill evaluation item, the standard evaluation module designs a three-stage execution process combining acoustic channel analysis, model structure response adaptation, and index-driven screening, ensuring that the classification model can highly fit the performance characteristics of different evaluation tasks and improving the overall evaluation accuracy and model stability.

[0067] First, the system constructs a set of reference samples for rating levels based on user samples and labeled samples covering all rating levels of the assessment item in the assessment database. On this basis, through frame-level feature extraction and channel response modeling, the system calculates the response intensity and distribution trend of each acoustic feature channel (such as fundamental frequency, formant, spectral centroid, energy change, etc.) in different rating levels. This process can calculate metrics such as frame mean, variance, time gradient, and response direction consistency after channel normalization to obtain a set of quantified "response gradient maps". The system marks the channels with the most significant changes and the clearest discrimination trends among the rating levels in this map as the "dominant acoustic channel set", which represents the most representative acoustic performance path in this assessment item.

[0068] Before binding the model structure for the evaluation of vocal basics, the system first constructs a set of reference samples for rating levels for each assessment item. This set consists of speech samples covering all rating levels of the assessment item and is divided into multiple level subgroups according to the teacher annotation information of each sample, such as "excellent", "good", "average", and "to be improved", etc. The system performs short-time frame division on each sample, with common parameters of 20 milliseconds per frame length and 10 milliseconds frame shift, and extracts acoustic feature channels with fixed dimensions on each frame. The channels preset by the system include but are not limited to the following: fundamental frequency (used to describe the vocal cord vibration period), formant position (used to reflect the shape of the oral cavity and pharynx), spectral density centroid (representing the location of sound energy concentration), energy intensity (reflecting the sound pressure change), fundamental frequency perturbation (measuring frequency fluctuation), amplitude perturbation (measuring energy fluctuation), and harmonic-to-noise ratio (reflecting the clarity and hoarseness of the sound).

[0069] To ensure the consistency of feature scales among samples in subsequent analysis, the system normalizes each channel feature of all samples within each rating level group. The specific operation is to calculate the mean and standard deviation of all frame values of each channel within each rating level group, and then subtract the mean from each frame value of this channel and divide by the standard deviation to obtain a normalized channel matrix. This processing ensures that the feature change trends among different level groups are more prominent.

[0070] Subsequently, based on the standardized results, the system calculates the characteristic changes of each acoustic channel between different scoring grade groups for generating the "channel response gradient map". The calculation of the channel response gradient map includes four types of metrics. The first type is the normalized mean difference value, that is, for each channel in all grade groups, the difference in the mean values between the current group and the next group of this channel is calculated in sequence, and the maximum difference is recorded as the average response difference metric of this channel. The second type is the direction consistency metric, that is, to check whether the direction of the mean value change of this channel is consistent among all scoring grades (always increasing or decreasing), if consistent, it is marked as "direction stable", otherwise marked as "direction inconsistent". The third type is the channel response stability metric, that is, the standard deviation of all frame values within each grade group of this channel is calculated respectively, and the standard deviations of the low-grade group and the high-grade group are compared to judge the stability degree of this channel in high-quality vocalization. The fourth type is the inter-frame change gradient metric, that is, for each channel, the adjacent difference sequence of all frame values is calculated, and its mean value is obtained in each grade group, and then this average inter-frame difference is compared between different groups to evaluate the vocalization smoothness or control ability of this channel.

[0071] Each acoustic channel will obtain a set of quantization scores in the above four types of metrics. The system assigns unified weight parameters to each channel. For example, the mean difference accounts for 40%, the direction consistency accounts for 20%, the response stability accounts for 20%, and the inter-frame change gradient accounts for 20%. These proportions can be adjusted according to the preferences of the evaluation task. The system aggregates the weighted scores into a total score and sorts all channels in descending order according to the total score.

[0072] After the sorting is completed, the system extracts the "dominant acoustic channel set" according to the set channel screening rules. The screening rules can be: select the top 30% scoring channels, or select all channels with scores higher than the average score of all channels plus one standard deviation. The channels included in this set are considered to be the most discriminative acoustic feature channels in the current evaluation item and will be used for subsequent model structure adaptation analysis and the final evaluation process.

[0073] After obtaining the set of dominant acoustic channels, the system enters the model structure adaptability evaluation stage. For multiple pre-set candidate model structures, including support vector machines, random forests, gradient boosting algorithms, multi-layer perceptrons, logistic regression, and decision trees, etc., the system analyzes the response stability of these models on the dominant acoustic channels one by one. This evaluation process includes but is not limited to the following criteria: the smoothness of the channel response surface of the model at different scoring levels, the sensitivity test to feature perturbations (such as the magnitude of prediction offset after adding small noise), and whether the distribution of the dependence strength of the model on the dominant channels shows highly non-linear fitting characteristics. For model types that show excessive sensitivity to channel drift, weak response gradient antagonism, or inconsistent feature directionality in the evaluation, the system removes them from the candidate set to obtain the structure adaptability screening results.

[0074] To ensure the objectivity and reproducibility of the above evaluation, the system adopts the following operation process in actual execution for the quantitative analysis of response stability. First, the system divides the scoring level control sample set into multiple subsets, each subset containing the proportion distribution of samples at high, medium, and low levels. For each candidate model, the system trains and validates the performance of the model on the dominant acoustic channels on each subset respectively, and records whether the continuity of the model prediction output with the change of sample levels and the discrimination interval remain stable. If there is an obvious shift in the prediction boundary between different subsets, or a discontinuous distribution between high and low levels, it indicates that the response surface stability of the model on this channel set is poor.

[0075] Furthermore, the system conducts a perturbation experiment on each dominant channel within each subset, that is, on the basis of the original channel features, a small perturbation is introduced channel by channel, for example, adding or subtracting a deviation amplitude not exceeding 1% to each channel value, and observing the variation range of the model output results. If the prediction results of the model change violently after adding a very small perturbation, then the model is too sensitive to channel noise and is not suitable for binding to this evaluation item. In addition, the system also calculates the feature importance or weight distribution of each dominant channel within each model. If there is a situation where the weight of a certain channel is extremely high while the weights of other channels are close to zero, and the high-weight channel is vulnerable to external interference, it indicates that the model may have non-linear overfitting and does not have structural stability.

[0076] Through the above analysis path, the system finally removes the models with unstable response structures and retains the model types that show stable performance, strong anti-interference ability, and consistent direction trends on the dominant channel set as the source of the structure adaptability screening results in the subsequent scoring process.

[0077] For example, when binding an evaluation model for the preset evaluation item of "breath stability", the system first constructs three subsets according to the scoring levels with reference to the sample set, corresponding to the samples of the "stable", "moderately stable", and "unstable" levels marked by teachers, and the number of samples in each group is kept consistent. The dominant acoustic channel set extracted by the system includes three feature channels: fundamental frequency perturbation, amplitude perturbation, and short-term energy fluctuation.

[0078] For the support vector machine (SVM), random forest (RF), multi-layer perceptron (MLP), and logistic regression (LR) in the candidate models, the system sequentially trains and cross-validates each model on these three subsets. Taking SVM as an example, after the system trains on each subset, it compares the trend curve of its prediction probability with the change of the scoring level. If the output probability distribution of SVM on the high-level samples highly overlaps with that on the medium-level samples, and the prediction boundary is concentrated on the edge of the samples, it indicates that there is a "flat area" or "local jump" in its channel response surface, and the stability is poor.

[0079] Subsequently, the system applies perturbation tests to each dominant channel. For example, all frame values in the fundamental frequency perturbation channel are increased by 0.5%, and the original samples are input into each model again, and the difference in prediction scores before and after the perturbation is recorded. If the score change of the RF model on most samples is less than 5%, while the change of the MLP model exceeds 15%, it can be inferred that the RF model has better anti-perturbation stability for this channel. The system also analyzes the channel dependence strength of each model during the training process and finds that the LR model almost only depends on the amplitude perturbation channel, and the weight coefficients of the other channels are close to zero, indicating that this model may be sensitive to feature deviations and have insufficient applicability in actual applications.

[0080] The system eliminates the MLP and LR models, and only retains the SVM and RF for the final evaluation. By uniformly scoring the scoring accuracy, F1 score, and AUC, it is determined that the RF model has the best performance in the "breath stability" evaluation item, and it is used as the final bound model for the scoring prediction of subsequent user samples.

[0081] Finally, the system deploys the model structures retained in the structural adaptability screening results on the historical sample set of this evaluation item for cross-validation to construct multiple rounds of evaluation experiments. The evaluation results of each model will include multiple performance indicators such as accuracy, recall rate, F1 score, and area under the curve (AUC). The system weights and scores these indicators to generate a unified performance evaluation score. The model with the highest final score is regarded as the optimal adapted structure for the current evaluation item and is bound to this evaluation item for automatic scoring and grade determination of new user voice samples in the actual evaluation process.

[0082] Through the above dynamic model selection mechanism, the standard evaluation module can not only achieve a deep adaptation between the evaluation task and the model structure, but also ensure stable model performance, strong anti-interference ability, and consistent channel response. It effectively avoids the problems of insufficient model generalization and task deviation caused by static performance optimization in traditional model selection methods, thus significantly enhancing the accuracy, professionalism, and flexibility of the system in multiple voice skill evaluations.

[0083] During the process of performing emotional or situational evaluation operations, the standard evaluation module will first call the scenario modeling component to generate a structured prompt word template that matches the user's settings. This template not only includes semantic descriptions of target emotions (such as anger, sadness, pleading, etc.), but also includes scene types (such as public speaking, role-playing, family communication) and user-defined role characteristics (such as age, personality, identity, etc.), and internally encodes rhythm control markers to ensure consistent time alignment references for subsequent fusion analysis. The system converts this structured prompt word template into a sequence of semantic guiding words, and then calls the segment semantic embedding mechanism to split this sequence of semantic guiding words into several segments. Each segment is embedded into a vector representation and bound to the time dimension, finally forming a time-aligned semantic coding stream.

[0084] After obtaining the semantic coding stream, the system synchronously inputs it, the preprocessed audio data, and the reference demonstration audio into the multimodal fusion engine. The core of the fusion engine is a computational framework built based on the cross-modal attention mechanism. In this engine, the system extracts the acoustic features of each frame in the preprocessed audio data, including but not limited to pitch, energy, pitch curvature, speech rate boundaries, and rhythm changes; at the same time, the reference demonstration audio is also processed into an equal-length frame structure to form a comparison baseline. Subsequently, the system establishes a correspondence between the two modalities by constructing a mutual attention weight matrix between the acoustic frame features and the semantic embedding, thereby extracting the correlation degree between pitch and semantic keywords, the consistency between speech rate rhythm and syntactic boundaries, and the dynamic coordination degree between pronunciation strength and emotional instructions. During this process, the system will automatically identify which acoustic change patterns have a stable linkage relationship with semantic expressions and construct a two-dimensional expression consistency map for visualizing the structural linkage structure of emotional expressions.

[0085] Based on this spectrogram, the system further extracts three types of intermediate metrics. The first type is the local inconsistency score, which refers to the acoustic expression deviation near some key semantic units, such as the contradictory area where the rhythm speeds up but the semantic representation is calm; the second type is the emotion channel missing marker, which is used to indicate the phenomenon that the lack of a specific acoustic feature channel (such as the lack of intensity change or the lack of pitch fluctuation) leads to insufficient emotion expression; the third type is the semantic drift degree, which evaluates whether the speech performance deviates from the intended direction set in the original prompt word template. For example, instead of presenting a pleading tone, it shows a threatening tendency. The system jointly models the above intermediate metrics with the global context parameters in the scenario setting, adopts a feature fusion and score aggregation strategy, and finally calculates three quantitative metrics in terms of emotion matching degree (i.e., whether the target emotion is accurately expressed), scenario restoration degree (i.e., whether the voice expression conforms to the set context), and emotion coherence (i.e., whether the emotion expression in the front and back sentences transitions naturally), and returns the evaluation results as a structured score feedback to the integration processing module for subsequent generation of the evaluation report.

[0086] The above process is automatically triggered every time the user submits an emotion or scenario-based voice sample, without relying on additional processing capabilities of the user terminal, ensuring the real-time and accuracy of the evaluation process, and enabling personalized expansion of multiple scenarios and emotions through configuration templates and reference demonstration content. The following is the reference implementation code for the aforementioned emotion and scenario multimodal evaluation process: import torch import torch.nn as nn import torch.nn.functional as F # Assume the maximum time step is T, the audio frame feature dimension is D_audio, and the semantic embedding dimension is D_text T = 100 D_audio = 64 D_text = 128 class SemanticEmbeddingEncoder(nn.Module): """Fragment semantic embedding module: Fragment the structured prompt word template and convert it into a time-aligned embedding representation""" def __init__(self, vocab_size, embedding_dim, hidden_dim): super(SemanticEmbeddingEncoder, self).__init__() self.embedding = nn.Embedding(vocab_size, embedding_dim) self.gru = nn.GRU(embedding_dim, hidden_dim, batch_first=True,bidirectional=True) def forward(self, template_seq): # Input is the index sequence after semantic template segmentation [B, L] embedded = self.embedding(template_seq) # [B, L, embedding_dim] outputs, _ = self.gru(embedded) # [B, L, 2 hidden_dim] return outputs class AcousticFeatureExtractor(nn.Module): """ Acoustic frame feature extractor: extracts features such as pitch, intensity, speech rate boundary, etc. from each frame """ def __init__(self, input_dim, output_dim): super(AcousticFeatureExtractor, self).__init__() self.linear = nn.Linear(input_dim, output_dim) def forward(self, acoustic_input): # Assume acoustic_input is the preliminary MFCC and other splicing features of each frame [B, T, input_dim] return self.linear(acoustic_input) # [B, T, output_dim] class CrossModalAttention(nn.Module): """ Cross-modal Attention Mechanism Module: Realize cross-attention between acoustic and semantic features """ def __init__(self, dim_audio, dim_text, hidden_dim): super(CrossModalAttention, self).__init__() self.query_proj = nn.Linear(dim_audio, hidden_dim) self.key_proj = nn.Linear(dim_text, hidden_dim) self.value_proj = nn.Linear(dim_text, hidden_dim) def forward(self, audio_feat, text_feat): # audio_feat: [B, T, dim_audio], text_feat: [B, L, dim_text] Q = self.query_proj(audio_feat) # [B, T, H] K = self.key_proj(text_feat) # [B, L, H] V = self.value_proj(text_feat) # [B, L, H] attention_scores = torch.matmul(Q, K.transpose(1, 2)) / (Q.size(-1) 0.5) # [B, T, L] attention_weights = F.softmax(attention_scores, dim=-1) # [B,T, L] context = torch.matmul(attention_weights, V) # [B, T, H] return context, attention_weights class MultimodalFusionEngine(nn.Module): """Multimodal Fusion Engine: Integrate Acoustic Features, Semantic Encoding, and Reference Audio to Form a Consistency Map""" def __init__(self, dim_audio, dim_text, fusion_dim): super(MultimodalFusionEngine, self).__init__() self.attention = CrossModalAttention(dim_audio, dim_text,fusion_dim) self.fusion_layer = nn.Linear(dim_audio + fusion_dim, fusion_dim) self.out_layer = nn.Linear(fusion_dim, 3) # Output three evaluation metrics def forward(self, audio_feat, text_feat): # audio_feat: [B, T, dim_audio], text_feat: [B, L, dim_text] context, weights = self.attention(audio_feat, text_feat) #context: [B, T, fusion_dim] fusion = torch.cat([audio_feat, context], dim=-1) # [B, T,dim_audio + fusion_dim] fusion_out = torch.tanh(self.fusion_layer(fusion)) # [B, T,fusion_dim] # Average pooling over the time dimension for the final prediction pooled = torch.mean(fusion_out, dim=1) # [B, fusion_dim] output = self.out_layer(pooled) # [B, 3], corresponding to sentiment matching degree, scene restoration degree, and emotional coherence return output, weights In the above implementation, the structured prompt template is first represented as a sequence of tokenized indices and vectorized by the SemanticEmbeddingEncoder module. This module uses a bidirectional GRU to perform temporal modeling on the semantic guiding words of each segment, thereby generating a semantic encoding stream with time-aligned features, providing an accurate semantic reference path for subsequent alignment analysis with acoustic frames.

[0087] After the preprocessed audio data passes through the AcousticFeatureExtractor module, the acoustic feature representation of each frame is output. These features can be composed of dimensions such as MFCC, pitch curvature, and rhythm boundaries, and are matched and fused in dimensions after linear transformation. Subsequently, the audio frames and the semantic encoding stream are jointly input into the CrossModalAttention module. By projecting them into query vectors and key-value vectors respectively, the system constructs a mutual attention weight matrix, extracts the linkage relationship between audio segments and semantic keywords, automatically identifies the mapping patterns between speech rate and emotional semantics, and between pitch level and keyword emphasis, and constructs an expression consistency map.

[0088] The fused context vector is sent to the MultimodalFusionEngine module. This module splices the acoustic features and the semantic attention output and performs non-linear fusion, and then through global pooling and the output layer, generates three structured evaluation metrics: emotional matching degree, scene restoration degree, and emotional coherence.

[0089] In the sound normalization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenes described in the present invention, the construction of the structured prompt template is a key pre-step in the emotional or scenario evaluation link. Its core role is to provide a unified, time-consistent, and scenario-complete semantic guiding benchmark for subsequent multi-modal fusion analysis. After receiving the evaluation target set by the user, the standard evaluation module first parses the setting information of the three dimensions input by the user: emotion labels (such as "angry", "sad", "imploring"), character identity parameters (such as "young female", "serious middle-aged male", etc.), and scenario setting information (such as "apologizing in public" or "comforting in a family conversation").

[0090] After parsing, the system performs semantic normalization processing on the input natural language description, and uses entity recognition and semantic normalization technology to unify the expressions into a standard label set. Each evaluation target is disassembled into independent fields such as emotion category, pragmatic scenario type, character gender, age range, and social identity. This label set not only ensures the structuring of information but also provides a fine-grained content scheduling basis for downstream tasks.

[0091] Next, the system constructs a prompt template with a clear hierarchical relationship based on the generated multi-field tag set. The template structure usually consists of three layers: the first layer is the global semantic main sentence, which is used to summarize the semantic goal of the entire evaluation task, such as "Please express the gentle persuasion emotion of an elderly father at a family gathering"; the second layer is the context guide, which is used to clearly express the situation and role background, such as "You are playing a male in his 60s, giving advice to relatives"; the third layer is the local semantic fragment, which is designed as a phrase or sentence unit corresponding one-to-one to the user's speech performance. The system will embed rhythm control markers (such as "[slow]", "[mid-stop]", "[emphasis]") and emotion regulation tips (such as "[low tone]", "[gentle intonation]") in this part. These information can not only guide the user's performance intention, but also provide time synchronization anchors and semantic weight indicators for the multimodal fusion module.

[0092] After construction, the system further converts this prompt template into a semantic coding stream that can be aligned with the audio frame-level timeline. Specifically, the system uses a pre-trained semantic embedding model (such as BERT, RoBERTa, etc.) to vectorize each semantic fragment and applies a time code corresponding to its position in the original speech expression to form a semantic coding stream tensor with time alignment attributes. Through this time alignment mechanism, the multimodal fusion analysis engine in the subsequent analysis stage can achieve frame-level interaction between acoustic features and semantic features, so as to achieve semantic-dominated alignment analysis and scoring inference when calculating key indicators such as emotional matching degree, scene restoration degree, and emotional coherence.

[0093] Through the above specific processing process, the structured prompt template not only completely covers the target emotion and scene setting of the evaluation task at the semantic level, but also provides a unified and schedulable representation method at the technical level, significantly improving the evaluation accuracy and interaction intelligence level of the system of the present invention in complex situations.

[0094] The integration processing module 104 is used to summarize the evaluation results generated by the standard evaluation module, generate a structured evaluation report, the evaluation report includes dimension scores, corresponding problem prompts, and personalized training suggestions, and send them back to the user terminal.

[0095] The integration processing module 104 plays an important role in integrating, analyzing, formatting, and feedback output of the evaluation results in this system. It is the interaction bridge between the user and the evaluation system, ensuring that the entire evaluation process forms a closed loop from input to feedback. The core function of this module is to process the multi-dimensional evaluation data output by the standard evaluation module 103, generate a complete structured evaluation report according to the system-set template, and accurately send this report back to the user terminal.

[0096] In specific implementation, the integration processing module first receives the data output from the standard evaluation module. This data includes the pronunciation accuracy rate of phonetic evaluation, the phoneme alignment results and error location; the scores of various acoustic features and their corresponding classification labels in basic skills evaluation; the matching degree scores, scene restoration indicators and coherence evaluations in emotion and scenario evaluation, etc. All these evaluation data are input in a structured format, including fields such as dimension identifiers, evaluation scores, comment template IDs, model confidence levels, etc. The integration processing module parses these structure fields, uniformly classifies and combines each evaluation dimension, and matches the corresponding natural language explanation statements to generate user-oriented explanatory content.

[0097] The system will automatically assign a quantitative level to each type of evaluation dimension according to the set scoring rules. For example, a pronunciation accuracy rate above 90% is marked as "excellent", 70%-90% is marked as "medium", and below 70% is marked as "needs improvement". Similarly, for dimensions with a low matching degree in emotion evaluation, the system will provide specific statements, such as "The intonation performance does not conform to the sad tone characteristics. It is recommended to slow down the speaking speed and increase the use of low-frequency intonation". Each scoring item is not only accompanied by a numerical value but also an explanatory label to assist users in understanding the deviation of their current speech characteristics.

[0098] To enhance the pertinence of feedback, the integration processing module will call the built-in expert rule library and training suggestion library, and generate personalized training suggestions by combining each evaluation score value and its context relationship. For example, if the user's "oral cavity resonance" score is low and the "sound intensity" fluctuates greatly, the system can output the suggestion: "It is recommended to increase the practice of oral cavity opening, try humming training to strengthen the anterior cavity resonance, and combine mirror-front pronunciation imitation to control the sound intensity stability". Such suggestions come from a pre-constructed label-training suggestion mapping table, which can ensure the executability and measurability of the suggestions.

[0099] In addition, the integration processing module supports the multi-level structure output of reports. The default output of the system is a standard format report, which includes four parts: score overview, dimension analysis, error prompt and improvement suggestions, and is suitable for most users. At the same time, an expert mode report is provided for professional users or teaching scenarios, which includes a complete feature index table, model scoring confidence level, error distribution heat map, etc., and supports in-depth analysis and comparison. In addition, this module also supports the historical record comparison function, which can call the user's past evaluation data and generate a progress curve or trend chart, so as to realize the tracking and quantification of the long-term vocal training results of users.

[0100] Finally, the integration and processing module packages the completed evaluation report into a visual data structure and transmits it back to the user device through the local interface or remote terminal. For APP or web users, the system can display the report in a combination of graphics and text, including bar charts, radar charts, text descriptions, and a summary block of suggestions. For API access users, the report data can be output in JSON or XML format for easy parsing and further processing by third-party systems. The module also supports exporting the report in PDF format for offline use or for teachers' review and reference.

[0101] Therefore, the integration and processing module not only realizes the standardized processing and user-friendly presentation of multi-source evaluation data, but also constructs a training guidance closed-loop through the feedback generation mechanism, endowing the entire evaluation system with the ability from diagnosis to intervention, and providing users with a practical guidance path for improving their voice expression ability.

[0102] In the sound standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios described in the present invention, the integration and processing module is not only responsible for sorting out the scoring results of various evaluation dimensions and generating reports, but also designs a set of structured, executable, and logically closed-loop training feedback paths. This path realizes an evaluation-tracking-training linkage mechanism aiming at the goal of improving user expression through accurate analysis of the current evaluation results, dynamic comparison of historical performance, and guided intervention of the expert knowledge base.

[0103] When the user completes a complete voice evaluation process, the standard evaluation module will output a structured evaluation result, which includes scoring information, scores, comment prompts, and types of expression problems detected by the system in multiple dimensions. The integration and processing module first receives this structured evaluation result, automatically analyzes and filters out the dimension labels with scores lower than the set threshold. The threshold can be preset by the system or flexibly set by the user according to their goals, such as below 70 points or the lowest 10% segment. These dimension labels will form the "weak item index set" in the current evaluation and serve as the starting point of the training guidance mechanism.

[0104] Subsequently, the system will retrieve the historical evaluation records from the local user database and conduct a time series comparison for each dimension corresponding to the current weak label one by one. During the comparison process, the system extracts the scoring trajectories of the same dimension in the historical N evaluation cycles and generates a trend curve covering the cycle. To ensure the interpretability of the trend, the system will also calculate statistical indicators such as the scoring stability coefficient, fluctuation range, extreme deviation points, and regression slope of each dimension in this cycle, and output a visual trend map, which intuitively presents the trend and the current score, helping users understand whether the problem exists for a long time or is a periodic fluctuation.

[0105] After obtaining the trend graph, the integration processing module will retrieve the training recommendation rule base and perform a structured matching for the weak dimensions. This rule base is constructed by speech teaching experts and organized in a ternary structure, that is, the mapping path of "dimension label - expression defect pattern - recommended training plan". For example, if the score of a certain user in the "speech rate rhythm" dimension is low, and the trend graph shows that its score fluctuates frequently and has poor stability, the system will identify it as the "rhythm control disorder" pattern and match the corresponding training nodes in the rule base, such as "breathing - stress marking training", "rhythm recognition training segment", etc. This matching process not only considers the dimension conformity, but also comprehensively analyzes the slope characteristics of the historical score trend, the score fluctuation range and the model confidence result to improve the accuracy and personalization degree of the training recommendation.

[0106] Finally, the integration processing module will embed the above training plan nodes into the user-visible evaluation report in a structured form. The report will display the current weak label, the specific performance description of this dimension, the typical errors or deficiencies detected by the system, the recommended training objectives, and the specific training task suggestions that can be directly executed in the form of pictures and texts. For example, if the problem of "unclear resonance position" is detected, specific practice suggestions such as "intensive reading training of single vowels", "experimental practice of cavity control" and operation instructions will be provided in the report. All contents will be formatted and displayed on the user terminal to ensure that the user can clearly understand the root cause of the problem, the training direction and the execution path, so as to achieve the logical closed-loop from evaluation determination to targeted training feedback.

[0107] In the voice standardization evaluation system described in the present invention, the training recommendation rule base is one of the core knowledge components in the integration processing module, and its function is to provide targeted training path suggestions according to the performance defects of users in specific dimensions after completing the structured evaluation. The construction of this rule base adopts a combination of expert experience and data-driven methods, and establishes a ternary binding mapping between the evaluation label, the expression defect and the recommended training plan according to the preset structure, so as to form a callable, extensible and inferable training recommendation system.

[0108] The original data sources of the training recommendation rule base include three aspects. First is the expert manual annotation data, that is, professional personnel with a speech teaching background perform semantic attribution annotation on a large number of historical evaluation records to identify the possible expression defect patterns behind each low-score performance. Second is the user evaluation log automatically archived by the system, including structured scoring results, acoustic feature distributions, expression consistency graphs and other information. Finally is the training task template data set collected by the platform, covering common and effective vocalization training methods in the current industry, such as rhythm control training, breath control training, emotion expression guidance, etc.

[0109] Based on the above data, the system first conducts a standardized definition of evaluation dimension tags and establishes a dimension index library, where each item identifies an independent sub-capability of voice expression, such as "pause and connection rhythm", "breath stability", "oral cavity resonance focus", etc. Subsequently, defect clustering is performed on historical records with scores lower than the preset threshold. According to the acoustic deviation characteristics and the expert feedback compared with the actual performance of the user, several reusable expression defect patterns are abstracted, such as "unstable pronunciation continuously", "insufficient ending sound of sentences", "rhythm acceleration not matching the target emotion", etc.

[0110] The system pairs the above dimension tags with the expression defect patterns one by one, and in the third stage, according to the feedback of teaching practice, binds each type of defect pattern to one or more training plan nodes. Each training plan node contains the following content: the name of the recommended training method (such as "breath fluctuation control exercise A"), the setting of the training scenario (such as "reading compound sentences with weak stress"), the execution method (such as "pause for 2 seconds for breath control between each sentence"), the recommended frequency and cycle (such as "5 times a day for 10 consecutive days"), and the evaluation standard prompt (such as "the rhythm stability score after retesting increases by no less than 15%").

[0111] To improve the system processing efficiency and the interpretability of rule invocation, the above ternary relationship is stored in the graph database in a structured format. Each node is identified by a unique index, and the edge relationship includes the association weight and the applicable scope annotation. When the system runs, it performs the shortest path inference according to the matching path between the current weak item label of the user and the defect graph, selects several training plan nodes with the highest adaptability as the recommended output, and displays them in a structured manner in the evaluation report.

[0112] The training recommendation rule library described in the present invention supports dynamic expansion and online update by experts. All ternary binding structures are maintained under version management. This mechanism can not only effectively enhance the interpretability and pertinence of the system in the feedback guidance process, but also build a stable and repeatable adaptive training closed-loop for users, improving the efficiency and accuracy of voice expression ability improvement.

[0113] In a typical application scenario of the present invention, after the user completes the voice recording on the client side, the system will automatically trigger the subsequent evaluation process, forming a complete execution link from user input to analysis and scoring. This link is manifested as a linkage communication mechanism between the front end, the back end and the algorithm service in actual deployment, and calls multiple evaluation subsystems to cooperate for processing to achieve a comprehensive evaluation of the user's audio. Please refer to Figure 2 , which is a timing schematic diagram of the user uploading voice and triggering the multi-path evaluation process in the sound standardization evaluation system for multi-dimensional acoustic parameters and dynamic analysis of emotional scenarios described in the present invention, showing the interaction logic and data call relationship between the user side, the back end, the algorithm side, and the pronunciation evaluation system, the emotion and scenario analysis system, and the basic skill evaluation system.

[0114] First, the user side performs voice recording and uploads the recorded audio to the backend service. After receiving the audio data, the backend system immediately conducts audio duration compliance detection, including determining whether it exceeds the specified duration range of the evaluation task and whether there are obvious missing, silent, or blank segments. The compliant audio will be temporarily stored and a unique resource locator (URL) will be generated for subsequent algorithm calls.

[0115] Next, the backend sends the generated audio URL to the algorithm side. After receiving the resource, the algorithm side first performs preliminary processing on the user audio, including operations such as standardizing noise reduction, frame-level slicing, and audio integrity verification, to ensure that subsequent feature extraction and model input have a unified time and amplitude scale. During this processing stage, the system will evaluate whether there are factors affecting the evaluation accuracy such as recording interruptions, frame skipping, and abnormal background noise, and may send back re-recording suggestions if necessary.

[0116] Subsequently, the system will automatically analyze the type of the item to be evaluated based on the user's evaluation configuration and the loaded demonstration audio content, and trigger the corresponding subsystem for special analysis: If the content to be evaluated includes indicators related to Mandarin pronunciation, comprehensive pronunciation, or reading skills, the pronunciation evaluation system will be called. This system first calls a third-party speech recognition component (such as the iFlytek pronunciation analysis interface) to obtain the preliminary phoneme recognition result of the user's pronunciation, and combines a self-developed algorithm to perform composite modeling on dimensions such as its accuracy, pronunciation fluency, and rhythm control, so as to output multi-dimensional scores.

[0117] If the evaluation content includes complex expression tasks such as emotional expression, role-playing, situational interpretation, or narrative interpretation, the system will call the emotion and situation analysis subsystem. Based on the multi-modal fusion mechanism described in the present invention, this module comprehensively analyzes the coupling degree between the acoustic parameters in the user's speech and the semantic information contained in the prompt word template, and calculates scores for dimensions such as its emotion matching degree, scene restoration degree, and emotion coherence, to determine whether the user's expression is consistent with the preset goal.

[0118] If the evaluation task in the audio involves voice basic skills, such as strong voice, breathy voice, breath stability, chest resonance and other indicators, the system will call the basic skills evaluation module to perform inference and determination on the user audio based on the trained classification model. The system will extract the dominant acoustic channel related to this task and use the feature screening and weight regulation strategy proposed in the present invention. Finally, by calculating the output probability distribution of the classification model and combining the positive class sample scoring criteria, the final score will be generated.

[0119] The evaluation results of all subsystems will be uniformly summarized in the integration processing module and form a structured scoring report. After completing the multi-dimensional analysis of the user's speech input, it will be transmitted back to the front-end interface so that the user can view the evaluation results, understand the deficiencies, and receive targeted training suggestions.

[0120] Through the orderly linkage and division of labor and cooperation among the above modules, the system constructs a highly extensible evaluation process framework, which not only realizes the end-to-end speech input processing, but also reflects the professionalism, real-time performance, and intelligent level of speech evaluation in multiple dimensions.

[0121] In the sound standardization evaluation system for multi-dimensional acoustic parameters and dynamic analysis of emotional scenes described in the present invention, in order to further improve the feature accuracy and model adaptability of the basic skill evaluation dimension, the system integrates a multi-stage analysis mechanism for acoustic signal features inside the basic skill evaluation system. Its core process includes key steps such as signal preprocessing, fundamental frequency and formant extraction, perturbation calculation, energy weighted analysis, and correlation matrix modeling. Please refer to Figure 3 , which is the correlation matrix heat map of acoustic features in the sound standardization evaluation system described in the present invention. This figure is used to show the correlation degree among various acoustic parameters (such as fundamental frequency, formant, perturbation, sound intensity, spectral center, etc.) in the basic skill evaluation. The colors from blue to red indicate the strength of negative correlation to positive correlation in turn, and can provide a reference basis for feature engineering and model feature selection.

[0122] First, the system performs pre-emphasis processing on the original audio signal input by the user to enhance high-frequency features and suppress low-frequency interference; then it performs frame segmentation operation, divides the continuous speech into multiple short frames with a duration of 20 milliseconds and overlaps them. Based on the autocorrelation algorithm, the system completes the fundamental frequency extraction of each frame, and at the same time combines the cepstrum analysis method to extract formant information and complete the source-filter model separation. In addition, the fundamental frequency perturbation (jitter) and amplitude perturbation (shimmer) are calculated through the relative difference of the inter-frame period change, and combined with the time-domain energy fluctuation to obtain a complete set of acoustic parameters.

[0123] In order to analyze the action relationship among various parameters under different evaluation dimensions, the system establishes a correlation matrix among multi-dimensional acoustic features and visualizes it as a heat map, as Figure 3 shown. This figure shows the pairwise correlation coefficients of various parameters including fundamental frequency (f0), formant (f1, f2tr, etc.), energy, jitter, shimmer, spectral center (gravity), and comprehensive perception features (such as energy_ratio), reflecting the coupling degree among features and the potential collinearity risk during modeling.

[0124] By analyzing the correlation structure in the heat map, the system systematically summarizes the core acoustic feature groups relied on by different basic skill assessment items. For example: The assessment of voice brightness / dullness relies on sound intensity, fundamental frequency, harmonic noise ratio, and energy ratio; The assessment of oral cavity resonance highly relies on the spectral center frequency and energy distribution; The assessment of breathy voice and breath stability mainly focuses on fundamental frequency perturbation and amplitude perturbation; The assessment of chest resonance and strong voice relies more on formant position, sound intensity, and perturbation indicators; The assessment of speech rate control and rhythm smoothness needs to combine short-time energy fluctuation and inter-frame frequency jump for comprehensive modeling.

[0125] After completing the construction of correlation features, the system adopts a data-driven strategy to screen feature combinations and conduct multiple rounds of model training on more than 28,000 training samples annotated by professional teachers. For different assessment items, mainstream machine learning algorithms such as support vector machine (SVM), random forest, gradient boosting, logistic regression, multi-layer perceptron (MLP), and decision tree are used respectively to evaluate the accuracy, recall rate, F1 score, and AUC index of each model, and automatically screen the model structure with the best performance for binding. The actual measurement results are as follows: For breathy voice, strong breath control, weak breath control, and dull voice: The support vector machine model performs best, with the accuracy ranging from 82.45% to 91.13%; For strong voice and chest resonance: The random forest model is used, with the accuracy being 88.89% and 91.00% respectively; For oral cavity resonance: The gradient boosting algorithm is adopted, with the accuracy reaching 83.00%; For voice brightness: The logistic regression model achieves a high accuracy of 94.35%; For full voice and weak voice: The multi-layer perceptron and decision tree are used respectively, with the accuracy being 91.14% and 82.64%.

[0126] In terms of emotion and scenario analysis, the system further divides the tasks into three categories: single emotion, scenario deduction, and role-playing, and constructs a supporting prompt word template generation mechanism, multi-modal fusion modeling path, and personalized evaluation algorithm respectively: The single emotion assessment covers emotional fullness and 15 specific emotions (such as joy, anger, sadness, fear, etc.). Among them, the emotional fullness is classified by logistic regression using acoustic features, with the accuracy reaching 84.47%; other emotions are analyzed for matching degree through the prompt word template + multi-modal audio comparison mechanism, and the accuracy is between 88% and 95%.

[0127] The scenario deduction assessment supports users to specify specific scenarios and target emotions, and the system dynamically generates a structured prompt word template. The fusion module outputs three-dimensional indicators of emotional matching degree, scenario restoration degree and emotion coherence, and the comprehensive accuracy rate reaches 87%.

[0128] The role-playing assessment superimposes role setting parameters (age, personality, etc.) on the basis of the scenario deduction, and introduces a feature coupling mechanism into the fusion model, jointly models with voice and text, and evaluates the role fitness. The measured accuracy rate reaches 85%-90%.

[0129] Although this application is disclosed above with preferred embodiments, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the protection scope of this application should be subject to the scope defined by the claims of this application.

Claims

1. A sound standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios, characterized in that, Including: An audio acquisition module, which is used to acquire the user's voice audio and generate raw audio data; An audio processing module, which is used to preprocess the raw audio data to obtain preprocessed audio data; wherein, the preprocessing includes noise reduction, sampling rate adjustment, pre-emphasis, framing and slicing operations; A standard evaluation module, which is used to determine the evaluation type according to the user's setting; receive the preprocessed audio data and perform an evaluation operation corresponding to the evaluation type, and the evaluation operation includes at least one of the following operations: Identify the speech units in the preprocessed audio data, calculate the accuracy rates of the initial consonants, final consonants, and phonetic flow change linguistic features, and generate a pronunciation evaluation result; Extract the acoustic features in the preprocessed audio data, including fundamental frequency, formant, sound intensity, fundamental frequency perturbation, and amplitude perturbation, and input the features into machine learning models separately trained for different evaluation items for classification prediction to generate a basic skills evaluation result; According to the user's emotion or scenario information, construct a prompt word template, and perform multi-modal fusion analysis on the preprocessed audio data and the scenario text, output the emotional matching degree, scene restoration degree, and emotion coherence index, and generate an emotion evaluation result; An integration processing module, which is used to summarize the evaluation results generated by the standard evaluation module, generate a structured evaluation report, and the evaluation report includes dimension scores, corresponding problem prompts, and personalized training suggestions, and send them back to the user terminal.

2. The sound normalization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios according to claim 1, characterized in that When the standard evaluation module performs the voice basic skills evaluation operation, for each preset evaluation item, the following feature selection process is performed: Based on the scoring dimension of the evaluation item, group the collected user samples and the labeled teacher samples, construct a control sample set covering all levels of performance, and extract an acoustic parameter set related to vocal stability, spectral change pattern, and energy distribution structure based on the control sample set. The acoustic parameter set includes: fundamental frequency perturbation, amplitude perturbation, spectral density centroid position, and non-steady energy fluctuation amplitude; Perform between-group difference enhancement analysis on the acoustic parameter set, use a convolutional structure based on a learnable weight kernel to generate a feature response map, calibrate the feature change patterns with significant discrimination for different levels of performance, and construct a feature-sensitive subspace for the corresponding evaluation item; In the feature-sensitive subspace, evaluate the interference risk of candidate features with respect to feature overlap in other evaluation items, and adjust the feature weights by introducing an overlap suppression factor to screen out a target feature group with high discrimination and minimum interference, and use it as the final feature corresponding to the evaluation item and input it into the trained exclusive classification model for evaluation determination.

3. The sound normalization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios according to claim 1, characterized in that When the standard evaluation module performs the voice basic skills evaluation operation, for each preset evaluation item, the following steps are sequentially performed to dynamically determine the optimal classification model bound to the evaluation item: Based on the control sample set covering all scoring levels of the evaluation item, extract the distribution response map of the acoustic feature channels, and analyze the channel response gradient between different levels to obtain the dominant acoustic channel set corresponding to the evaluation item; Based on the set of dominant acoustic channels, evaluate the response stability of each candidate model structure on this channel set, and eliminate model types that have a tendency of non-linear overfitting to the dominant channels or are sensitive to channel drift, forming the structural adaptability screening results; Deploy the models retained in the structural adaptability screening results to the historical sample set of this evaluation item for cross-validation respectively. Calculate the unified evaluation index based on accuracy, recall rate, F1 score, and area under the curve, and use the model with the highest evaluation index as the final classification model and bind it to the evaluation item for subsequent sample evaluation and scoring.

4. The sound normalization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios according to claim 1, characterized in that When performing emotion or scenario evaluation operations, the standard evaluation module sequentially performs the following steps to achieve multi-modal fusion path construction and evaluation calculation: Synchronously align the preprocessed audio data with the structured prompt word template, where the structured prompt word template is automatically generated according to the target emotion, scene description, or role setting selected by the user, and contains a semantic guiding word sequence and rhythm control marks. The system converts it into a time-aligned semantic coding stream through a segment semantic embedding mechanism; Input the semantic coding stream, the reference demonstration audio, and the preprocessed audio data into the multi-modal fusion engine together. The fusion engine uses a cross-modal attention mechanism to extract the dynamic coupling pattern between pronunciation intensity, pitch change, speech rate rhythm, and keyword semantic emphasis based on the mutual attention weight matrix between acoustic frame features and semantic embedding sequences, and constructs an expression consistency map; Extract multi-dimensional intermediate indicators including local inconsistency scores, emotion channel missing marks, and semantic drift degrees according to the expression consistency map, and fuse the global scene constraint information to calculate the evaluation results in three dimensions of final emotion matching degree, scene restoration degree, and emotion coherence, for forming a structured scoring feedback under the emotion dimension.

5. The sound standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios according to claim 1, characterized in that, When performing emotion or scenario evaluation operations, the standard evaluation module performs the following steps to construct a structured prompt word template for the evaluation target set by the user: Receive and parse the target emotion label, role identity parameter, and scene setting information input by the user, and map them to a multi-field label set through semantic standardization processing. The label set includes emotion category, pragmatic scene type, role gender, age range, and social identity; Based on the multi-field label set, construct a structured prompt word template with a hierarchical nested structure. The structured prompt word template includes: a global semantic main sentence defining the task objective, a context guiding language carrying role attribute constraints, and local semantic segments for prompting expression details; among them, rhythm control marks and emotion regulation prompts are embedded in the local semantic segments to support multi-modal alignment processing; Encode the structured prompt word template into a time-aligned semantic coding stream, and bind it to the target speech sample in the frame-level time dimension through a position embedding mechanism for the multi-modal fusion analysis engine to call, thereby supporting the association reasoning of expression consistency judgment and scenario index scoring.

6. The sound standardization evaluation system for dynamic analysis of multi-dimensional acoustic parameters and emotional scenarios according to claim 1, characterized in that, During the process of generating a structured evaluation report, the integration processing module sequentially performs the following steps for each evaluation dimension involved in the current evaluation result to construct a training feedback closed-loop path: Receive and parse the current structured evaluation result output by the standard evaluation module, extract all dimension labels with scores lower than the preset threshold, form a targeted training index set, and perform item-by-item comparison at the dimension level with the historical evaluation records stored in the local database; Based on the comparison result, call the progress archiving module to generate a quantitative training curve covering the most recent N evaluation cycles. The training curve reflects the time-series score change trend of each index. At the same time, calculate the stability coefficient and fluctuation range of each dimension to form a visual user performance trend map; Match the trend map with the training suggestion rule base, which is constructed based on the experience of speech teaching experts and adopts a ternary binding mechanism of "dimension label - expression defect pattern - recommended training plan". Through map structure alignment and defect pattern recognition, locate the most valuable expression shortcoming of the current user and screen out several personalized training plan nodes with the highest matching degree; Embed the training plan nodes into the final evaluation report in a structured manner, and display a five-dimensional feedback result including the current weak item label, specific performance description, typical problem prompt, recommended training target and specific training task suggestion on the user terminal, thus constituting a complete feedback closed-loop path from evaluation determination, historical tracking, trend modeling to training guidance.

Citation Information

Patent Citations

  • Audio quality comprehensive evaluation method and system

    CN109147765A

  • Generative AI emotion propagation prediction and guidance large model construction method and system

    CN119047512A

  • Service quality monitoring method and device based on multi-modal model sentiment analysis

    CN119477090A

  • System for automatic assessment of fluency in spoken language and a method thereof

    WO2021074721A2

  • Multi-modal model generation method, multi-modal processing method, and device

    WO2025031090A1

Cited By

  • Emotion analysis method and system

    CN120526807A

  • Speech recognition authentication method and system based on multi-modal features and dynamic evaluation

    CN120748413A

  • Intelligent call auxiliary system for old people

    CN121281527A

  • Emotion recognition method and device based on large model, medium, equipment and product

    CN121483311A

  • Audio processing method and device

    CN121641072A