Business english oral practice method
By collecting and analyzing speech signals, and combining acoustic and natural language processing technologies, the system identifies business scenario categories and sets evaluation criteria, thus addressing the shortcomings of traditional evaluation methods. This enables intelligent and personalized assessment and improvement of business English speaking skills, thereby enhancing students' business English communication abilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN UNIV
- Filing Date
- 2026-04-01
- Publication Date
- 2026-07-10
AI Technical Summary
Traditional business English oral proficiency assessment methods struggle to accurately identify, deeply understand, and comprehensively evaluate students' business English oral proficiency. Furthermore, the assessment results often deviate from actual abilities, and the methods fail to dynamically adjust assessment priorities based on specific business scenarios, thus affecting the accuracy and practicality of the assessment.
By collecting speech signals, extracting spectral feature data using acoustic models, and combining natural language processing and a business context dictionary for semantic analysis, we can determine the business scenario category, set evaluation criteria, and generate personalized improvement plans, including evaluation and improvement of pronunciation quality, grammatical accuracy, and fluency.
It enables intelligent and personalized assessment of business English speaking skills, improves students' business English communication abilities, and provides targeted and accurate improvement solutions.
Smart Images

Figure CN122369455A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of information processing technology, and in particular relates to a method for training business English oral communication. Background Technology
[0002] Business English oral proficiency assessment faces complex challenges in terms of accuracy, personalization, and practicality. Traditional assessment methods often fail to comprehensively capture students' language performance in real-world business scenarios, leading to discrepancies between assessment results and actual abilities. The limitations of speech recognition technology in handling non-standard pronunciation, accents, and background noise can result in errors in transcribed text, affecting the accuracy of subsequent analysis. Furthermore, the widespread use of professional terminology, industry slang, and context-dependent expressions in business English increases the difficulty of semantic comprehension. Developing assessment criteria also faces the challenge of balancing language accuracy with business communication effectiveness; single-dimensional scoring cannot accurately reflect a student's comprehensive abilities. In addition, different business scenarios have varying requirements for language skills; dynamically adjusting assessment focus based on specific contexts and organically integrating assessment results with targeted training are crucial for improving the practicality of the assessment. These intertwined factors form a multi-faceted and multi-layered technical challenge: how to accurately identify, deeply understand, comprehensively assess, and effectively improve users' business English oral proficiency. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a business English oral communication training method that realizes intelligent and personalized assessment of business English oral communication skills, and can effectively improve students' business English communication skills.
[0004] This application provides a method for practical training in business English oral communication, including: The user's voice signal is collected, and the acoustic features of the voice signal are extracted using an acoustic model to obtain spectral feature data; The spectral feature data is phoneme decoded to obtain text transcription data and confidence score data for each word, and the precise semantic information of the speech signal is determined. Based on the precise semantic information, the business scenario category to which the voice signal belongs and the corresponding evaluation criteria for each indicator are determined; The voice signal is evaluated based on the evaluation criteria of various indicators under the aforementioned business scenario category, and corresponding improvement plans are generated.
[0005] Furthermore, the precise semantic information of the speech signal is determined in the following manner: Natural language processing algorithms are used to analyze the syntactic structure and semantic components of the text transcription data in order to determine the semantic association information between words. The confidence score data for each word is compared with a pre-set first threshold. If the confidence score of any word is lower than the first threshold, a business context dictionary is used to perform secondary analysis on the semantic association information between words in order to perform semantic disambiguation and obtain accurate semantic information.
[0006] Furthermore, the step of determining the business scenario category to which the voice signal belongs and the corresponding evaluation criteria for various indicators based on the precise semantic information includes: Keyword extraction is performed on the precise semantic information, and feature values that can characterize the business scenario are calculated according to the pre-set weights to determine the business scenario category to which the voice signal belongs; The business scenario category of the current voice signal is compared with the business scenario category of the previous voice signal to determine whether the dialogue scenario has changed. When the dialogue scenario changes, the business scenario code corresponding to the current business scenario category is obtained to activate the corresponding evaluation framework. Based on the aforementioned evaluation framework, various indicators and corresponding evaluation standards related to the current business scenario category are obtained.
[0007] Furthermore, the indicators include: pronunciation quality; the pronunciation quality of the speech signal is evaluated in the following ways: The spectral feature data is compared with a standard business English pronunciation template to obtain a phoneme-level similarity score. The similarity score for each phoneme is compared with a second threshold; Phonemes with similarity scores below the second threshold are identified as deviated phonemes, and the positions of the deviated phonemes in the phoneme sequence are marked. A pronunciation quality score is obtained by weighting the similarity score of each phoneme and the importance of each phoneme in the current business scenario.
[0008] Furthermore, the improvement scheme for the pronunciation quality is generated in the following manner: The phoneme categories to which each deviated phoneme belongs are statistically analyzed to obtain the deviation ratio of each phoneme category; Based on the phoneme categories with high deviation ratios, spectral feature difference data are obtained. Based on spectral feature difference data, a pronunciation quality improvement plan is generated by matching it with a pre-set pronunciation correction database.
[0009] Furthermore, the metrics include: grammatical accuracy; the grammatical accuracy of the speech signal is evaluated in the following ways, and an improvement scheme for the grammatical accuracy is generated: The parser is used to identify the types and locations of grammatical errors in the text transcription data, and the number of errors is counted. The grammatical accuracy was calculated based on the total number of sentences and the number of errors in the text transcription data. Determine whether the number of errors exceeds a preset third threshold; wherein, the third threshold is the maximum number of errors allowed in the current business scenario; If so, then an improvement plan for grammatical accuracy will be generated based on the grammatical error categories with the most frequent errors and the grammatical usage habits in the current business scenario.
[0010] Furthermore, the metrics include: fluency; the fluency of the speech signal is evaluated in the following ways, and an improvement scheme for the fluency is generated: Based on the prosodic features and pause patterns in the speech signal, a speech rate stability index and a intonation naturalness index are calculated, and a fluency score is generated by combining the frequency and repetition rate of filler words. Determine if the fluency score is lower than a pre-set fourth threshold; If so, the problematic areas affecting fluency, such as speech rate, intonation, and filler words, will be located, and the problem categories will be determined. Then, a fluency improvement plan will be generated based on the contextual data of the current business scenario.
[0011] Furthermore, based on the prosodic features and pause patterns in the speech signal, a speech rate stability index and a intonation naturalness index are calculated, and combined with the frequency and repetition rate of filler words, a fluency score is generated, including: The speech signal is segmented to extract speech rate variation data, pitch variation data, and pause duration data; Based on the speech rate change data and pause duration data, the speech rate stability index is obtained. Combined with the pitch change data, the pitch fluctuation and emotional expression are quantitatively evaluated to determine the tone naturalness index. Based on the text transcription data, the occurrence frequency and distribution location of filler words are obtained, and the frequency and repetition rate of filler words are statistically analyzed; After standardizing the speech rate stability index, intonation naturalness index, and the frequency and repetition rate of filler words, a multi-dimensional fusion calculation is performed to generate a fluency score.
[0012] The business English oral communication training method provided in this application enables intelligent and personalized assessment of business English oral communication skills, which can effectively improve students' business English communication abilities. Attached Figure Description
[0013] Figure 1 A flowchart of the business English oral training method provided in the embodiments of this application is shown. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this technical solution clearer, the following detailed description, in conjunction with specific embodiments, further illustrates this technical solution. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of this technical solution.
[0015] Please see as follows Figure 1 The flowchart shown illustrates a business English oral communication training method. Figure 1 As shown, the method includes: S101. Acquire the user's voice signal and extract the acoustic features of the voice signal using an acoustic model to obtain spectral feature data.
[0016] In this step, a deep learning speech recognition algorithm is used to process the business English speech signal input by the student. Spectral feature data is extracted from the original speech data using an acoustic model to obtain preliminary analysis results containing acoustic characteristics. As an example, a convolutional neural network (CNN) and a long short-term memory network (LSTM) from the deep learning speech recognition algorithm are used to process the student's input business English speech signal into frames. The frame length is set to 25 milliseconds, and the frame shift is 10 milliseconds. 40-dimensional spectral feature data is extracted using Mel-frequency cepstral coefficients (MFCC) to form the acoustic characteristic analysis results.
[0017] S102. The spectral feature data is phoneme decoded to obtain text transcription data and confidence score data for each word, and the precise semantic information of the speech signal is determined.
[0018] In this step, phoneme sequence decoding is performed based on the extracted spectral feature data and a pre-established business English language model to obtain initial text transcription data and confidence scores for each word. As an example, based on the spectral feature data and a Transformer-based business English language model, phoneme sequence decoding is performed, and the Viterbi algorithm is used to calculate the optimal path, outputting the initial text transcription data and a confidence score for each word, ranging from 0 to 1. Specifically, a Transformer-based acoustic model is used to extract 40-Vimel frequency cepstral coefficient features from a 16kHz sampled speech signal, combined with a language model containing 50,000 business English words for Viterbi decoding, outputting the initial text transcription data and a confidence score for each word; for example, the confidence score for "negotiation" is 0.92 while the confidence score for "proposal" is 0.78.
[0019] In practical implementation, the precise semantic information of the speech signal can be determined in the following ways: Step 1021: Use natural language processing algorithms to parse the syntactic structure and semantic components of the text transcription data to determine the semantic association information between words.
[0020] In this step, natural language processing algorithms are used to parse the syntactic structure and semantic components of the initial text transcription data to determine the semantic association information of each word in the text.
[0021] Step 1022: Compare the confidence score data of each word with the pre-set first threshold.
[0022] Step 1023: If the confidence score of any word is lower than the first threshold, the semantic association information between words is analyzed a second time using a business context dictionary to perform semantic disambiguation and obtain accurate semantic information.
[0023] In this step, if the confidence score of any word is lower than the pre-set first threshold of 0.8, a semantic disambiguation mechanism is activated. This mechanism combines a business context dictionary with secondary analysis to obtain the precise semantic information expressed by the student. For example, for the initial text transcription data, the semantic understanding module uses a dependency parser and a BERT model to parse the syntactic structure, achieving a part-of-speech tagging accuracy of 95%. If the confidence score of "presentation" is detected to be 0.75, which is lower than the threshold of 0.8, the business context dictionary is used to match the definition of "product demonstration" instead of "gift," generating precise semantic information. Specifically, using the BERT model for dependency parsing, "submit" in "submit the report" is identified as a predicate verb with a confidence score of 0.85. When the confidence score of "quarterly" is detected to be 0.75, which is lower than the threshold of 0.8, the business context dictionary is used to match the time range "Q1-Q4," correcting the semantics to "Q3 financial report."
[0024] S103. Based on the precise semantic information, determine the business scenario category to which the voice signal belongs and the corresponding evaluation criteria for each indicator.
[0025] In practical implementation, the business scenario category to which the voice signal belongs and the corresponding evaluation criteria for each indicator can be determined through the following methods: Step 1031: Extract keywords from the precise semantic information and calculate feature values that can characterize the business scenario according to the pre-set weights to determine the business scenario category to which the voice signal belongs.
[0026] In this step, based on precise semantic information, the business scenario category of the current dialogue is determined through keyword extraction and preset weight calculation. For example, the TF-IDF algorithm is used to extract keywords such as "quarterly report" and "revenue," and a scenario matching weight of 0.92 is calculated, classifying it as a financial reporting scenario. Specifically, keywords such as "contract" and "delivery" are extracted, and when the total weight reaches 0.89, it is classified as a logistics scenario. When "payment terms" appears, the status parameters are updated to switch to the financial scenario code F-1032.
[0027] Step 1032: Compare the business scenario category of the current voice signal with the business scenario category of the previous voice signal to determine whether the dialogue scenario has changed.
[0028] Step 1033: When the dialogue scenario changes, obtain the business scenario code corresponding to the current business scenario category to activate the corresponding evaluation framework.
[0029] In this step, the scene recognition code is parsed, a pre-built business dialogue evaluation framework is invoked, and a hash table index matching method is used to retrieve the scene classification information corresponding to the code "MEETING001" from the framework, determining that the current dialogue belongs to a business meeting context. If a scene transition signal is detected, the dialogue state parameters are updated to obtain the corresponding scene code. As an example, when the topic is detected to have shifted to "marketing strategy" and the weight change exceeds 15%, the scene recognition code is updated to MKT02.
[0030] Step 1034: Based on the evaluation framework, obtain the various indicators and corresponding evaluation standards related to the current business scenario category.
[0031] In this step, a pre-established business dialogue evaluation framework is activated based on the scene recognition code. Evaluation standard data relevant to the current scene is then retrieved to determine the reference basis for subsequent analysis. For example, the evaluation framework is loaded using the MKT02 code to obtain evaluation indicators for this scene, such as a standard speech rate of 120 words / minute and a 30% proportion of technical terms. Specifically, the evaluation framework loads 12 standards from the F-1032 scene, including a terminology accuracy requirement of ≥90%, and a cosine similarity calculation showing a match of 0.82 between the student's expression "wire transfer" and the standard term "TT payment".
[0032] S104. Evaluate the voice signal based on the evaluation criteria of various indicators under the business scenario category, and generate corresponding improvement schemes.
[0033] The indicators include: pronunciation quality.
[0034] The pronunciation quality of the speech signal is evaluated using the following methods: Step 201: Compare the spectral feature data with the standard business English pronunciation template to obtain a phoneme-level similarity score.
[0035] In this step, the pronunciation accuracy assessment algorithm uses the Dynamic Time Warping (DTW) algorithm to compare the extracted spectral feature data with the standard business English pronunciation template, calculate the similarity score for each phoneme, with the score range from 0 to 1, and retain two decimal places of precision.
[0036] Step 202: Compare the similarity score of each phoneme with the second threshold.
[0037] In this step, when the similarity of "negotiation" is 0.72, which is lower than the threshold of 0.75, the vowel deviation of the second syllable is marked.
[0038] Step 203: Identify phonemes with similarity scores below the second threshold as deviated phonemes and mark the position of the deviated phonemes in the phoneme sequence.
[0039] In this step, if the similarity score is lower than the dynamically adjusted threshold of 0.75, a phoneme alignment technique based on a Hidden Markov Model (HMM) is used to mark the specific position of the deviated phoneme on the time axis and generate deviation label data.
[0040] Step 204: Based on the similarity score of each phoneme and the importance of each phoneme in the current business scenario, a weighted average is calculated to obtain the pronunciation quality score.
[0041] In this step, a weighted average is calculated for the similarity scores of all phonemes. The weights are dynamically adjusted based on the importance of the phonemes in business English to generate pronunciation quality score data, with a score range of 0 to 100.
[0042] Furthermore, the improvement scheme for the pronunciation quality is generated in the following manner: Step 205: Calculate the phoneme category to which each deviated phoneme belongs, so as to obtain the deviation ratio of each phoneme category.
[0043] Step 206: Obtain spectral feature difference data based on the phoneme categories with higher deviation ratios containing the deviated phonemes.
[0044] In this step, the proportion of phonemes with a deviation ratio below 0.75 is counted. If the deviation ratio of a specific phoneme category exceeds a preset threshold of 0.3, the K-means clustering algorithm is used to group the deviating phonemes, generating phoneme deviation classification data. The spectral feature differences between the deviating phonemes and the standard template are compared, and the spectral envelope differences are calculated using Euclidean distance to obtain feature difference analysis data.
[0045] Step 207: Based on the spectral feature difference data, generate a pronunciation quality improvement scheme by matching it with a pre-set pronunciation correction database.
[0046] In this step, based on the feature difference analysis data, a predefined pronunciation correction strategy library is matched, such as recommending tongue position adjustment exercises for vowel deviations, to generate targeted pronunciation improvement suggestions.
[0047] In addition, the metrics include: grammatical accuracy.
[0048] The grammatical accuracy of the speech signal is evaluated using the following methods, and an improvement scheme for the grammatical accuracy is generated: Step 301: Use a parser to identify the types and locations of grammatical errors in the text transcription data, and count the number of errors.
[0049] In this step, a syntactic tree structure is generated using a dependency parser, which annotates components such as subject, verb, and object. For example, it identifies subject-verb disagreement errors in "He go to school." The parser detects grammatical errors based on rules and statistical models, such as tense errors, singular / plural errors, or misuse of prepositions, and records the error location. For instance, it detects subject-verb disagreement in "They is working" where the error occurs in the second word.
[0050] Step 302: Calculate the grammatical accuracy based on the total number of sentences and the number of errors in the text transcription data.
[0051] In this step, the number of errors of each type is accumulated, such as 2 tense errors and 1 missing article, for a total of 3 errors. The formula for calculating grammatical accuracy is (1 - total number of errors / total number of words) × 100. If the total number of words is 50 and the number of errors is 3, then the accuracy rate is 94%.
[0052] Step 303: Determine whether the number of errors exceeds a preset third threshold; wherein the third threshold is the maximum number of errors allowed in the current business scenario.
[0053] In this step, the fault tolerance limit for the current business scenario (such as a meeting invitation) is obtained.
[0054] Step 304: If so, generate an improvement plan for grammatical accuracy based on the grammatical error categories with the most errors and the grammatical usage habits in the current business scenario.
[0055] As an example, if the third threshold is 3 and the actual number of errors reaches 4, then the tense error with the highest error rate is marked as the focus of improvement. Improved annotation information is associated with the error location (e.g., the third word in sentence 2) and correction suggestions (the past tense should be used), generating structured evaluation data. The evaluation criteria for a "formal email" scenario are applied, requiring a grammatical accuracy of ≥95%. If the actual measured value is 94%, a difference marker is triggered. The training module extracts corresponding practice materials from the resource library based on the error type (e.g., tense errors), such as tense conversion exercises, generating personalized learning path data to be pushed to the user.
[0056] In addition, the metrics include: smoothness.
[0057] The fluency of the speech signal is evaluated using the following methods, and a fluency improvement plan is generated: Step 401: Calculate the speech rate stability index and intonation naturalness index based on the prosodic features and pause patterns in the speech signal, and generate a fluency score by combining the frequency and repetition rate of filler words.
[0058] In practice, fluency scores can be generated in the following ways: Step 4011: Perform segmentation processing on the speech signal to extract speech rate change data, pitch change data, and pause duration data.
[0059] Step 4012: Based on the speech rate change data and pause duration data, obtain the speech rate stability index. Combined with the pitch change data, quantitatively evaluate the pitch fluctuation and emotional expression to determine the tone naturalness index.
[0060] In this step, short-time energy analysis and fundamental frequency extraction techniques are used to analyze the prosodic features and pause patterns in the original speech data. The speech rate variation curve is calculated with a frame length of 20ms, and the standard deviation of the number of syllables per second is statistically analyzed. If the standard deviation exceeds 0.5 syllables, it is marked as a speech rate instability interval, yielding preliminary calculation results for the speech rate stability index. Combining the pitch variation data in the speech signal, intonation analysis is performed based on a hidden Markov model. The fundamental frequency mean of five consecutive speech segments is normalized, and its cosine similarity to a standard business English intonation template is calculated. If the similarity is below 0.65, the intonation is deemed unnatural, thus determining the numerical value of the intonation naturalness index.
[0061] Step 4013: Obtain the occurrence frequency and distribution location of filler words based on the text transcription data, and count the frequency and repetition rate of filler words.
[0062] In this step, when processing the original speech data into text, a sequence labeling model based on an attention mechanism is used to detect filler words such as "um" and "ah", and their frequency of occurrence in every 100 words is counted. If the frequency exceeds 8%, filler word marking is triggered. At the same time, the edit distance algorithm is used to detect repeated phrases and calculate the repetition rate data.
[0063] Step 4014: After standardizing the speech rate stability index, intonation naturalness index, frequency and repetition rate of filler words, perform multi-dimensional fusion calculation to generate a fluency score.
[0064] In this step, based on the speech rate stability index, intonation naturalness index, filler word frequency and repetition rate data, the entropy weight method is used to calculate the weight of each index. Speech rate stability is assigned a weight of 0.4, intonation naturalness a weight of 0.3, and filler words and repetition rate each a weight of 0.15. The fluency score data of 0-100 is generated by weighted summation.
[0065] Step 402: Determine whether the fluency score is lower than the preset fourth threshold.
[0066] Step 403: If so, locate the problem areas of speech rate, intonation, and filler words that affect fluency, determine the problem categories, and generate a fluency improvement plan based on the contextual data of the current business scenario.
[0067] In this step, the fluency score data is compared with a preset threshold of 75 points for business negotiation scenarios. If the score is lower than the threshold, it is determined that the business communication standard is not met. As an example, when locating the problem area, the isolated forest algorithm is used to detect outliers in the speech rate indicator. Combined with the low similarity intervals output by the intonation analysis model, the problem classification label is determined to be "speech rate fluctuation" or "flat intonation". Combining the contextual data of the "customer reception" branch in the business scenario, the standard dialogue template for this scenario is extracted as improvement reference content, and the targeted improvement suggestion data includes specific indicators such as "maintaining a speech rate of 120-140 words per minute". The improvement suggestion data is then stored in association with the fluency score data.
[0068] Furthermore, the system can set pronunciation weights (40%), grammar weights (30%), and fluency weights (30%) according to different scenarios. A comprehensive score of 68 points triggers targeted training, pushing learning content including minimal opposition exercises and business conversation templates. For example, a weighted fusion algorithm is used to integrate pronunciation quality scoring data (e.g., phoneme similarity score of 0.68), grammar evaluation data (e.g., 4 errors), and fluency scoring data (e.g., speech rate stability index of 0.72). The weights are dynamically adjusted based on a scenario-customized set of evaluation rules (e.g., business meeting scenario weight configuration: pronunciation 40%, grammar 30%, fluency 30%), resulting in a comprehensive evaluation result (e.g., 65 points) and preliminary improvement suggestions (e.g., focusing on correcting phoneme deviations). The comprehensive evaluation result is compared to a preset standard of 70 points. If the score is below the threshold, the system is deemed unqualified, triggering targeted training and extracting training requirement data (e.g., marking weak pronunciation elements as vowels / i / and consonants / θ / ). Weaknesses are analyzed by combining training requirements data with a set of evaluation rules (e.g., a tolerance limit of 3 errors), and category labels are determined (e.g., pronunciation label is "vowel correction," grammar label is "tense error"). Resource data is matched from the training content library based on these category labels (e.g., standard phoneme comparison audio for pronunciation training, and interactive exercises on tense rules for grammar training). A push algorithm (e.g., collaborative filtering algorithm based on historical learning records) is used to sort the resource data, and the push order is determined by combining business scenario recognition codes (e.g., SC101 negotiation scenario) (e.g., prioritizing pronunciation training). The list of push content (e.g., JSON format data packets) is transmitted via an interface; if transmission fails, a backup channel (e.g., FTP protocol) is switched and resent. The learning profile is updated based on the transmission results (e.g., recording training content ID and push time), and subsequent learning plan parameters are generated based on the comprehensive evaluation results (e.g., the next stage focuses on grammar and tense training).
[0069] The above content is only a preferred embodiment of the present invention. For those skilled in the art, many changes can be made in the specific implementation and application scope based on the ideas of the present invention. As long as these changes do not depart from the concept of the present invention, they all fall within the protection scope of the present invention.
Claims
1. A method for practical training in business English oral communication, characterized in that, The method includes: The user's voice signal is collected, and the acoustic features of the voice signal are extracted using an acoustic model to obtain spectral feature data; The spectral feature data is phoneme decoded to obtain text transcription data and confidence score data for each word, and the precise semantic information of the speech signal is determined. Based on the precise semantic information, the business scenario category to which the voice signal belongs and the corresponding evaluation criteria for each indicator are determined; The voice signal is evaluated based on the evaluation criteria of various indicators under the aforementioned business scenario category, and corresponding improvement plans are generated.
2. The method as described in claim 1, characterized in that, The precise semantic information of the speech signal is determined by the following method: Natural language processing algorithms are used to analyze the syntactic structure and semantic components of the text transcription data in order to determine the semantic association information between words. The confidence score data for each word is compared with a pre-set first threshold. If the confidence score of any word is lower than the first threshold, a business context dictionary is used to perform secondary analysis on the semantic association information between the words in order to perform semantic disambiguation and obtain accurate semantic information.
3. The method as described in claim 1, characterized in that, The process of determining the business scenario category to which the voice signal belongs and the corresponding evaluation criteria for various indicators based on the precise semantic information includes: Keyword extraction is performed on the precise semantic information, and feature values that can characterize the business scenario are calculated according to the pre-set weights to determine the business scenario category to which the voice signal belongs; The business scenario category of the current voice signal is compared with the business scenario category of the previous voice signal to determine whether the dialogue scenario has changed. When the dialogue scenario changes, the business scenario code corresponding to the current business scenario category is obtained to activate the corresponding evaluation framework. Based on the aforementioned evaluation framework, various indicators and corresponding evaluation standards related to the current business scenario category are obtained.
4. The method as described in claim 1, characterized in that, The metrics include: pronunciation quality; the pronunciation quality of the speech signal is evaluated in the following ways: The spectral feature data is compared with a standard business English pronunciation template to obtain a phoneme-level similarity score. The similarity score for each phoneme is compared with a second threshold; Phonemes with similarity scores below the second threshold are identified as deviated phonemes, and the positions of the deviated phonemes in the phoneme sequence are marked. A pronunciation quality score is obtained by weighting the similarity score of each phoneme and the importance of each phoneme in the current business scenario.
5. The method as described in claim 4, characterized in that, The improvement scheme for the pronunciation quality is generated in the following way: The phoneme categories to which each deviated phoneme belongs are statistically analyzed to obtain the deviation ratio of each phoneme category; Based on the phoneme categories with high deviation ratios, spectral feature difference data are obtained. Based on spectral feature difference data, a pronunciation quality improvement plan is generated by matching it with a pre-set pronunciation correction database.
6. The method as described in claim 1, characterized in that, The metrics include: grammatical accuracy; the grammatical accuracy of the speech signal is evaluated in the following ways, and improvement schemes for the grammatical accuracy are generated: The parser is used to identify the types and locations of grammatical errors in the text transcription data, and the number of errors is counted. The grammatical accuracy was calculated based on the total number of sentences and the number of errors in the text transcription data. Determine whether the number of errors exceeds a preset third threshold; wherein, the third threshold is the maximum number of errors allowed in the current business scenario; If so, then an improvement plan for grammatical accuracy will be generated based on the grammatical error categories with the most frequent errors and the grammatical usage habits in the current business scenario.
7. The method as described in claim 1, characterized in that, The metrics include: fluency; the fluency of the speech signal is evaluated in the following ways, and improvement schemes for the fluency are generated: Based on the prosodic features and pause patterns in the speech signal, a speech rate stability index and a intonation naturalness index are calculated, and a fluency score is generated by combining the frequency and repetition rate of filler words. Determine if the fluency score is lower than a pre-set fourth threshold; If so, the problematic areas affecting fluency, such as speech rate, intonation, and filler words, will be located, and the problem categories will be determined. Then, a fluency improvement plan will be generated based on the contextual data of the current business scenario.
8. The method as described in claim 7, characterized in that, The process involves calculating speech rate stability and intonation naturalness indices based on prosodic features and pause patterns in the speech signal, and combining these with the frequency and repetition rate of filler words to generate a fluency score, including: The speech signal is segmented to extract speech rate variation data, pitch variation data, and pause duration data; Based on the speech rate change data and pause duration data, the speech rate stability index is obtained. Combined with the pitch change data, the pitch fluctuation and emotional expression are quantitatively evaluated to determine the tone naturalness index. Based on the text transcription data, the occurrence frequency and distribution location of filler words are obtained, and the frequency and repetition rate of filler words are statistically analyzed; After standardizing the speech rate stability index, intonation naturalness index, and the frequency and repetition rate of filler words, a multi-dimensional fusion calculation is performed to generate a fluency score.