An english oral English multi-dimensional evaluation method based on double-engine cooperation and storage medium

CN122551828APending Publication Date: 2026-08-11SHANGHAI CIVIL AVIATION VOCATIONAL & TECH COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-14
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

现有口语自动评估系统多采用单一语音分析架构,仅能对发音、流利度、语调韵律等声学特征进行量化评分,难以对民航场景下的词汇规范、语法严谨性、内容逻辑性、服务切题性及交互得体性等语义能力实现有效评测,评估维度单一、覆盖不完整

Benefits of technology

[0016] The beneficial effects of this invention are as follows: This invention adopts an automatic speech scoring engine and a semantic evaluation engine to work together, and at the same time realizes a comprehensive evaluation of speech features such as pronunciation, fluency, intonation and rhythm, as well as semantic capabilities such as vocabulary and grammar, content logic, scenario relevance and service appropriateness. This solves the problem of single dimension and incomplete coverage in civil aviation service oral evaluation, and improves the completeness and professionalism of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551828A_ABST
    Figure CN122551828A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-dimensional English oral proficiency assessment method and storage medium based on dual-engine collaboration, belonging to the field of English oral proficiency assessment technology. It constructs a heterogeneous dual assessment engine; takes the audio of the spoken language to be tested as input, and uses the heterogeneous dual assessment engine to output speech scoring results and semantic scoring results; creates a multi-dimensional oral proficiency assessment system, mapping the speech scoring results and semantic scoring results to corresponding ability dimensions in the multi-dimensional oral proficiency assessment system; performs dimensional alignment and collaborative fusion processing on the speech scoring results and semantic scoring results to generate independent quantitative assessment results for each ability dimension; summarizes the independent quantitative assessment results of each ability dimension, and outputs a combined multi-dimensional oral proficiency diagnostic result to reflect the English oral proficiency status of the tested subject. This invention uses an automatic speech scoring engine and a semantic assessment engine working collaboratively to comprehensively assess and solve the problems of single dimensions and incomplete coverage in civil aviation service oral proficiency assessment, improving the completeness and professionalism of the assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of English oral assessment, and in particular relates to a multi-dimensional English oral assessment method and storage medium based on dual-engine collaboration. Background Technology

[0002] With the widespread application of artificial intelligence technology in language assessment, automated English oral proficiency evaluation has become an important technical means for the professional competence certification, job qualification testing, and service standardization assessment of civil aviation service personnel. Civil aviation service English oral proficiency has strict requirements for pronunciation standards, fluency, semantic accuracy, scenario adaptability, and appropriate interaction, necessitating a comprehensive and professional evaluation that balances pronunciation standardization and semantic expression. Existing technologies have significant shortcomings when applied to civil aviation service scenarios: Existing automatic oral assessment systems mostly adopt a single speech analysis architecture, which can only quantify and score acoustic features such as pronunciation, fluency, intonation and rhythm. They are unable to effectively evaluate semantic capabilities such as vocabulary standardization, grammatical rigor, content logic, service relevance and interaction appropriateness in civil aviation scenarios. The assessment dimensions are single and the coverage is incomplete.

[0003] Meanwhile, voice assessment and semantic assessment operate independently, and a unified capability dimension system and dimension mapping rules adapted to civil aviation service scenarios have not been established. The dual-modal data cannot achieve temporal alignment, modality normalization and collaborative fusion under the same framework. There is a lack of conflict identification and scenario-based correction mechanisms, resulting in insufficient assessment accuracy and poor scenario adaptability.

[0004] Furthermore, the system outputs mostly single comprehensive scores, which cannot form structured, multi-dimensional capability diagnostic results that meet the needs of civil aviation services. This makes it difficult to support the accurate assessment, personalized guidance, and standardized improvement of civil aviation service oral communication skills, and fails to meet the actual needs of professional and standardized assessment in the civil aviation industry. Summary of the Invention

[0005] To address the problems existing in the background technology, this invention provides a multi-dimensional assessment method and storage medium for spoken English in civil aviation services based on dual-engine collaboration.

[0006] This invention adopts the following technical solution: a multi-dimensional assessment method for spoken English based on dual-engine collaboration, comprising the following steps: A heterogeneous dual evaluation engine consisting of an automatic speech scoring engine and a semantic evaluation engine is constructed; the spoken audio to be tested is used as input, and the heterogeneous dual evaluation engine outputs the corresponding speech scoring results. and semantic scoring results ; Dynamically create a multi-dimensional oral proficiency assessment system and use the speech scoring results and semantic scoring results Each is mapped to the corresponding ability dimension in the multi-dimensional oral communication ability assessment system; In terms of ability, the voice scoring results and semantic scoring results Perform dimensional alignment and collaborative fusion processing to generate independent quantitative evaluation results for each capability dimension; The independent quantitative assessment results of each ability dimension are summarized to output a combined multidimensional oral proficiency diagnostic result, which reflects the English oral proficiency status of the tested subjects.

[0007] In a further embodiment, based on the combined multidimensional oral ability diagnosis results, at least one of the following is output: pronunciation correction guidance, fluency optimization suggestions, and scenario-based expression improvement schemes; The independent quantitative evaluation results of the same subject during iterative learning and repeated evaluation are stored in a time sequence, generating and outputting a trend curve of changes in oral proficiency dimension, so as to visually represent the improvement trend of the subject's oral proficiency and learning effect.

[0008] In a further embodiment, the voice scoring result The output process is as follows: The spoken audio to be tested is input into an automatic speech scoring engine to extract multi-level speech feature sets; The multi-level speech feature set includes Each first-level speech feature contains [number] speech features, and each first-level speech feature includes [number] speech features. Two-dimensional features of speech; Acoustic feature matching and normalization quantization are performed on the secondary features of each speech level to generate atomic feature linear score vectors between 0 and 100 partitions. ; Based on a pre-defined first-dimensional speech weight vector with second-level dimension weight vector Calculate the hierarchical score for each first-level dimension of speech. ; Combining atomic feature linear score vectors Hierarchical scores for all speech levels in the first dimension. Perform linear weighted fusion to generate the total score of the speech layer evaluation results. ; Hierarchical scores for each first-level dimension of speech. With total score The combined output forms a structured speech layer evaluation result that includes scores for each dimension and a total score. .

[0009] In a further embodiment, the semantic layer evaluation result The output process is as follows: The speech recognition text of the spoken language to be tested is input into the semantic evaluation engine to extract a multi-level semantic feature set; the multi-level semantic feature set includes Each semantic first-level feature contains [number] semantic first-level features, and each semantic first-level feature includes [number] semantic first-level features. Two semantic secondary dimension features; Text feature matching and normalization quantization are performed on each semantic second-level dimension feature to generate a linear score vector of semantic atomic features between 0 and 100 partitions. ; Based on a pre-defined semantic first-level dimension weight vector with second-level dimension weight vector Calculate the hierarchical score for each semantic first-level dimension. ; Combining semantic atomic feature linear score vector Hierarchical scores for all semantic first-level dimensions Perform linear weighted fusion to generate the total score of the semantic layer evaluation results. ; Hierarchical scores for each semantic first-level dimension With the total score Together as semantic layer evaluation results Output.

[0010] In a further embodiment, the dynamic creation process of the multi-dimensional oral proficiency assessment system is as follows: Based on the spoken content type and spoken application scenario corresponding to the spoken audio to be tested, an assessment system containing multiple levels of ability dimensions is adaptively constructed. The capability dimensions are divided into at least three categories: voice capability dimensions determined solely by voice scoring results, semantic capability dimensions determined solely by semantic scoring results, and comprehensive capability dimensions determined jointly by voice scoring results and semantic scoring results. Furthermore, corresponding weight allocation rules and quantitative calculation methods are configured for each ability dimension to form a dynamic multi-dimensional oral ability assessment system adapted to the current assessment task.

[0011] In a further embodiment, the voice scoring result The mapping steps to a multi-dimensional oral proficiency assessment system include: Establish hierarchical scores for the primary dimension of speech. Mapping function to the speech capability dimension: In the formula, This indicates the first dimension in the multi-dimensional oral proficiency assessment system. Initial quantification scores for each voice capability dimension. This is a speech component mapping function; The total score of the speech layer evaluation results Through mapping function A comprehensive scoring item mapped to the dimension of voice ability , For the overall speech term mapping function; Make the speech scoring results The scores at each level are fixedly bound to the voice ability dimensions, thus completing the dimensional mapping of all voice scoring results.

[0012] In a further embodiment, the semantic scoring result The mapping steps to a multi-dimensional oral proficiency assessment system include: Establish hierarchical scores for the first-level semantic dimension. Mapping function to semantic capability dimension: In the formula, This indicates the first dimension in the multi-dimensional oral proficiency assessment system. Initial quantization score for each semantic capability dimension It is a semantic item mapping function; The total score of the semantic layer evaluation results Through mapping function A comprehensive scoring item mapped to the semantic ability dimension , This is a semantic general term mapping function; Make semantic scoring results The scores at each level are fixedly bound to the semantic ability dimensions, thus completing the dimensional mapping of all semantic scoring results.

[0013] In a further embodiment, the dimension alignment and collaborative fusion process includes: The speech scoring results and semantic layer scoring results are aligned with time frames to unify the time granularity. Synchronizing the speech analysis window and semantic analysis window allows for the initial quantization score of the speech capability dimension to be obtained through matching within the same temporal frame. Initial quantization score of semantic capability dimension ; Initial quantitative score for the voice ability dimension Initial quantization score of semantic capability dimension Perform modal normalization to map to the same standard range. The normalized score is obtained. , ; For the comprehensive capability dimension, scenario-related weights are introduced. , ,satisfy Construct a fusion model: ;in, For the first A comprehensive assessment value integrating multiple capability dimensions. This serves as an identifier for the evaluation scenario type; The comprehensive ability dimension is evaluated based on preset threshold rules. Perform deviation calibration to generate the final independent quantitative evaluation results for each capability dimension.

[0014] In a further embodiment, the speech capability dimension includes: a speech expression dimension and a speech fluency dimension; The semantic capability dimension includes: language structure dimension, professional vocabulary dimension, and scene understanding dimension; The comprehensive capability dimension includes: the interactive communication dimension, which is obtained by coupling and weighting the naturalness of tone and the stability of speech rate of the speech capability dimension with the politeness, timeliness of response and appropriateness of expression of the semantic capability dimension, and carries the alignment fusion and comprehensive quantitative evaluation of bimodal scores.

[0015] A computer-readable storage medium storing at least one executable instruction that, when executed on an electronic device, causes the electronic device to perform the method described above.

[0016] The beneficial effects of this invention are as follows: This invention adopts an automatic speech scoring engine and a semantic evaluation engine to work together, and at the same time realizes a comprehensive evaluation of speech features such as pronunciation, fluency, intonation and rhythm, as well as semantic capabilities such as vocabulary and grammar, content logic, scenario relevance and service appropriateness. This solves the problem of single dimension and incomplete coverage in civil aviation service oral evaluation, and improves the completeness and professionalism of the evaluation.

[0017] This invention constructs a multi-dimensional oral proficiency assessment system adapted to civil aviation service scenarios. Through unified dimension mapping, time frame alignment, and modality normalization processing, it achieves accurate alignment and efficient fusion of speech and semantic dual-modal data, significantly improving the accuracy and reliability of civil aviation service oral proficiency assessment.

[0018] This invention employs a scenario-adaptive weight allocation and modal conflict correction mechanism, which can dynamically adjust the fusion strategy and calibrate abnormal results according to different civil aviation service positions and evaluation scenarios, greatly improving the adaptability and robustness of the evaluation system in civil aviation professional scenarios and avoiding evaluation distortion.

[0019] This invention ultimately outputs a combined multidimensional oral communication ability diagnostic result, forming a structured and interpretable ability profile, supporting the standardized assessment, personalized guidance, and refined ability improvement of civil aviation service oral communication, and better meeting the standardized and professional assessment needs of the civil aviation industry. Attached Figure Description

[0020] Figure 1This is a flowchart of a multi-dimensional assessment method for spoken English based on dual-engine collaboration, as described in Example 1.

[0021] Figure 2 The top view of the multidimensional evaluation result visualization interface in Example 1 (comprehensive score, ICAO level, 6-dimensional capability distribution radar chart and dual-engine consistency detection).

[0022] Figure 3 This is a visualization interface of the multidimensional evaluation results in Example 1 (comprehensive score, ICAO level, 6-dimensional evaluation dimension distribution - 2-dimensional speech ability dimension, 3-dimensional semantic ability dimension and 1-dimensional comprehensive ability dimension - and dual-engine modal conflict detection).

[0023] Figure 4 This is a visualization interface diagram of the first-level dimensional quantitative scores and personalized diagnostic feedback of the speech and semantic layers in Example 1.

[0024] illustrate: Figures 2-4 This is a visual interface diagram of the present invention. The UI elements are for illustrative purposes only and do not constitute a limitation on the claims. Detailed Implementation

[0025] The present invention will now be further described with reference to the accompanying drawings and embodiments.

[0026] Example 1 This embodiment targets spoken English in civil aviation services, employing the dual-engine collaborative multi-dimensional assessment method for spoken English described in this invention. It achieves a comprehensive assessment across multiple dimensions, including pronunciation accuracy, fluency, semantic accuracy, and scenario suitability. The visualization interface for the multi-dimensional assessment results and personalized diagnostic feedback is shown below. Figure 2 , Figure 3 , Figure 4 As shown.

[0027] like Figure 1 As shown, this embodiment discloses a multi-dimensional assessment method for English oral communication based on dual-engine collaboration, including the following steps: Construct a heterogeneous dual evaluation engine consisting of an automatic speech scoring engine and a semantic evaluation engine; The audio recording of the spoken language to be tested is used as input, and the heterogeneous dual evaluation engine is used to output the corresponding speech score. and semantic scoring results ; Dynamically create a multi-dimensional oral proficiency assessment system and use the speech scoring results and semantic scoring results Each is mapped to the corresponding ability dimension in the multi-dimensional oral communication ability assessment system; In terms of ability, the voice scoring results and semantic scoring results Perform dimensional alignment and collaborative fusion processing to generate independent quantitative evaluation results for each capability dimension; The independent quantitative assessment results of each ability dimension are summarized to output a combined multidimensional oral proficiency diagnostic result, which reflects the English oral proficiency status of the tested subjects.

[0028] It should be noted that the automatic speech scoring engine described in this embodiment is built upon acoustic feature analysis and speech signal processing. It uses spoken audio as direct input and quantifies speech-level features such as speech expression and fluency through methods including phoneme extraction, prosody analysis, duration detection, and pause recognition. This engine processes acoustic signals without relying on text semantics, achieving hierarchical scoring solely through audio waveforms, spectral features, and phoneme matching.

[0029] The semantic evaluation engine is built on a natural language processing and semantic understanding model. It first converts spoken audio into text, and then analyzes vocabulary usage, grammatical structure, content completeness, logical coherence, scene adaptability, and service interaction rationality.

[0030] While some existing oral assessment tools output scores for individual components such as pronunciation and fluency, as well as a total score, they lack a "secondary feature" system. First-level dimension Level Score The structured and configurable calculation system for the "total score" also fails to achieve a formal mapping between the speech scoring results and the multi-dimensional oral proficiency assessment system. These scores are mostly directly assigned or simply weighted, and cannot be dynamically adjusted according to different assessment scenarios such as words, sentences, and dialogues. Nor can they provide standardized dimensional inputs for subsequent bimodal fusion calculations, resulting in poor scenario adaptability of the assessment results and low data reusability.

[0031] The voice scoring results described in this embodiment The output process is as follows: The spoken audio to be tested is input into an automatic speech scoring engine to extract a multi-level speech feature set; the multi-level speech feature set includes Each first-level speech feature contains [number] speech features, and each first-level speech feature includes [number] speech features. The speech features are divided into two secondary dimensions. For example, the primary dimensions of speech features in this embodiment include: pronunciation accuracy, fluency, intonation and prosody, and completeness. 4; The corresponding second-level speech features are shown in Table 1.

[0032] To further illustrate, such as the current example Acoustic feature matching and normalization quantization are performed on the secondary features of each speech level to generate atomic feature linear score vectors between 0 and 100 partitions. The expression is as follows: ;in, For the first The first-level feature of speech A vector composed of the atomic scores of each second-dimensional speech feature. ; To adapt to the acoustic features of different spoken language tasks, a pre-defined first-level speech dimension weight vector is used. with second-level dimension weight vector Calculate the hierarchical score for each first-level dimension of speech. For ease of understanding, the speech first-level dimension weight vector ,satisfy Second-level dimension weight vector ,satisfy The corresponding weights are illustrated in Table 1, where the spoken audio samples are used for sentence evaluation.

[0033] The hierarchical scores for each first-level dimension of speech The calculation formula is as follows: .

[0034] Combining atomic feature linear score vectors Hierarchical scores for all speech levels in the first dimension. Perform linear weighted fusion to generate the total score of the speech layer evaluation results. : .

[0035] Hierarchical scores for each first-level dimension of speech. With total score The combined output forms a structured speech layer evaluation result that includes scores for each dimension and a total score. : .

[0036] The results reflect the individual performance of the test subjects in each assessment dimension, and also provide a quantitative conclusion on the overall oral proficiency level, which facilitates subsequent analysis and feedback.

[0037] Table 1 Using a similar expression, the semantic layer evaluation results described in this embodiment The output process is as follows: The speech recognition text of the spoken language to be tested is input into the semantic evaluation engine to extract a multi-level semantic feature set; the multi-level semantic feature set includes Each semantic first-level feature contains [number] semantic first-level features, and each semantic first-level feature includes [number] semantic first-level features. Two semantic secondary dimension features; Text feature matching and normalization quantization are performed on each semantic second-level dimension feature to generate a linear score vector of semantic atomic features between 0 and 100 partitions. ; A semantic first-level dimension weight vector based on presets (text content, topic type, and contextual logic). with second-level dimension weight vector Calculate the hierarchical score for each semantic first-level dimension. Semantic first-level dimension weight vector and second-level dimension weight vector Each of them satisfies the condition that they combine to form 1.

[0038] Combining semantic atomic feature linear score vector Hierarchical scores for all semantic first-level dimensions Perform linear weighted fusion to generate the total score of the semantic layer evaluation results. ; Hierarchical scores for each semantic first-level dimension With the total score Together as semantic layer evaluation results Output, .

[0039] Given that existing oral assessment methods do not dynamically adjust assessment dimensions according to the type of oral content and application scenario, they cannot adapt to the differentiated needs of different assessment tasks such as words, sentences, and dialogues. The assessment system is fixed and has poor scenario adaptability, resulting in a lack of targeted scoring results. In word assessment, there is an overemphasis on semantic dimensions, and in dialogue assessment, the interaction logic is weakened, ultimately affecting the accuracy and professionalism of the assessment.

[0040] Therefore, the dynamic creation process of the multi-dimensional oral proficiency assessment system described in this embodiment is as follows: Based on the spoken content type and spoken application scenario corresponding to the spoken audio to be tested, an assessment system containing multiple levels of ability dimensions is adaptively constructed. The capability dimensions are divided into at least three categories: voice capability dimensions determined solely by voice scoring results, semantic capability dimensions determined solely by semantic scoring results, and comprehensive capability dimensions determined jointly by voice scoring results and semantic scoring results. Furthermore, corresponding weight allocation rules and quantitative calculation methods (generally weighted calculation) are configured for each ability dimension to form a dynamic multi-dimensional oral ability assessment system adapted to the current assessment task.

[0041] The speech ability dimension includes: speech expression dimension and speech fluency dimension; The semantic capability dimension includes: language structure dimension, professional vocabulary dimension, and scene understanding dimension; The comprehensive capability dimension includes: the interactive communication dimension, which is obtained by coupling and weighting the naturalness of tone and the stability of speech rate of the speech capability dimension with the politeness, timeliness of response and appropriateness of expression of the semantic capability dimension, and carries the alignment fusion and comprehensive quantitative evaluation of bimodal scores.

[0042] The coupling mentioned in this embodiment specifically refers to weighted coupling, such as the comprehensive dimension of "professional expression", which is composed of pronunciation standard, oral fluency and the appropriateness of professional terminology. By weighting and integrating the clarity of pronunciation and fluency of expression at the speech level and the accuracy of professional terminology at the semantic level, the effectiveness of conveying professional instructions in civil aviation scenarios is evaluated. The comprehensive dimension of "communication effectiveness" is composed of tone naturalness, speech rate stability, content completeness, and logical coherence. It integrates tone and rhythm control at the voice level with information completeness and logical coherence at the semantic level to evaluate the efficiency and clarity of information transmission in a conversation. The comprehensive dimension of "scenario adaptability" is composed of tone naturalness, speech rate stability, scenario adaptability, and expression appropriateness. It combines tone adaptability and speech rate control at the speech level with scenario understanding and language appropriateness at the semantic level to assess the ability to cope in specific civil aviation scenarios.

[0043] In other words, if the spoken audio being tested is a word pronunciation assessment, then the corresponding first-level speech features are pronunciation accuracy, fluency, and completeness, and the corresponding first-level speech weight vector is: The weighting of pronunciation accuracy has been increased, while the weighting of completeness is adapted to the characteristic of words without pauses.

[0044] The aforementioned weighting and feature dimensions can be dynamically adjusted according to the assessment scenario, and are also mapped to the speech ability dimension in the multi-dimensional oral ability assessment system, providing structured data support for subsequent temporal alignment, modal fusion and comprehensive ability assessment.

[0045] When the spoken audio being tested is a dialogue assessment in a civil aviation scenario, this invention constructs a six-dimensional civil aviation service spoken language assessment framework covering speech, language, and interaction capabilities, adapting to the professional characteristics of the dialogue scenario: Voice expression dimension (weight 20%): including word pronunciation accuracy, intonation naturalness, and the degree of influence of accent; Fluency (weight 20%): includes moderate speaking speed, reasonable pauses, and natural and fluent delivery; Language structure dimension (weight 15%): includes grammatical correctness, sentence structure completeness, and logical coherence; Specialized vocabulary dimension (weight 20%): This includes the use of specialized terminology, vocabulary richness, and word accuracy; Scenario understanding dimension (weight 15%): includes correctly understanding the question, responding appropriately, and providing complete information; Interactive communication dimension (weight 10%): including politeness of expression, timeliness of response, and effectiveness of communication.

[0046] The corresponding evaluation dimension weight vectors are: speech expression 20%, language structure 15%, professional vocabulary 20%, fluency of expression 20%, scenario understanding 15%, and interactive communication 10%. By dynamically adjusting the dimension weights and quantification rules, a professional and multi-dimensional evaluation of civil aviation service dialogue scenarios can be achieved, solving the shortcomings of existing technologies that cannot cover interactive logic and scenario understanding.

[0047] Based on the structured speech layer evaluation results given in this embodiment... : and semantic layer evaluation results Considering that existing oral assessment methods lack a standardized and formalized mapping mechanism between speech and semantic scoring results and multi-dimensional oral ability assessment systems, resulting in poor dimensional adaptability across different scenarios, low data reusability, and an inability to provide a unified dimensional input for subsequent bimodal fusion computation, therefore, speech scoring results... The mapping steps to a multi-dimensional oral proficiency assessment system include: Establish hierarchical scores for the primary dimension of speech. Mapping function to the speech capability dimension: In the formula, This indicates the first dimension in the multi-dimensional oral proficiency assessment system. Initial quantification scores for each voice capability dimension. This is a speech component mapping function. It is generally a linear scaling function, in the form of: ,in, This is the scaling factor. This is the bias coefficient. It can be dynamically configured according to different assessment scenarios such as words, sentences, and dialogues. For example, in a word pronunciation assessment scenario, the pronunciation accuracy dimension... Set it to 1.2 to amplify the weight of pronunciation accuracy and adapt to the needs of the scenario.

[0048] The total score of the speech layer evaluation results Through mapping function A comprehensive scoring item mapped to the dimension of voice ability , For the overall speech term mapping function; The standardized mapping function has the following form: ,in, The average of the total speech scores from historical samples. The standard deviation of the total speech score in historical samples is used to map the total speech score of different evaluation scenarios to a standardized range, eliminating the differences in scoring standards between scenarios and providing a unified quantitative benchmark for subsequent dual-modal fusion calculation.

[0049] Make the speech scoring results The scores at each level are fixedly bound to the voice ability dimensions, thus completing the dimensional mapping of all voice scoring results.

[0050] Correspondingly, semantic scoring results The mapping steps to a multi-dimensional oral proficiency assessment system include: Establish hierarchical scores for the first-level semantic dimension. Mapping function to semantic capability dimension: In the formula, This indicates the first dimension in the multi-dimensional oral proficiency assessment system. Initial quantization score for each semantic capability dimension For semantic sub-item mapping functions, the same function as the speech sub-item mapping function can be selected.

[0051] The total score of the semantic layer evaluation results Through mapping function A comprehensive scoring item mapped to the semantic ability dimension , The semantic total term mapping function is selected, using the same function as the semantic total term mapping function.

[0052] Make semantic scoring results The scores at each level are fixedly bound to the semantic ability dimensions, thus completing the dimensional mapping of all semantic scoring results.

[0053] Finally, to overcome the shortcomings of existing oral assessment methods, such as the mismatch of time windows and inconsistent units between speech and semantic scoring results, which prevents effective collaborative integration and the lack of scenario-specific weights in comprehensive ability dimension assessment, the dimension alignment and collaborative integration processing in this embodiment includes: The speech scoring results and semantic layer scoring results are aligned with time frames to unify the time granularity. Synchronizing the speech analysis window and semantic analysis window allows for the initial quantization score of the speech capability dimension to be obtained through matching within the same temporal frame. Initial quantization score of semantic capability dimension Specifically, the fixed duration window for speech analysis and the sentence / dialogue turn window for semantic analysis are interpolated and aligned along a unified time axis to ensure that quantitative data of both speech and semantic dimensions exist simultaneously at each time point.

[0054] Initial quantitative score for the voice ability dimension Initial quantization score of semantic capability dimension Perform modal normalization to map to the same standard range. The normalized score is obtained. , This eliminates the differences in scoring scales between different modalities, providing a unified quantitative benchmark for subsequent fusion calculations.

[0055] For the comprehensive capability dimension, scenario-related weights are introduced. , By dynamically configuring the evaluation scenario type identifier, for example, in a civil aviation dialogue scenario, the weight of semantic dimensions related to interactive communication is increased. And it meets the following requirements: Construct a fusion model: ;in, For the first A comprehensive assessment value integrating multiple capability dimensions. This serves as an identifier for the evaluation scenario type; The comprehensive ability dimension is evaluated based on preset threshold rules. Perform deviation calibration to generate the final independent quantitative evaluation results for each capability dimension.

[0056] In a further embodiment, the deviation calibration includes modal conflict identification and scene-specific deviation correction: Preset modal conflict threshold Calculate the normalized score difference in frames within the same time frame. The embodiments described herein The calculation formula is: ; Based on modal conflict threshold Perform conditional judgment: If If no modal conflict is found, the fusion value is directly used. As the final evaluation value; like If a modal conflict is detected, the process will proceed to the scenario-specific correction process.

[0057] The scenario-based correction described in this embodiment refers to the scenario type. They are divided into three categories: speech-first, semantic-first, and balanced. When voice-priority is enabled, set an upper limit for voice weight. ,make , And substitute it into the fusion model to recalculate the fusion evaluation value of the comprehensive capability dimension; When semantic priority is prioritized, a semantic weight cap is set. ,make , And substitute it into the fusion model to recalculate the fusion evaluation value of the comprehensive capability dimension; When it is a balanced type, a deviation correction factor is introduced. and symbolic functions Construct a corrected model: ; Output the final quantitative evaluation value of the corrected comprehensive capability dimensions.

[0058] To verify the practical effect of the dual-engine collaborative English oral communication multidimensional assessment method described in this invention in civil aviation service scenarios, this embodiment uses a typical passenger seat service dialogue as an example for evaluation and explanation.

[0059] 1. Evaluation Scenarios and User Responses The evaluation scenario involved a passenger requesting a window seat from a flight attendant, and the user's response was: “Let me check…Yes, I can give you 23A, a window seat.” 2. Overall Evaluation Results The above answers were quantitatively evaluated using the multi-dimensional oral communication skills assessment system of this invention, and the following results were obtained: Total score: 78 / 100; Civil aviation language proficiency level: ICAO Level 4 (Operational); It should be noted that the ICAO Level 4 (Operational) mentioned in this embodiment is a job proficiency level constructed based on the ICAO Language Proficiency Requirements Framework (LPR) and combined with the characteristics of civil aviation service positions, and is not directly equivalent to the official ICAO language proficiency level applicable to pilots / air traffic controllers.

[0060] Overall rating: Good, meets the language requirements for civil aviation work; 3. Quantitative Results by Item Dimension Based on the classification of speech ability dimension, semantic ability dimension and comprehensive ability dimension according to the present invention, the corresponding scores of the six dimensions are shown in Table 2.

[0061] According to the data in Table 2, this invention, through a dual-engine collaborative scoring, structured mapping, and scenario-related fusion model, can achieve multi-dimensional and quantitative evaluation of dialogue in civil aviation service scenarios. In the semantic ability dimension, the scenario understanding dimension scored the highest (88%), indicating that users can accurately understand passenger needs and give a complete response; In the voice ability dimension, the voice expression dimension scored relatively high (82%), but the expression fluency dimension scored relatively low (72%), which is consistent with the hesitation and pause characteristics in the user's answers; The high score (85%) in the interactive communication dimension verifies that the present invention effectively evaluates complex communication abilities such as politeness and timely response through weighted coupling.

[0062] The comprehensive ability score (78%) is a fusion score after coupling and associating the voice ability dimension and the semantic ability dimension according to preset weights. It reflects not only the user's performance in the voice level such as pronunciation and fluency, but also the comprehensive level of the semantic level such as content understanding, scenario adaptation and appropriateness of interaction. Finally, it forms an overall quantitative evaluation of the civil aviation service English oral ability, which is completely consistent with the total score given in this embodiment.

[0063] The data from this embodiment demonstrates that the evaluation method of the present invention can cover the full-dimensional evaluation of voice, semantics, and comprehensive communication capabilities in civil aviation service scenarios, providing reliable quantitative support for the scenario-based and professional evaluation of English speaking ability.

[0064] Table 2 Example 2 Based on the dual-engine collaborative English oral proficiency multidimensional assessment method disclosed in Embodiment 1, this embodiment further discloses the application extension and iterative learning function of the method. That is, based on the combined multidimensional oral proficiency diagnosis results, at least one of the following is output: pronunciation correction guidance, fluency optimization suggestions, and scenario-based expression improvement schemes. At the same time, the independent quantitative assessment results of the same test subject in the iterative learning and repeated evaluation process are stored in a time sequence, and the oral proficiency dimension change trend curve is generated and output to visualize the oral proficiency improvement trend and learning effect of the test subject.

[0065] Taking the dialogue evaluation of civil aviation service scenarios in Example 1 as an example, the process of generating personalized feedback based on multi-dimensional evaluation results is explained.

[0066] The tested user responded: "Let me check...Yes, I can give you 23A, a windoweat." The combined diagnostic results obtained through the method described in Example 1 triggered the system to output the following multi-dimensional improvement guidance: Pronunciation correction guidance: Based on the diagnostic result of 82% in the speech expression dimension of speech ability, the problem of insufficient clarity of the / tʃ / sound in the word "check" was identified, and the following pronunciation practice suggestions were given: practice the / tʃ / sound more (check, choose, change), and compare and distinguish the pronunciation differences between / e / and / ɪ / (check vs. chick).

[0067] Fluency Optimization Suggestions: Based on the diagnostic result of a fluency score of 72%, the issue of an excessively long 2.5-second pause after "Let me check" was identified, and the following fluency improvement plan was proposed: shorten thinking time, use transitional phrases such as "One moment, please," and practice quick responses in common scenarios.

[0068] Contextualized Expression Improvement Solution: Based on the diagnostic results of a 75% score in the language structure dimension and contextual adaptability, the solution identifies issues such as overly colloquial and lacking formality in expression, and provides a reference expression for civil aviation scenarios: "Certainly. Let me check the availability… I can offer you seat 23A, which is a window seat. "Would that be suitable for you?", and suggested adding a confirmation question to enhance interactivity.

[0069] The above improvement suggestions correspond one-to-one with the diagnostic results of the speech, semantic and comprehensive capability dimensions in Example 1, achieving a precise match between problem points and improvement paths.

[0070] This embodiment stores iterative learning and repeated evaluation data of the same test subject over 30 days in a time-series format, generating a trend curve of oral language ability improvement. The specific representation is as follows: Overall improvement trend: The test subjects' total oral proficiency score increased from the initial 45 points to 80 points within 30 days, an improvement of +35 points. The test subjects practiced for 28 days. Their current civil aviation language proficiency level is ICAO Level 4 (Operational), with an average score of 72 points. The overall trend is stable upward.

[0071] Weakness Dimension Tracking: Through continuous analysis of time-series data, the top 3 weakest dimensions of the tested object were identified, namely: Fluency score -72%: The problem is excessive pauses and slow speaking speed. Targeted suggestions include practicing answering 5 common questions quickly every day, using transition words, and recording and comparing your speech to the standard speaking speed. Language structure dimension -75%: The problem is that the sentence structure is simple and lacks complex sentences. It is recommended to learn common complex sentence structures, practice using conjunctions (however, therefore), and refer to standard sentence patterns.

[0072] Through the aforementioned time-series storage and visualization, the tested individuals can intuitively grasp the improvement path, weaknesses, and learning outcomes of their oral communication skills, achieving a closed-loop iterative optimization of "assessment-feedback-learning-reassessment." This embodiment verifies that the present invention can not only achieve multi-dimensional quantitative assessment of a single dialogue, but also provide users with targeted learning guidance and continuous improvement paths through personalized diagnostic feedback and time-series trend representation, thereby enhancing the practicality and educational value of oral communication assessment methods.

[0073] Example 3 To implement the dual-engine collaborative multidimensional English speaking assessment method disclosed in Examples 1 and 2, this embodiment further discloses an English speaking multidimensional assessment system, including: The first module is configured to construct a heterogeneous dual evaluation engine consisting of an automatic speech scoring engine and a semantic evaluation engine; it takes the spoken audio to be tested as input and uses the heterogeneous dual evaluation engine to output the corresponding speech scoring results. and semantic scoring results ; The second module is configured to dynamically create a multi-dimensional oral proficiency assessment system and use the voice scoring results. and semantic scoring results Each is mapped to the corresponding ability dimension in the multi-dimensional oral communication ability assessment system; The third module is set up to score the voice results at the ability level. and semantic scoring results Perform dimensional alignment and collaborative fusion processing to generate independent quantitative evaluation results for each capability dimension; The fourth module is set to summarize the independent quantitative assessment results of each ability dimension and output a combined multidimensional oral ability diagnostic result to reflect the English oral ability status of the tested subject.

[0074] The fifth module is configured to output at least one of the following based on the combined multidimensional oral proficiency diagnostic results: pronunciation correction guidance, fluency optimization suggestions, and scenario-based expression improvement schemes. The sixth module is set to store the independent quantitative evaluation results of the same test subject in a time sequence during the iterative learning and repeated evaluation process, generate and output the trend curve of the change in the oral ability dimension, so as to visualize the improvement trend of the test subject's oral ability and the learning effect.

[0075] It also includes: a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the English oral multidimensional assessment method based on dual-engine collaboration as described in Embodiments 1 and 2.

Claims

1. A multidimensional assessment method for English oral communication based on dual-engine collaboration, characterized in that, Includes the following steps: A heterogeneous dual evaluation engine consisting of an automatic speech scoring engine and a semantic evaluation engine is constructed; the spoken audio to be tested is used as input, and the heterogeneous dual evaluation engine outputs the corresponding speech scoring results. and semantic scoring results ; Dynamically create a multi-dimensional oral proficiency assessment system and use the speech scoring results and semantic scoring results Each is mapped to the corresponding ability dimension in the multi-dimensional oral communication ability assessment system; In terms of ability, the voice scoring results and semantic scoring results Perform dimensional alignment and collaborative fusion processing to generate independent quantitative evaluation results for each capability dimension; The independent quantitative assessment results of each ability dimension are summarized to output a combined multidimensional oral proficiency diagnostic result, which reflects the English oral proficiency status of the tested subjects.

2. The English oral proficiency multidimensional assessment method based on dual-engine collaboration according to claim 1, characterized in that, Based on the combined multidimensional oral proficiency diagnosis results, at least one of the following is output: pronunciation correction guidance, fluency optimization suggestions, and scenario-based expression improvement schemes; The independent quantitative evaluation results of the same subject during iterative learning and repeated evaluation are stored in a time sequence, generating and outputting a trend curve of changes in oral proficiency dimension, so as to visually represent the improvement trend of the subject's oral proficiency and learning effect.

3. The English oral proficiency multidimensional assessment method based on dual-engine collaboration according to claim 1, characterized in that, The voice scoring results The output process is as follows: The spoken audio to be tested is input into an automatic speech scoring engine to extract multi-level speech feature sets; The multi-level speech feature set includes Each first-level speech feature contains [number] speech features, and each first-level speech feature includes [number] speech features. Two-dimensional features of speech; Acoustic feature matching and normalization quantization are performed on the secondary features of each speech level to generate atomic feature linear score vectors between 0 and 100 partitions. ; Based on a pre-defined first-dimensional speech weight vector with second-level dimension weight vector Calculate the hierarchical score for each first-level dimension of speech. ; Combining atomic feature linear score vectors Hierarchical scores for all speech levels in the first dimension. Perform linear weighted fusion to generate the total score of the speech layer evaluation results. ; Hierarchical scores for each first-level dimension of speech. With total score The combined output forms a structured speech layer evaluation result that includes scores for each dimension and a total score. .

4. The English oral proficiency multidimensional assessment method based on dual-engine collaboration according to claim 1, characterized in that, The semantic layer evaluation results The output process is as follows: The speech recognition text of the spoken language to be tested is input into the semantic evaluation engine to extract a multi-level semantic feature set; the multi-level semantic feature set includes Each semantic first-level feature contains [number] semantic first-level features, and each semantic first-level feature includes [number] features. Two semantic secondary dimension features; Text feature matching and normalization quantization are performed on each semantic second-level dimension feature to generate a linear score vector of semantic atomic features between 0 and 100 partitions. ; Based on the pre-defined semantic first-level dimension weight vector with second-level dimension weight vector Calculate the hierarchical score for each semantic first-level dimension. ; Combining semantic atomic feature linear score vector Hierarchical scores for all semantic first-level dimensions Perform linear weighted fusion to generate the total score of the semantic layer evaluation results. ; Hierarchical scores for each semantic first-level dimension With the total score Together as semantic layer evaluation results Output.

5. The English oral proficiency multidimensional assessment method based on dual-engine collaboration according to claim 1, characterized in that, The dynamic creation process of the multi-dimensional oral proficiency assessment system is as follows: Based on the spoken content type and spoken application scenario corresponding to the spoken audio to be tested, an assessment system containing multiple levels of ability dimensions is adaptively constructed. The capability dimensions are divided into at least three categories: voice capability dimensions determined solely by voice scoring results, semantic capability dimensions determined solely by semantic scoring results, and comprehensive capability dimensions determined jointly by voice scoring results and semantic scoring results. Furthermore, corresponding weight allocation rules and quantitative calculation methods are configured for each ability dimension to form a dynamic multi-dimensional oral ability assessment system adapted to the current assessment task.

6. The English oral proficiency multidimensional assessment method based on dual-engine collaboration according to claim 1, characterized in that, The voice scoring results The mapping steps to a multi-dimensional oral proficiency assessment system include: Establish hierarchical scores for the primary dimension of speech. Mapping function to the speech capability dimension: In the formula, This indicates the first dimension in the multi-dimensional oral proficiency assessment system. Initial quantification scores for each voice capability dimension. This is a speech component mapping function; The total score of the speech layer evaluation results Through mapping function A comprehensive scoring item mapped to the dimension of voice ability , For the overall speech term mapping function; Make the speech scoring results The scores at each level are fixedly bound to the voice ability dimensions, thus completing the dimensional mapping of all voice scoring results.

7. The English oral proficiency multidimensional assessment method based on dual-engine collaboration according to claim 1, characterized in that, The semantic scoring results The mapping steps to a multi-dimensional oral proficiency assessment system include: Establish hierarchical scores for the first-level semantic dimension. Mapping function to semantic capability dimension: In the formula, This indicates the first dimension in the multi-dimensional oral proficiency assessment system. Initial quantization score for each semantic capability dimension It is a semantic item mapping function; The total score of the semantic layer evaluation results Through mapping function A comprehensive scoring item mapped to the semantic ability dimension , This is a semantic general term mapping function; Make semantic scoring results The scores at each level are fixedly bound to the semantic ability dimensions, thus completing the dimensional mapping of all semantic scoring results.

8. The English oral proficiency multidimensional assessment method based on dual-engine collaboration according to claim 1, characterized in that, The dimension alignment and collaborative fusion process includes: The speech scoring results and semantic layer scoring results are aligned with time frames to unify the time granularity. Synchronizing the speech analysis window and semantic analysis window allows for the initial quantization score of the speech capability dimension to be obtained through matching within the same temporal frame. Initial quantization score of semantic capability dimension ; Initial quantitative score for the voice ability dimension Initial quantization score of semantic capability dimension Perform modal normalization to map to the same standard range. The normalized score is obtained. , ; For the comprehensive capability dimension, scenario-related weights are introduced. , ,satisfy Construct a fusion model: ;in, For the first A comprehensive evaluation value integrating multiple capability dimensions. This serves as an identifier for the evaluation scenario type; The comprehensive ability dimension is integrated and evaluated based on preset threshold rules. Perform deviation calibration to generate the final independent quantitative evaluation results for each capability dimension.

9. The English oral proficiency multidimensional assessment method based on dual-engine collaboration according to claim 5, characterized in that, The speech ability dimension includes: speech expression dimension and speech fluency dimension; The semantic capability dimension includes: language structure dimension, professional vocabulary dimension, and scene understanding dimension; The comprehensive capability dimension includes: the interactive communication dimension, which is obtained by coupling and weighting the naturalness of tone and the stability of speech rate of the speech capability dimension with the politeness, timeliness of response and appropriateness of expression of the semantic capability dimension, and carries the alignment fusion and comprehensive quantitative evaluation of bimodal scores.

10. A computer-readable storage medium, characterized in that, The storage medium stores at least one executable instruction, which, when executed on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 9.