AI interactive classroom learning state evaluation method based on voice large model, computer program product and system

Through the AI ​​interactive classroom learning status evaluation method based on speech big model, combined with multi-dimensional voice scoring and multi-modal information integration, the problem of insufficient evaluation and insufficient multi-modal information integration in the online learning system is solved, and a more comprehensive learning quality evaluation and user experience improvement is achieved.

CN120564765APending Publication Date: 2025-08-29GUANGDONG JINHONG DIGITAL TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510891496.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The existing online learning system is limited to a single mode in terms of voice rating, and fails to fully consider multi-dimensional language features, resulting in insufficient objective and in-depth evaluation, and insufficient integration of multi-modal information, affecting user experience and teaching quality.

Method used

The AI ​​interactive classroom learning status evaluation method based on the speech model is adopted. Through multi-dimensional speech score construction and dynamic weight adjustment, combined with voice, visual and text information, the evaluation dimensions of intonation, rhythm and emotional expression are added, and standardized processing and timestamp alignment are carried out through the multi-modal fusion interface module to achieve synchronization and complementarity of multi-modal data.

Benefits of technology

It provides more detailed and comprehensive learning quality feedback, improves the teaching quality and user experience of the online education platform, and enhances the system's perception ability and the richness of the interactive scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564765A_ABST
    Figure CN120564765A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of online learning, in particular to an AI interactive classroom learning state evaluation method based on a voice large model, a computer program product and an online learning system. According to the AI interactive classroom learning state evaluation method based on the voice large model, multi-dimensional voice scoring comprises two dimensions of language and emotion, the former dimension covers indexes such as accuracy, fluency and integrity, the latter dimension comprises intonation, rhythm and emotion expression, the indexes are firstly scored in the scoring process, and the scoring result is obtained. Then, the ratio of each score value to a preset standard value is calculated, a total score is obtained according to a preset weight, and the weight of each dimension is adjusted according to a specific scene recognition result and the importance proportion thereof, so that the comprehensive evaluation of the learning state is enhanced, and more detailed feedback is provided by increasing the evaluated emotion dimension, and the learning efficiency is improved. And the user experience and the perception capability of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of online learning technology, and in particular to an AI interactive classroom learning status assessment method based on a large speech model, a computer program product, and an online learning system. Background Art

[0002] With the promotion and application of digital intelligent technologies such as big data, artificial intelligence, cloud computing, and blockchain, online learning, as a form of learning through the internet, has broken the time and space limitations of traditional teaching, allowing students to access rich course resources based on their needs. However, online learning and its related technologies still face some challenges:

[0003] Limitations of Voice Scoring Systems: Current mainstream voice scoring systems primarily rely on automatic speech recognition technology, focusing their evaluation on basic metrics such as accuracy, fluency, and completeness. These systems fail to fully consider deeper linguistic characteristics, resulting in an incomplete and in-depth assessment of voice interaction quality. Furthermore, existing scoring methods often rely on manual scoring, which is highly subjective and difficult to provide objective and authentic feedback on learning quality.

[0004] Insufficient multimodal information integration: Existing educational products and technical solutions are often limited to a single feedback mode, such as supporting only client-side voice input. This single-modality approach limits the richness and interactivity of the user experience and fails to fully utilize diverse sensory information to evaluate learning outcomes.

[0005] To address the above challenges, it is necessary to study a new method that can combine multi-dimensional information and achieve a comprehensive and objective evaluation of learning quality. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an AI interactive classroom learning status evaluation method based on a large speech model, a computer program product storing a computer program that implements the steps of the method when executed, and an online learning system. The AI ​​interactive classroom learning status evaluation method based on a large speech model can effectively integrate multiple perceptual information such as speech, vision and text in the AI ​​interactive classroom, thereby improving the teaching quality and user experience of the online education platform.

[0007] In order to solve the above technical problems, in a first aspect, the present invention provides an AI interactive classroom learning status assessment method based on a large speech model, comprising the following steps:

[0008] The multi-dimensional speech scoring step includes a language dimension and an emotional dimension. The language dimension includes multiple of the following speech indicators: accuracy, fluency, and completeness; the emotional dimension includes one or more of the following emotional indicators: intonation, rhythm, and emotional expression;

[0009] Scoring step: obtaining the scoring values ​​of the voice index and the emotional index of the current class respectively, calculating the ratio of the obtained scoring values ​​to the preset standard values ​​of each index, and calculating the total scoring value according to the preset weights;

[0010] The dynamic weight adjustment step identifies the current classroom scene, obtains the dimensions corresponding to the classroom scene and their importance ratios, and adjusts the weights of each dimension so that the sum of the weights of all dimensions is the preset value.

[0011] Furthermore, in the multi-dimensional speech scoring construction step, accuracy includes the degree of conformity of the user's pronunciation and grammar relative to the standard value. Specifically, the degree of conformity of the pronunciation relative to the standard value refers to the pronunciation error rate, and the degree of conformity of the grammar relative to the standard value refers to the number of grammatical errors and the number of missing key terms or vocabulary.

[0012] Furthermore, in the multi-dimensional speech scoring construction step, fluency includes the coherence and natural fluency of the user's language expression, the coherence refers to the number of abnormal pauses, and the natural fluency refers to the proportion of repeated content in the sentence and the fluctuation range of the speaking speed.

[0013] Furthermore, in the multi-dimensional speech scoring construction step, the completeness includes content coverage, specifically the repetition rate of the speech recognition content relative to the preset content.

[0014] Furthermore, in the multi-dimensional speech scoring construction step, intonation refers to: identifying pitch changes and judging the degree of compliance with preset change rules.

[0015] Furthermore, in the multi-dimensional speech scoring construction step, rhythm refers to: identifying the distribution of stress and pause positions, and judging the degree of compliance with preset change rules.

[0016] Furthermore, in the multi-dimensional voice scoring construction step, emotional expression refers to: obtaining the current classroom context, identifying the emotions conveyed by the user during the expression process, and judging the degree of conformity of the emotion with the context.

[0017] On the second aspect, a computer program product is also provided, which stores a computer program. When the computer program is executed by a processor, it can implement the steps of the above-mentioned AI interactive classroom learning status assessment method based on a large speech model.

[0018] On the third aspect, an online learning system is also provided, including a processor and a client and a server connected to the processor respectively, and also including the above-mentioned computer program product, the computer program on which can be executed by the processor.

[0019] Furthermore, it includes a multimodal fusion interface module, which performs standardized processing on data of different modalities; the data of different modalities include multiple types among voice, vision, and text; the standardized processing includes modal registration and timestamp alignment; the modal registration refers to: for an identified new data interface access request, calling a registration interface to identify the metadata of the access request, the metadata includes multiple types among modal type identifiers, data format requirements and processing function entries; the timestamp alignment refers to: timestamp confirmation of all accessed modal data so that the clock synchronization error of all data does not exceed a preset threshold.

[0020] The online learning system of the present invention executes the steps of the AI ​​interactive classroom learning status assessment method based on a large voice model. By adding assessment dimensions for intonation, rhythm, and emotional expression, it provides more detailed and comprehensive feedback, adds key factors affecting voice quality to the assessment, and improves the comprehensiveness of the scoring system. The online learning system supports richer interactive scenarios by integrating an interactive framework of voice, vision, and text information, allowing information from different modalities to complement each other, thereby enhancing the system's perception capabilities and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments.

[0022] Figure 1 This is a flowchart of the steps of the AI ​​interactive classroom learning status evaluation method based on the large voice model.

[0023] Figure 2 This is a schematic diagram of the system architecture of the AI ​​interactive classroom learning status assessment method based on the large voice model. DETAILED DESCRIPTION

[0024] The present invention is further described in detail below in conjunction with specific embodiments.

[0025] The online learning system of this embodiment includes a processor and a client and a server connected to the processor respectively. The learning object of the online learning system is oral practice. The server stores the content of all online courses. The user's client stores a computer program. The computer program is executed by the processor to achieve the following Figure 1The AI ​​interactive classroom learning status assessment method based on a large voice model adds assessment dimensions of intonation, rhythm, and emotional expression to the assessment process, improving the comprehensiveness of the scoring system.

[0026] The specific steps of the learning status evaluation method of this embodiment are described as follows.

[0027] The multi-dimensional speech scoring construction step includes a language dimension and an emotional dimension. The language dimension includes multiple of the following speech indicators: accuracy, fluency and completeness; the emotional dimension includes one or more of the following emotional indicators: intonation, rhythm, and emotional expression.

[0028] The main dimensions of the multi-dimensional speech scoring system constructed in this embodiment include accuracy, fluency and completeness in the language dimension, and intonation, rhythm and emotional expression in the emotional dimension. The specific goals and corresponding scoring criteria for each dimension are described below.

[0029] Accuracy includes the degree of conformity between the user's pronunciation and the grammar relative to the standard value. Specifically, the degree of conformity between the pronunciation and the standard value refers to the pronunciation error rate, and the degree of conformity between the grammar and the standard value refers to the number of grammatical errors and the number of missing key terms or vocabulary.

[0030] Specifically, the accuracy of the user's pronunciation, the correctness of grammatical structure, and the appropriateness of word choice are quantified. If the pronunciation error rate does not exceed 5%, the user's pronunciation is rated as excellent, with a score ranging from 90 to 100. If there is no more than one grammatical error in each sentence, the sentence's grammar is rated as good, with a score between 75 and 89. If key terms or vocabulary are omitted, the corresponding points will be directly deducted from the total score.

[0031] Fluency includes the coherence and natural fluency of the user's language expression. The coherence refers to the number of abnormal pauses, and the natural fluency refers to the proportion of repeated content in a sentence and the fluctuation range of speaking speed.

[0032] Specifically, the coherence and natural fluency of language expression are assessed. Abnormal pauses should be limited to two per minute; repetitions within sentences should not exceed 3%; and overall speaking speed should fluctuate within a reasonable range, with the assessment standard being a fluctuation of no more than ±15%.

[0033] Completeness includes content coverage, specifically the repetition rate of the speech recognition content relative to the preset content.

[0034] Specifically, the user's answer should be evaluated for comprehensive coverage of the required content, ensuring that no information is omitted while avoiding unnecessary redundancy. The user's answer must cover at least 80% of the core proposition. No more than two instances of supporting details should be omitted.

[0035] Intonation refers to: identifying pitch changes and judging the degree of compliance with preset change rules.

[0036] Specifically, analyze whether pitch changes conform to the basic laws of language communication, especially whether the pitch changes correctly between questions and statements. At the end of questions, the pitch should rise by at least 20Hz. For statements, the overall pitch fluctuation should remain within a small range, generally no more than 30Hz.

[0037] Rhythm refers to: identifying the distribution of stress and pause positions, and judging the degree of conformity with preset change rules.

[0038] Specifically, the distribution of stress and pause placement within the speech, along with their consistency with the overall rhythmic pattern, should be assessed. The coefficient of variation of the stress interval should be less than or equal to 0.25. When emphasizing a particular section, the speech rate should be reduced by at least 20%.

[0039] Emotional expression refers to: obtaining the current classroom context, identifying the emotions conveyed by the user during the expression process, and judging the degree to which the emotion conforms to the context.

[0040] Specifically, the system detects and evaluates whether the emotion conveyed by the user is appropriate for the current context and whether the intensity of the emotion is within a preset range. The emotion type must match the expected scenario, and the confidence level of the emotion intensity must be at least 0.7.

[0041] Scoring steps: obtain the scoring values ​​of the voice index and emotional index of the current class respectively, calculate the ratio of the obtained scoring values ​​to the preset standard values ​​of each index, and calculate the total scoring value according to the preset weights.

[0042] Set weights for each dimension and calculate the total score:

[0043]

[0044] ∑w i =1

[0045] i is the dimension of the current class.

[0046] The dynamic weight adjustment step identifies the current classroom scene, obtains the dimensions corresponding to the classroom scene and their importance ratios, and adjusts the weights of each dimension so that the sum of the weights of all dimensions is the preset value. The weights of different dimensions are dynamically adjusted through scene recognition and rule engines. For specific teaching scenarios (such as English speeches or word dictation), the weight values ​​of each dimension are automatically adjusted according to preset rules. For example, in a speech scenario, the weight of "emotional expression" is increased to 30%. At the same time, the model is trained by collecting teacher manual scoring data through machine learning methods to further optimize the weight setting.

[0047] The AI ​​interactive classroom learning status assessment method based on a large speech model in this embodiment provides more detailed and comprehensive feedback by adding assessment dimensions for intonation, rhythm, and emotional expression, adds key factors affecting speech quality to the assessment, and improves the comprehensiveness of the scoring system.

[0048] The online learning system of this embodiment includes a multimodal fusion interface module, which performs standardized processing on data of different modalities; the data of different modalities include multiple types of voice, vision, and text; the standardized processing includes modality registration and timestamp alignment; the modality registration refers to: for an identified new data interface access request, calling a registration interface to identify the metadata of the access request, the metadata including multiple types of modality type identifiers, data format requirements, and processing function entries; the timestamp alignment refers to: timestamp confirmation of all accessed modality data so that the clock synchronization error of all data does not exceed a preset threshold.

[0049] Modality registration in this embodiment refers to a modular registration mechanism. New modality processors, such as an electroencephalogram (EEG) analyzer, are created on demand by calling the framework's registration interface and submitting relevant metadata, such as the modality type identifier, data format requirements, and processing function entry points. Once the framework verifies the compatibility of the new module, it is loaded into the runtime container and the modality routing table is updated to facilitate subsequent calls.

[0050] The timestamp alignment module of this embodiment effectively aligns data from different modalities in the time dimension through the dynamic time warping (DTW) algorithm, so that all modal data follow a unified timestamp rule, ensuring that data from different sources can be accurately aligned on the time axis, and the clock synchronization error of all devices shall not exceed 50 milliseconds.

[0051] Specifically, after receiving timestamped speech clips and visual keyframes, the system extracts corresponding feature sequences and calculates the similarity between them. Based on this similarity, the DTW algorithm can find the optimal time alignment path to associate data from different modalities.

[0052] In addition to temporal alignment, the online learning system of this embodiment also includes a spatial alignment module, which transforms the acquired visual data from the camera coordinate system—with the upper left corner of the blackboard as the origin—to the blackboard coordinate system. Specifically, a 3×3 homography matrix H is established to describe the mapping relationship from one coordinate system to another. Through this transformation, the visual data is accurately positioned on the blackboard plane.

[0053] The standardized processing for constructing a multimodal fusion interface in this embodiment also includes a step of pre-processing multimodal data such as speech, vision, and text by the acquisition device, as described in detail below.

[0054] The speech preprocessing step includes marking silent segments, collecting speech data through a microphone array with a signal-to-noise ratio greater than 30dB, performing voice activity detection (VAD) on the collected speech data, performing streaming feature extraction if the detection result is a valid speech segment, and discarding silent frames if the detection result is invalid speech.

[0055] Visual preprocessing steps include face detection and key point tracking. A camera with a resolution of at least 720p captures the user's facial expressions and gestures, and outputs key point coordinates at a rate of at least 25 frames per second to capture facial landmarks and gesture key points. Precise projection transformation technology is used to ensure that the conversion error from the camera coordinate system to the blackboard coordinate system does not exceed 5 pixels.

[0056] Text preprocessing includes the recognition of text and its position bounding box information. The information on the OCR (optical character recognition) blackboard or electronic whiteboard is obtained in JSON format through a document camera or electronic whiteboard. Perspective correction and text area segmentation are used on the obtained information to generate position bounding box information to define the text range, and a general text recognition program is called to recognize the text information.

[0057] The online learning system of the present invention supports richer interactive scenarios by integrating an interactive framework of voice, vision and text information, so that information of different modalities can complement each other, thereby enhancing the system's perception ability and user experience.

[0058] The online learning system of the present invention also includes an exception handling module to switch to handle situations where modal data is lost or conflicting. If the loss rate of visual data exceeds 50%, the system will automatically switch to pure speech enhancement mode and reduce the weight of other modalities (such as gestures) accordingly. In addition, a confidence threshold is set. Only when the confidence of each modal output exceeds 0.7 will it be included in the final decision. Such a design not only improves the robustness of the system, but also ensures that basic functional operation can be maintained even in the event of partial modal failure.

[0059] This embodiment implements the above-mentioned AI interactive classroom learning status assessment method based on a large voice model through a computer program. The computer program is stored in a computer program product and is executed by a computer processor to implement the above-mentioned AI interactive classroom learning status assessment method based on a large voice model. The online learning system embodiment described above is only schematic, wherein the modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art can understand and implement it without paying any creative labor.

[0060] Finally, it should be noted that the AI ​​interactive classroom learning status assessment method based on a large speech model disclosed in the embodiment of the present invention only discloses a preferred embodiment of the present invention, which is only used to illustrate the technical solution of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, ordinary technicians in this field should understand that it is still possible to modify the technical solutions recorded in the aforementioned embodiments, or to make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for evaluating AI interactive classroom learning status based on a large speech model, characterized by: The following steps are involved: The multi-dimensional speech scoring step includes a language dimension and an emotional dimension. The language dimension includes multiple of the following speech indicators: accuracy, fluency, and completeness; the emotional dimension includes one or more of the following emotional indicators: intonation, rhythm, and emotional expression; Scoring step: obtaining the scoring values ​​of the voice index and the emotional index of the current class respectively, calculating the ratio of the obtained scoring values ​​to the preset standard values ​​of each index, and calculating the total scoring value according to the preset weights; The dynamic weight adjustment step identifies the current classroom scene, obtains the dimensions corresponding to the classroom scene and their importance ratios, and adjusts the weights of each dimension so that the sum of the weights of all dimensions is the preset value.

2. The AI ​​interactive classroom learning status assessment method based on a large speech model as claimed in claim 1 is characterized in that: In the multi-dimensional speech scoring construction step, accuracy includes the degree of conformity of the user's pronunciation and grammar relative to the standard value. Specifically, the degree of conformity of the pronunciation relative to the standard value refers to the pronunciation error rate, and the degree of conformity of the grammar relative to the standard value refers to the number of grammatical errors and the number of missing key terms or vocabulary.

3. The AI ​​interactive classroom learning status assessment method based on a large speech model as claimed in claim 1 is characterized in that: In the multi-dimensional speech scoring construction step, fluency includes the coherence and natural fluency of the user's language expression, the coherence refers to the number of abnormal pauses, and the natural fluency refers to the proportion of repeated content in the sentence and the fluctuation range of the speaking speed.

4. The AI ​​interactive classroom learning status assessment method based on a large speech model as claimed in claim 1 is characterized in that: In the multi-dimensional speech scoring construction step, the completeness includes content coverage, specifically the repetition rate of the speech recognition content relative to the preset content.

5. The AI ​​interactive classroom learning status assessment method based on a large speech model as claimed in claim 1 is characterized in that: In the multi-dimensional speech scoring construction step, intonation refers to: identifying pitch changes and judging the degree of compliance with preset change rules.

6. The AI ​​interactive classroom learning status assessment method based on a large speech model as claimed in claim 1 is characterized in that: In the multi-dimensional speech scoring construction step, rhythm refers to: identifying the distribution of stress and pause positions, and judging the degree of compliance with preset change rules.

7. The AI ​​interactive classroom learning status assessment method based on a large speech model as claimed in claim 1 is characterized in that: In the multi-dimensional voice scoring construction step, emotional expression refers to: obtaining the current classroom context, identifying the emotions conveyed by the user during the expression process, and judging the degree of conformity of the emotion with the context.

8. A computer program product storing a computer program, characterized in that: When the computer program is executed by a processor, it can implement the steps of the AI ​​interactive classroom learning status assessment method based on a large speech model according to any one of claims 1 to 7.

9. Online learning system, characterized by, The computer program product comprises a processor, a client and a server respectively connected to the processor, and also comprises the computer program product according to claim 8, wherein the computer program on the computer program product can be executed by the processor.

10. The online learning system according to claim 9, wherein: The system includes a multimodal fusion interface module that standardizes data in different modalities, including voice, visual, and text. The standardization includes modality registration and timestamp alignment. Modality registration involves calling a registration interface to identify metadata associated with a new data interface access request, including a modality type identifier, data format requirements, and a processing function entry. The timestamp alignment refers to: performing timestamp confirmation on all received modal data so that the clock synchronization error of all data does not exceed a preset threshold.

Citation Information

Patent Citations

  • Objective standard based automatic oral evaluation system

    CN101826263A

  • English phonetic pronunciation quality evaluation system with emotion recognition function and method thereof

    CN104050965A

  • Speech comprehensive evaluation method and device, equipment and storage medium

    CN117238321A

  • Multi-modal classroom emotion recognition method and system based on modal adaptive learning

    CN119418725A

  • Method and device for automatic evaluation of Korean speaking ability based artificial intelligence

    KR102793891B1