Classroom teaching real-time intelligent interactive communication system based on large language model

Through the classroom teaching system with multimodal data fusion and edge computing optimization, the problems of unidirectional communication, insufficient intelligence and poor real-time performance of the existing classroom teaching interaction system are solved, and real-time interaction with low latency and high accuracy are achieved.

CN120336484APending Publication Date: 2025-07-18JIUTU TECH (NANJING) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510444935.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing classroom teaching interaction system has problems such as one-way communication, insufficient intelligence, poor real-time performance and lack of feedback mechanisms.

Method used

A real-time intelligent interactive communication system for classroom teaching based on large language models is adopted, and voice, text and gesture data are collected in real time through multimodal sensors, combined with lightweight LLM (DistilBERT+GPT-3.5) to perform localized real-time inference at edge computing nodes, cloud-based collaborative modules perform model training and privacy data processing, and interactive terminals provide real-time Q&A and personalized learning suggestions.

Benefits of technology

Real-time interaction with low latency is achieved, improving student engagement and complex problem solving rates, and ensuring data privacy compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336484A_ABST
    Figure CN120336484A_ABST
Patent Text Reader

Abstract

The invention discloses a classroom teaching real-time intelligent interactive communication system based on a large language model, which realizes low-delay and high-precision classroom real-time questions and answers through multi-modal data fusion, edge calculation optimization and a dynamic feedback mechanism, has a self-adaptive teaching content generation capability, and is high in practicability. And the blank in the aspects of real-time performance, multi-modal processing and personalized feedback in the prior art is filled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent teaching, and particularly to a real-time intelligent interactive communication system for classroom teaching based on large language models, which is applicable to improving the personalization and real-time feedback efficiency of classroom teaching. Background Art

[0002] The existing classroom teaching interactive systems have the following limitations:

[0003] 1. Predominantly one-way communication: Traditional systems rely on the one-way output of teachers. Students need to raise their hands or use fixed terminals to ask questions, which is inefficient and easily disrupts the teaching rhythm.

[0004] 2. Insufficient intelligence: Most existing intelligent question-and-answer systems are based on pre-set knowledge bases or simple semantic matching, unable to deeply understand complex questions and lacking the ability to fuse and process multi-modal data (voice, text, gestures).

[0005] 3. Poor real-time performance: Most systems use cloud-based LLM inference, resulting in high latency and difficult to meet the requirements of instant classroom interaction.

[0006] 4. Lack of feedback mechanism: Unable to dynamically adjust teaching content according to the real-time status of students (such as confusion level, participation degree). Summary of the Invention

[0007] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title. However, such simplifications or omissions shall not be used to limit the scope of the present invention.

[0008] In view of the problems existing in the above-mentioned existing classroom teaching interactive systems, the present invention is proposed.

[0009] Therefore, the technical problem solved by the present invention is to solve the problems of predominantly one-way communication, insufficient intelligence, poor real-time performance, and lack of feedback mechanism in the existing classroom teaching interactive systems.

[0010] To solve the above technical problem, the present invention provides the following technical solution: A real-time intelligent interactive communication system for classroom teaching based on large language models, including the following architecture components: Data acquisition layer: Real-time collection of voice, text, gestures, and blackboard writing content through multi-modal sensors (microphone array, camera, electronic whiteboard); Edge computing node: Deployment of a lightweight LLM (DistilBERT+GPT-3.5 hybrid architecture) and a multi-modal fusion module to achieve local real-time inference; Cloud collaboration module: Used for continuous model training, cross-class knowledge sharing, and privacy data desensitization; Interactive terminal: The student / teacher terminal displays real-time Q&A, knowledge point association graphs, and personalized learning suggestions.

[0011] As a preferred solution of the real-time intelligent interactive communication system for classroom teaching based on large language models of the present invention, wherein: the implementation of local real-time inference by the lightweight LLM and multimodal fusion module specifically includes the following steps:

[0012] S1: Speech-Text Spatiotemporal Alignment

[0013] Input: Speech signal stream S = {s1, s2,..., s m}, text sequence T = {t1, t2,..., t n};

[0014] Output: Text-speech mapping relationship π after time alignment * ;

[0015] Wherein:

[0016]

[0017] Wherein, s i ∈R ds represents the MFCC (Mel Frequency Cepstral Coefficient) feature vector of speech frame i; t j

[0018] ∈R dt represents the BERT embedding vector of text word j; λ ∈ [0, 1] represents the path continuity constraint weight, defined as 0.3; Penalty(π) = ∑ (i,j)∈π |i - j| represents the penalty for non-monotonic alignment paths;

[0019] S2: Multimodal Attention Fusion

[0020] Input:

[0021] Text embedding matrix H t ∈R n×dt

[0022] Image embedding matrix H v ∈R k×dv (ResNet-50 features of blackboard writing / gestures);

[0023] Cross-modal attention calculation:

[0024]

[0025] Wherein, d a represents the attention dimension; W q , W k , W v represent trainable linear transformation matrices;

[0026] S3: Dynamic Model Distillation

[0027] Input: Device computing power evaluation metrics C = {CPU utilization, memory occupancy, Δt allowed}

[0028] Model switching rule:

[0029]

[0030] where, Δt allowed represents the maximum allowed response latency; θ = f(C) = 0.7 * CPU idle rate + 0.3 * remaining memory ratio;

[0031] S4: Mixed Precision Caching Mechanism

[0032] Cache hit condition:

[0033]

[0034] Cache update policy:

[0035] c new = α·c old +(1 - α)·e q (α = 0.9)

[0036] where, e q ∈R 768 represents the Sentence - BERT embedding of question q; α represents the cache vector smoothing coefficient.

[0037] As a preferred solution of the real - time intelligent interactive communication system for classroom teaching based on large - language models of the present invention, it further includes S5 Priority Score Calculation;

[0038] where, the scoring formula is specifically:

[0039] Priority i = α·U i +β·(1 - P i )+γ·C i

[0040] where, U i represents the urgency; P i represents the historical performance; C i represents the semantic complexity.

[0041] As a preferred solution of the real - time intelligent interactive communication system for classroom teaching based on large - language models of the present invention, it further includes S6 Student Perplexity Detection, specifically including:

[0042] Input: Facial expression feature f face ∈R128 , the number of follow-up questions N follow ;

[0043] Detection model:

[0044] c t = σ(w1·N follow + w2·||f face - f confused ||)

[0045] where f confused represents the benchmark feature vector of the predefined confused expression; w1 = 0.6, w2 = 0.4, representing the weight coefficients; σ represents the Sigmoid function;

[0046] Among them:

[0047] σ(x) = 1 / (1 + e -x ).

[0048] As a preferred solution of the real-time intelligent interactive communication system for classroom teaching based on the large language model of the present invention, wherein: it further includes dynamic adjustment of S7 generation parameters;

[0049] Among them, the temperature coefficient update rule:

[0050] T new = T base ·(1 + η·c t )

[0051] where T base represents the base temperature coefficient, defined as 0.7; η = 0.5, representing the perplexity influence factor.

[0052] The present invention provides a real-time intelligent interactive communication system for classroom teaching based on the large language model, having the following

[0053] beneficial effects:

[0054] 1. The average response delay is reduced from 2.3 s of the traditional system to 0.8 s;

[0055] 2. The student participation rate is increased by 65%, and the complex problem-solving rate is increased by 40%;

[0056] 3. 100% data privacy compliance is achieved through localization processing. Description of the Drawings

[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:

[0058] Figure 1 This is the overall method flowchart for the lightweight LLM and multi-modal fusion module provided by the present invention to achieve local real-time inference. Detailed implementation manners

[0059] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the drawings of the specification. Obviously, the described embodiments are some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0060] Existing classroom teaching interaction systems have the following limitations:

[0061] 1. Mainly one-way communication: Traditional systems rely on the one-way output of teachers. Students need to raise their hands or use fixed terminals to ask questions, which is inefficient and easily interrupts the teaching rhythm.

[0062] 2. Insufficient intelligence: Existing intelligent question-and-answer systems are mostly based on preset knowledge bases or simple semantic matching, unable to deeply understand complex questions, and lacking the ability to fuse and process multi-modal data (voice, text, gesture).

[0063] 3. Poor real-time performance: Most systems use cloud LLM inference, with high latency and difficult to meet the immediate interaction needs in the classroom.

[0064] 4. Lack of feedback mechanism: Unable to dynamically adjust teaching content according to the real-time status of students (such as confusion level, participation degree).

[0065] Therefore, the present invention provides a real-time intelligent interactive communication system for classroom teaching based on a large language model, including the following architecture components:

[0066] Data acquisition layer: Real-time collect voice, text, gesture, and blackboard writing content through multi-modal sensors (microphone array, camera, electronic whiteboard).

[0067] Edge computing node: Deploy a lightweight LLM (DistilBERT+GPT-3.5 hybrid architecture) and a multi-modal fusion module to achieve local real-time inference.

[0068] Cloud Collaboration Module: Used for continuous model training, cross-class knowledge sharing, and privacy data desensitization;

[0069] Interactive Terminal: The student / teacher terminal displays real-time Q&A, knowledge point association graphs, and personalized learning suggestions.

[0070] It should be noted that:

[0071] 1. Data Acquisition Layer

[0072] Objective: To collect multi-modal data (such as voice, text, images, blackboard writing content, etc.) in the classroom in real-time and provide input for subsequent multi-modal fusion and reasoning.

[0073] Technical Details:

[0074] ① Voice Acquisition:

[0075] Equipment: Deploy a microphone array (such as a circular array composed of 8 - 16 microphones), supporting 360-degree voice acquisition.

[0076] Technology: Use the Beamforming algorithm to extract the target voice signal and suppress environmental noise and reverberation.

[0077] Implementation:

[0078] Adopt an array signal processing library (such as Python's Sounddevice or C++'s PortAudio) for real-time capture and preprocessing of voice signals.

[0079] Use a deep learning-based voice enhancement model (such as Wav2Vec 2.0) to denoise the audio.

[0080] The audio sampling rate is set to 16kHz to meet the requirements of clear voice acquisition.

[0081] ② Text Acquisition:

[0082] Source: Text questions input by students / teachers through the terminal, class notes, or instant messages.

[0083] Implementation: Obtain text data in real-time through the API interface and use natural language processing (NLP) tools (such as spaCy) for syntactic analysis and keyword extraction.

[0084] ③ Image Acquisition:

[0085] Equipment: Cameras deployed in the classroom (at least 3 to cover the entire classroom perspective) and the built-in camera of the electronic whiteboard.

[0086] Technology: Use OpenCV or TensorFlow to detect dynamic changes in students' gestures (such as raising hands and writing) and the content of the blackboard writing.

[0087] Implementation:

[0088] For gesture detection, adopt an object detection model based on YOLOv5 to identify the actions of raising hands or writing.

[0089] For the content of the blackboard writing, use OCR technology (such as Tesseract) to extract handwritten text and combine it with a handwriting recognition model (such as IAM) to analyze the structure of the blackboard writing.

[0090] ④ Environmental information collection:

[0091] Equipment: Deploy environmental sensors (such as temperature, humidity, and light sensors) to monitor the classroom environment status in real time.

[0092] Implementation: Transmit environmental data to the edge computing node through the MQTT protocol for dynamically adjusting the interaction strategy.

[0093] 2. Edge computing node

[0094] Goal: Complete multimodal data fusion and lightweight LLM inference locally to ensure real-time performance and privacy protection.

[0095] Technical details:

[0096] ① Lightweight LLM model:

[0097] Architecture: Adopt a hybrid architecture that combines DistilBERT (a pre-trained language representation model) and GPT-3.5 (a generative language model).

[0098] DistilBERT: Used for text understanding and keyword extraction.

[0099] GPT-3.5: Used for generating natural language answers.

[0100] Optimization: Use model distillation technology (such as knowledge distillation) to compress GPT-3.5 to a size suitable for edge devices (such as 700M parameters).

[0101] Implementation:

[0102] Use the HuggingFace library to load the pre-trained model and perform quantization through the ONNX format (such as mixed FP16 / FP32 calculation) to reduce the computational overhead.

[0103] Deploy on edge devices (such as NVIDIA Jetson Nano or Intel NUC) to ensure processing at least 2 questions per second.

[0104] ② Multimodal Fusion Module:

[0105] Implementation:

[0106] Encode speech, text, and image data into embedding vectors respectively (such as BERT embeddings, ResNet-50 image features).

[0107] Use the attention mechanism (such as Cross-Attention) to fuse multimodal data into a unified embedding vector (vector dimension: 512).

[0108] Implement the multimodal fusion network using the PyTorch or TensorFlow framework.

[0109] ③ Real-time Inference:

[0110] Process:

[0111] Receive multimodal data (e.g., student questions + blackboard writing content).

[0112] Align speech and text through the DTW algorithm (such as matching the timestamps corresponding to speech signals with the text content).

[0113] Generate answers using a lightweight LLM (such as "The eddy current phenomenon of electromagnetic induction can be understood in the following way...").

[0114] Latency Optimization:

[0115] Use local caching technology to store frequently asked questions and their answers, reducing repeated calculations.

[0116] Optimize computing resource allocation through dynamic task scheduling (such as Docker containerization technology).

[0117] 3. Cloud Collaboration Module

[0118] Goal: Achieve continuous training of the model, cross-class knowledge sharing, and privacy data desensitization, improving the overall system performance and knowledge coverage.

[0119] Technical Details:

[0120] ① Model Continuous Training:

[0121] Data: Classroom interaction data collected from different classes (such as questions, answers, blackboard writing content).

[0122] Technology: Use Federated Learning to update the model, ensuring data privacy.

[0123] Implementation: Each edge node (school or class) updates the locally trained model and uploads the encrypted model parameters to the cloud. The cloud server aggregates these parameters to generate a global model and then distributes it back to the edge nodes.

[0124] Frameworks used: PyTorchLightning or TensorFlow Federated.

[0125] ② Cross-class knowledge sharing:

[0126] Implementation: By constructing a knowledge graph (such as using the Ubergraph tool), common problems and knowledge points in different classes are associated to form a global knowledge graph.

[0127] Optimization: Using graph embedding techniques (such as Node2Vec) to enhance the semantic associations between knowledge points and improve the accuracy of cross-scenario question answering.

[0128] ③ Privacy data desensitization:

[0129] Techniques: Adopting data encryption (such as AES-256) and anonymization processing (such as randomly replacing student names) to protect sensitive information.

[0130] Implementation: Using the cryptography library in Python for data encryption and ensuring strict access control (such as role-based access control, RBAC) when storing in the cloud.

[0131] 4. Interactive terminal

[0132] Goal: Provide students and teachers with real-time question answering, knowledge point association graphs, and personalized learning suggestions to improve classroom interaction efficiency.

[0133] Technical details:

[0134] ① Real-time question answering:

[0135] Implementation: Maintain real-time communication with edge computing nodes through the WebSocket protocol and quickly display the answers generated by the LLM.

[0136] Techniques: Use a front-end framework (such as React or Vue.js) to build a responsive interface that supports dynamic updates (such as scrolling the answers on the screen).

[0137] ② Knowledge point association graph:

[0138] Techniques: Use a graph database (such as Neo4j) to build the association relationships between knowledge points and dynamically display them through a network visualization tool (such as Gexf.js).

[0139] Implementation: When a student asks a question, the system displays in real time the knowledge points related to the question and their contextual relationships to help the student build a knowledge network.

[0140] ③ Personalized learning suggestions:

[0141] Algorithm: Dynamically generate a learning path based on the student's real-time performance (such as answer accuracy and perplexity).

[0142] Use recommendation system techniques (such as content-based recommendation or collaborative filtering) to recommend relevant practice questions, video tutorials, etc. to the student.

[0143] Implementation: Call the resources of an online learning platform (such as Khan Academy) through an API and display the recommended content on the student terminal in real time.

[0144] Furthermore, refer to Figure 1 , the implementation of local real-time inference for the lightweight LLM and multi-modal fusion module specifically includes the following steps:

[0145] S1: Speech-Text Spatiotemporal Alignment

[0146] Input: Speech signal stream S = {s1, s2,..., s m}, text sequence T = {t1, t2,..., t n};

[0147] Output: Time-aligned text-speech mapping relationship π * ;

[0148] Among them, the dynamic time warping (DTW) formula:

[0149]

[0150] Among them, s i ∈R ds , represents the MFCC (Mel Frequency Cepstral Coefficient) feature vector of the i-th speech frame. MFCC is a feature extraction method for speech signal processing, reflecting the time-frequency characteristics of speech, and is usually used for speech recognition and classification; t j ∈R dt , represents the BERT embedding vector of the j-th text word. BERT is a pre-trained language representation model that can convert text words into high-dimensional vectors and capture language context and semantic information; λ ∈ [0, 1], represents the path continuity constraint weight, defined as 0.3. This parameter is used to balance the weights of the distance term between speech and text features and the path penalty term. For example, λ = 0.3 means that the importance of path penalty is relatively low, and more attention is paid to the distance alignment of features; Penalty(π) = ∑ (i,j)∈π|i - j| represents the penalty for non - monotonic alignment paths. This function calculates the sum of the absolute values of the differences between the speech frame indices and text word indices in the alignment path, preventing excessive path jumps and maintaining the continuity of alignment.

[0151] Physical meaning: By minimizing the Euclidean distance between speech frames and text words and constraining the continuity of the alignment path, it solves the timing misalignment problem caused by speed differences. This formula ensures the temporal alignment of speech and text by minimizing the total Euclidean distance between speech frames and text words and combining a path penalty term. The path penalty term reduces non - monotonic jumps in the alignment process, thus more accurately reflecting the corresponding relationship between speech and text back and forth. For example, when a student asks a question, even if his speech speed is fast or slow, the system can accurately match the speech signal with the relevant text content, improving subsequent reasoning and answering effects.

[0152] S2: Multi - modal attention fusion

[0153] Input:

[0154] Text embedding matrix H t ∈R n×dt

[0155] Image embedding matrix H v ∈R k×dv (ResNet - 50 features of the blackboard writing / gestures);

[0156] Cross - modal attention calculation:

[0157]

[0158] Among them, d a represents the attention dimension (default 256); W q , W k , W v represents a trainable linear transformation matrix;

[0159] Physical meaning: By attentively weighting the image features with the text, it extracts visual information related to the current semantics. For example, when a student asks "electromagnetic induction", the system automatically focuses on the magnetic induction line diagram on the blackboard.

[0160] It should be noted that

[0161] ① In the text embedding matrix H t ∈R n×dt , n is the length of the text sequence (i.e., how many words or word segments); d t is the embedding dimension of each text word. For example, when using BERT embeddings, d t is usually 768.

[0162] Example: Suppose there is a text with 5 words and the embedding dimension of each word is 768, then H t is a 5×768 matrix.

[0163] ② Image embedding matrix H v ∈R k×dv , where k is the number of features in the image (e.g., the size of the feature map output by the last convolutional layer of ResNet-50); d v is the dimension of each image feature. For example, the feature dimension of ResNet-50 is usually 2048.

[0164] Example: For an image with ResNet-50 features of 7×7×2048, then H v can be flattened into a 49×2048 matrix.

[0165] ③ Query matrix: Q = H t W q

[0166] Wq ∈ R dt×da : A linear transformation matrix that maps the text embedding to the attention dimension.

[0167] Example: If d t = 768 and d a = 256, W q is a 768×256 matrix.

[0168] Then Q will be an n×256 matrix.

[0169] ④ Key matrix: K = H v W k

[0170] W k ∈R dv×da : A linear transformation matrix that maps the image embedding to the attention dimension.

[0171] Example: If d v = 2048 and d a = 256, Wk is a 2048×256 matrix.

[0172] Then K will be a k×256 matrix.

[0173] ⑤ Value matrix: V = H v W v

[0174] W v ∈R dv×da : Also a linear transformation matrix that maps the image embedding to the attention dimension.

[0175] Example: same as W k has the same structure as W v is also a 2048×256 matrix, and V is also a k×256 matrix.

[0176] ⑥ QK T : The similarity matrix between the query and the key, with a size of n×k.

[0177] Softmax operation: Converts the similarity matrix into a probability distribution such that the sum of probabilities in each row is 1.

[0178] Finally, the final attention-weighted feature is obtained by multiplying by V).

[0179] Matrix example:

[0180] Assumption:

[0181] H t is a 5×768 matrix (5 words, each word with a 768-dimensional embedding).

[0182] H v is a 49×2048 matrix (a 7×7 feature map, each feature with 2048 dimensions).

[0183] W q is a 768×256 matrix.

[0184] W k and W v are both 2048×256 matrices.

[0185] Calculation process:

[0186] Q = H t × W q : The result is a 5×256 matrix.

[0187] K = H v × W k : The result is a 49×256 matrix.

[0188] V = H v × W v : The result is also a 49×256 matrix.

[0189] QK T : A 5×49 matrix, and the (i, j)-th element represents the correlation between the i-th word in the text and the j-th feature in the image.

[0190] After Softmax, each text word will get a 49-dimensional probability vector, indicating which parts of the image features are most relevant to the current text.

[0191] Finally, multiply by V to obtain the attention-weighted image features of 5×256, that is, each text word can obtain a weighted image feature representation.

[0192] S3: Dynamic Model Distillation

[0193] Input: Device computing power evaluation metrics C = {CPU utilization, memory occupancy, Δt allowed}

[0194] Model switching rule:

[0195]

[0196] Among them, Δt allowed represents the maximum allowed response delay (unit: second), which is set to 1.2 seconds in the classroom scenario to ensure real-time interaction; θ = f(C) = 0.7 * CPU idle rate + 0.3 * remaining memory ratio;

[0197] Physical meaning: On the premise of ensuring the delay constraint, dynamically select the model scale according to the device resources (the number of parameters of the Base model is 500M, and the Large model is 1.5B). The dynamic distillation mechanism intelligently selects the model suitable for the current computing power by monitoring the device resources in real time, ensuring that the system can operate efficiently in different environments. For example, when the device resources are tense, switch to the Base model to reduce the computational overhead and avoid response delays caused by insufficient resources; while when the resources are sufficient, select the Large model to provide better inference results.

[0198] It should be noted that:

[0199] ① CPU idle rate: 1 - CPU utilization, reflecting whether the computing resources of the device are sufficient.

[0200] Remaining memory ratio: free memory / total memory, measuring the available memory resources of the system.

[0201] ② According to the comparison between Δt allowed and θ, determine whether to use the Base model (lightweight) or the Large model (heavyweight).

[0202] M base : The number of parameters is 500M, with fast inference speed, suitable for scenarios with tight resources.

[0203] M large : The number of parameters is 1.5B, with more accurate inference, suitable for scenarios with sufficient resources.

[0204] S4: Mixed Precision Caching Mechanism

[0205] Cache hit condition:

[0206]

[0207] Among them, e q is the Sentence-BERT embedding vector of the current question q, and e ci is the embedding vector of the question c i in the cache.

[0208] Threshold τ:

[0209] When sim(q, c i ) > τ, it is regarded as a cache hit. The default value τ = 0.85 means that only when the similarity exceeds 85% is it considered a cache hit.

[0210] Cache update strategy:

[0211] c new = α · c old + (1 - α) · e q (α = 0.9)

[0212] Among them, e q ∈ R 768 represents the Sentence-BERT embedding of the question q; α represents the cache vector smoothing coefficient.

[0213] Among them,

[0214] c new : The updated cache vector.

[0215] c old : The cache vector before update.

[0216] α = 0.9: The smoothing coefficient, which controls the weight ratio between the old cache and the new data.

[0217] Smoothing coefficient α:

[0218] α = 0.9 means that the new cache vector retains more content of the old cache while appropriately integrating the new embedding vector. This strategy helps to maintain the stability of the cache, prevent drastic changes, and gradually introduce the role of new data.

[0219] Parameter explanation:

[0220] e q : The Sentence-BERT embedding vector of the question q, with a dimension of 768. Sentence-BERT can map the entire sentence to a high-dimensional vector, reflecting the semantic information of the sentence.

[0221] α: The smoothing coefficient, used to control the smoothness of cache updates. α = 0.9 indicates that the weight of the old cache is larger, maintaining the stability of the system.

[0222] τ: Similarity threshold, used to determine cache hits. τ = 0.85 ensures the accuracy of cache hits.

[0223] Physical meaning: By matching high-frequency questions through cosine similarity, the number of LLM calls is reduced by more than 60%. When a new question enters the system, it first checks if there are similar questions and their corresponding answers in the cache. If the similarity exceeds τ, the question is considered cached, and the cached result is directly used to avoid repeated calls to the LLM for reasoning. At the same time, the cache update strategy ensures that the cache vectors are gradually updated over time, always maintaining relevance to the latest data and reducing cache ineffectiveness.

[0224] Furthermore, it also includes the calculation of S5 priority scores;

[0225] Among them, the scoring formula is specifically:

[0226] Priority i = α·U i + β·(1 - P i ) + γ·C i

[0227] Among them, U i represents the urgency; P i represents the historical performance; C i represents the semantic complexity.

[0228] Among them:

[0229] Urgency:

[0230]

[0231] e topic : Embedding vector of the current teaching topic, such as "electromagnetic induction"; e q : Embedding vector of the question q, generated using a pre-trained language model (such as BERT).

[0232] Physical meaning: Calculate the cosine similarity between the question q and the current teaching topic, representing the urgency of the question.

[0233] Performance:

[0234] (Normalized to [0, 1])

[0235] Among them, CorrectRate i : Answer correct rate of student i (normalized to [0, 1]).

[0236] ResponseSpeed i: The question response speed of student i (normalized to [0, 1]).

[0237] Physical meaning: Reflects the historical performance of student i. Students with high accuracy and fast response have high historical performance scores.

[0238] Semantic complexity (Complexity):

[0239] (Information entropy based on n-gram language model)

[0240] Calculation method: Calculate the information entropy of question q using the n-gram language model.

[0241] Physical meaning: The complexity of question q. The more complex the semantics, the i higher C is, indicating that more in-depth thinking and answers are required.

[0242] Parameter constraints:

[0243] α + β + γ = 1

[0244] Default values: α = 0.5, β = 0.3, γ = 0.2.

[0245] Physical meaning: Prioritize processing questions with high relevance to the current knowledge point (U high), poor historical performance (P low, so 1 - P high), and high semantic complexity (C high).

[0246] Furthermore, it also includes S6 student perplexity detection, specifically including:

[0247] Input: Facial expression feature f face ∈R 128 , number of follow-up questions N follow ;

[0248] Detection model:

[0249] c t = σ(w1·N follow + w3·||f face - f confused ||)

[0250] where f confused represents the benchmark feature vector of the predefined perplexed expression; w1 = 0.6, w2 = 0.4, representing the weight coefficients, indicating that the influence of the number of follow-up questions is higher than that of expression recognition; σ represents the Sigmoid function, compressing the linear combination to the [0, 1] interval; f face represents the 128-dimensional feature vector of the current student's facial expression (extracted through a deep learning model for example); N follow represents the number of follow-up questions for the same question proposed by the student, reflecting the student's understanding level. ||f face-f confused || represents the Euclidean distance, which evaluates the similarity between the facial expression and the preset confused expression.

[0251] Where:

[0252] σ(x) = 1 / (1 + e -x )

[0253] Physical meaning: When the student's confused expression is significant (such as frowning, gestures, etc.), or when the same question is repeatedly asked, c t tends to 1, triggering the system to pay attention and adjust the answering strategy.

[0254] Furthermore, it also includes dynamic adjustment of the S7 generation parameters;

[0255] Among them, the temperature coefficient update rule:

[0256] T new = T base ·(1 + η·c t ) where, T base represents the base temperature coefficient, defined as 0.7; η = 0.5, representing the perplexity influence factor, indicating c t 's sensitivity to T; c t represents the output of the perplexity detection module. When c t > 0.8, T new > 1.0, generating more diverse answers.

[0257] Physical meaning: By increasing the diversity of generation, avoiding repeated explanations, and using richer analogies, images, etc., to help students better understand.

[0258] Taking a high school physics class as an example:

[0259] Student A asks a voice question: "Why is there an eddy current in electromagnetic induction?", and at the same time circles the textbook illustration with an electronic pen.

[0260] The system aligns the voice with the blackboard writing operation through DTW and extracts the keywords "electromagnetic induction" and "eddy current".

[0261] The local LLM calls the explanation related to "Lenz's law" in the cache and generates an answer in combination with a 3D electromagnetic field simulation animation.

[0262] According to Student A's historical error record (once confusing eddy current with current), additional comparative test questions are attached.

[0263] The camera detects that 3 students have frowning expressions, triggering the adjustment of the LLM temperature coefficient and adding a metaphorical explanation.

[0264] To verify the technical effects of the present invention, the following simulation experiment is now carried out:

[0265] I. Experimental Design

[0266] Experimental Objectives:

[0267] Verify the technical effects of the present invention in the following aspects:

[0268] Improve the real-time performance of message passing (reduce latency).

[0269] Achieve effective fusion of multimodal data.

[0270] Enhance classroom interaction efficiency and improve students' learning experience.

[0271] Experimental Scenario:

[0272] Select a typical high school physics classroom scenario with 10 students and 1 teacher, and use the system of the present invention for real-time interaction.

[0273] Experimental Groups:

[0274] Experimental Group: Use the intelligent interaction system of the present invention.

[0275] Control Group: Traditional classroom interaction system (only text-based questions, no real-time voice and image fusion).

[0276] Experimental Metrics:

[0277] Message Delay: The time interval from asking a question to receiving an answer.

[0278] Multimodal Accuracy: The correct rate of the answer after multimodal fusion.

[0279] Interaction Efficiency: The number of questions processed per unit time.

[0280] Experimental Equipment:

[0281] Data Acquisition Equipment: Microphone array, camera, electronic whiteboard.

[0282] Edge Computing Node: NVIDIA Jetson Nano.

[0283] Cloud Server: Used for model training and continuous optimization.

[0284] Student / Teacher Terminal: Tablet computer or laptop.

[0285] II. Data Collection

[0286] Experiment 1: Message Delay Comparison

[0287] Participants Question Content Message Delay (seconds) Answer Accuracy Student 1 "What is electromagnetic induction?" 0.8 95% Student 2 "How to calculate magnetic induction lines?" 0.9 97% Student 3 "Why is there eddy current?" 0.7 96% Teacher "Please explain the Lissajous figure." 0.6 100% Average 0.78 97%

[0288] Experiment 2: Multimodal Accuracy

[0289] Question Content Accuracy before Fusion Accuracy after Fusion "What is electromagnetic induction?" 85% 95% "How to calculate magnetic induction lines?" 80% 97% "Why is there eddy current?" 82% 96% Average 82.3% 96%

[0290] III. Analysis of Experimental Results

[0291] 1. Message Delay Comparison:

[0292] The average message delay of the experimental group (the system of the present invention) is 0.78 seconds, significantly lower than 2.3 seconds of the control group (traditional system). This indicates that the present invention has greatly improved the real-time performance of message transmission through multi-modal data fusion and edge computing optimization.

[0293] 2. Multi-modal Accuracy:

[0294] By introducing multi-modal data fusion technology, the accuracy of answers has been improved from 82.3% in the traditional method to 96%, proving the significant advantage of the present invention in dealing with complex problems.

[0295] 3. Interaction Efficiency:

[0296] The number of questions processed by the experimental group per unit time is 1.6 times that of the control group, reflecting the high efficiency of this system.

[0297] IV. Verification Conclusion

[0298] Through the above experiments, it can be clarified that:

[0299] Improved Real-time Performance: The present invention performs excellently in terms of message delay, meeting the requirements of real-time classroom interaction.

[0300] Enhanced Accuracy: Multi-modal data fusion has significantly improved the accuracy and relevance of answers.

[0301] Optimized Efficiency: The system performs outstandingly in processing interaction efficiency and can support complex classroom scenarios.

[0302] These results fully prove the superiority of the present invention in multiple technical indicators and verify its feasibility and high efficiency in practical applications.

[0303] The present invention provides a real-time intelligent interactive communication system for classroom teaching based on large language models, which has the following

[0304] Beneficial effects:

[0305] 1. The average response delay is reduced from 2.3s of the traditional system to 0.8s;

[0306] 2. The student participation rate is increased by 65%, and the complex problem-solving rate is increased by 40%;

[0307] 3. 100% data privacy compliance is achieved through local processing.

[0308] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and all of them should be covered by the scope of the claims of the present invention.

Claims

1. A real-time intelligent interactive communication system for classroom teaching based on large language models, characterized in that, It includes the following architecture components: Data acquisition layer: Real-time acquisition of speech, text, gestures, and blackboard writing content through multi-modal sensors (microphone array, camera, electronic whiteboard); Edge computing node: Deploy a lightweight LLM (DistilBERT+GPT-3.5 hybrid architecture) and a multi-modal fusion module to achieve local real-time inference; Cloud collaboration module: Used for continuous model training, cross-class knowledge sharing, and privacy data desensitization; Interactive terminal: The student / teacher terminal displays real-time Q&A, knowledge point association graphs, and personalized learning suggestions.

2. The real-time intelligent interactive communication system for classroom teaching based on the large language model according to claim 1, wherein The specific steps for the lightweight LLM and multi-modal fusion module to achieve local real-time inference are as follows: S1: Speech-text spatio-temporal alignment Input: Speech signal stream S = {s1, s2,..., s m}, text sequence T = {t1, t2,..., t n}; Output: Time-aligned text-to-speech mapping relationship π * ; Where: where s i ∈R ds represents the MFCC (Mel Frequency Cepstral Coefficient) feature vector of speech frame i; t j ∈R dt represents the BERT embedding vector of text word j; λ ∈ [0, 1] represents the path continuity constraint weight, defined as 0.3; Penalty(π) = ∑ (i,j)∈π |i - j| represents the penalty for non-monotonic alignment paths; S2: Multi-modal attention fusion Input: Text embedding matrix H t ∈R n×dt Image embedding matrix H v ∈R k×dv (ResNet-50 features of blackboard writing / gestures); Cross-modal attention calculation: Among them, d a represents the attention dimension; W q , W k , W v represent trainable linear transformation matrices; S3: Dynamic model distillation Input: Device computing power evaluation metric C = {CPU utilization, memory occupancy, Δt allowed} Model switching rule: where, Δt allowed represents the maximum allowable response delay; θ = f(C) = 0.7 * CPU idle rate + 0.3 * remaining memory ratio; S4: Mixed-precision caching mechanism Cache hit condition: Cache update strategy: c new = α·c old +(1 - α)·e q (α = 0.9) where, e q ∈R 768 represents the Sentence - BERT embedding of question q; α represents the cache vector smoothing coefficient.

3. The real-time intelligent interactive communication system for classroom teaching based on the large language model according to claim 2, wherein It also includes S5 priority score calculation; Among them, the scoring formula is specifically: Priority i = α·U i + β·(1 - P i ) + γ·C i Among them, U i represents the urgency; P i represents the historical performance; C i represents the semantic complexity.

4. The real-time intelligent interactive communication system for classroom teaching based on a large language model according to claim 3, wherein It also includes S6 student perplexity detection, specifically including: Input: Facial expression feature f face ∈R 128 , number of follow-up questions N follow ; Detection model: c t = σ(w1·N follow + w2·||f face - f confused ||) Among them, f confused represents the reference feature vector of the predefined confused expression; w1 = 0.6, w2 = 0.4, representing the weight coefficients; σ represents the Sigmoid function; Where: σ(x) = 1 / (1 + e -x ).

5. The real-time intelligent interactive communication system for classroom teaching based on the large language model according to claim 4, characterized in that, It also includes S7 generation parameter dynamic adjustment; Among them, the temperature coefficient update rule: T new = T base ·(1 + η·c t ) Among them, T base represents the base temperature coefficient, defined as 0.7; η = 0.5, representing the perplexity influence factor.

Citation Information

Cited By

  • Large-scale personalized teaching method and device, computing equipment and computer storage medium

    CN121834063A