A method and device for analyzing interactive activities in Chinese language classrooms based on line-following interactive modeling

CN122573646APending Publication Date: 2026-08-14WENHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

现有高中语文课堂教学评价方法存在效率低下、主观性强、评价维度缺失且结果片面等问题

Benefits of technology

本发明通过同步获取课堂视频流的视觉模态特征与音频流的听觉模态特征,基于预设时间窗口内的多模态特征数据构建循迹式时空师生动态交互图,并将图特征转化为大语言模型(LLM)可解析的结构化文本序列。通过执行层次化多模态融合策略,将所述结构化文本序列输入LLM以输出多模态交互特征嵌入,再将此嵌入与原始的视觉、听觉模态特征进行深度协同与融合,生成统一的多模态交互表征,使得LLM所学习到的高级语义信息(如交互意图、逻辑链条)能够作为引导信号,反向优化和调制原始的视觉与听觉特征,实现了语义与感知信息的深度耦合与互补。然后,将生成的统一多模态交互表征与语文教学场景专属提示模版拼接,再次输入LLM并通过多任务指令设计解码输出目标字段的结构化语义元组,充分利用LLM的语义生成与逻辑推理能力,将融合后的多模态信息精准映射到语文教学评价所关切的特定维度,从而自动生成富含语义的结构化课堂交互记录。最后,基于连续输出的结构化语义元组统计生成课堂交互密度与文本关联度等量化评估指标,使得对课堂教学过程的定性观察转化为可量化、可比较的客观数据,能够清晰呈现互动频率、互动与教学内容的关联程度等核心质量维度,从而为实现自动化、精细化、数据驱动的课堂教学评价提供了直接、可靠的依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122573646A_ABST
    Figure CN122573646A_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for analyzing classroom interaction in Chinese language teaching based on line-following interactive modeling. The method includes: simultaneously acquiring visual modal features of a classroom video stream and auditory modal features of an audio stream; constructing a line-following spatiotemporal dynamic interaction graph of teachers and students based on multimodal feature data within a preset time window, and converting the graph features into a structured text sequence that can be parsed by a large language model; embedding the structured text sequence into an LLM (Language Modeling) to output multimodal interaction features, and then deeply collaborating and fusing this embedding with the original visual and auditory modal features to generate a unified multimodal interaction representation; concatenating the generated unified multimodal interaction representation with a prompt template specific to the Chinese language teaching scenario, inputting it again into an LLM, and decoding and outputting structured semantic tuples of the target field through multi-task instruction design; and statistically generating quantitative evaluation indicators such as classroom interaction density and text relevance based on the continuously output structured semantic tuples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of educational informatization technology, and in particular to a method and device for analyzing interactive classroom interactions in Chinese language based on line-following interactive modeling. Background Technology

[0002] Classroom interaction is the core of high school Chinese language teaching. Teachers' nonverbal behaviors, such as body language and vocal cord movements, are deeply coupled with verbal behaviors like text interpretation and questioning. Learners' interactive behaviors, such as classroom responses and group discussions, are also closely related to teachers' guidance. Accurately capturing and analyzing these teacher-student interactions is a crucial prerequisite for evaluating the quality of Chinese language classroom teaching and guiding young teachers to improve their teaching skills. Existing high school Chinese language classroom teaching evaluation methods suffer from problems such as low efficiency, strong subjectivity, lack of evaluation dimensions, and one-sided results. Specifically, traditional evaluation methods rely heavily on manual coding and subjective judgment, resulting in extremely low evaluation efficiency. Furthermore, there are significant subjective differences in judging teachers' nonverbal behaviors and ambiguous teacher-student interactions, making it difficult to comprehensively cover complex teaching scenarios and provide personalized teaching suggestions.

[0003] In addition, existing multimodal classroom analysis technologies mostly extract single-modal features independently and then simply splice or vote on them, failing to delve into the deep connections between cross-modal features such as visual, auditory, and vocal cord movement. This results in the inability to accurately identify the core direction of teachers' teaching guidance, judge the emotional state of teachers and students, and the true intention of teaching interactions.

[0004] Meanwhile, existing technologies lack generative LLM capabilities tailored to the characteristics of the Chinese language and literature discipline. This makes it difficult to transform multimodal classroom features into structured semantic information relevant to the Chinese language and literature teaching scenario. Consequently, it fails to accurately capture the linguistic logic and intellectual clashes in text interpretation and discussions of literary viewpoints, and also struggles to provide targeted teaching optimization suggestions. These issues collectively result in insufficient ability to accurately capture interactive intentions and states in the two core scenarios of high school Chinese language interactive classrooms (text interpretation and group problem discussion), severely restricting the accuracy and interpretability of interactive understanding, and failing to provide efficient data support for smart classroom supervision and automated teaching evaluation. Summary of the Invention

[0005] This invention provides a method and device for analyzing interactive activities in Chinese language classrooms based on line-following interactive modeling, in order to overcome the deficiencies in existing technologies and provide efficient data support for smart classroom supervision and automated teaching evaluation.

[0006] In a first aspect, the present invention provides a method for analyzing language classroom interactions based on line-following interaction modeling, comprising: The visual modal features of the classroom video stream and the auditory modal features of the classroom audio stream are obtained to obtain multimodal feature data. Based on multimodal feature data within a preset time window, a tracking spatiotemporal dynamic interaction graph of teachers and students is constructed, and the graph features of the tracking spatiotemporal dynamic interaction graph of teachers and students are transformed into a structured text sequence that can be parsed by a large language model. A hierarchical multimodal fusion strategy is implemented, inputting the structured text sequence into a large language model and outputting multimodal interaction feature embeddings; The multimodal interaction features are embedded as cross-modal guidance signals, and the visual modal features are dynamically modulated through a cross-attention mechanism and weighted and fused with the auditory modal features to generate a unified multimodal interaction representation. After combining the unified multimodal interaction representation with the language teaching scenario-specific prompt template, the data is input into the large language model again. Through multi-task instruction design, the structured semantic tuple of the target field is decoded and output. Based on the continuous output of the structured semantic tuples, a quantitative evaluation index for classroom interaction density and text association is generated to realize classroom teaching evaluation.

[0007] Furthermore, the visual modal features of the acquired classroom video stream include: The visual modal features include the two-dimensional joint coordinates and limb direction unit vectors of individual teachers and students; The video stream is input into the visual modality feature extraction subnetwork for frame-by-frame processing to obtain the two-dimensional key point coordinates of each teacher and student at each time step; Based on the teacher and student's nose tip joint coordinates and neck joint coordinates in the two-dimensional joint coordinates, calculate the limb direction unit vector representing the head orientation, and based on the wrist joint coordinates and shoulder joint coordinates in the two-dimensional joint coordinates, calculate the limb direction unit vector representing the arm orientation.

[0008] Furthermore, the auditory modal features of the classroom audio stream include: The auditory modal features include vocal cord motion feature vectors and speaker state coding features; The classroom video stream is input into the vocal cord motion feature processing network to perform laryngeal region localization and vocal cord motion feature extraction, resulting in a vocal cord motion feature vector. The classroom audio stream is input into the auditory modality feature extraction subnetwork for feature extraction, speech activity detection, speaker separation and recognition, to obtain the speaker state coding features at each time step.

[0009] Furthermore, the step of converting the graph features of the tracking-based spatiotemporal dynamic interaction graph between teachers and students into a structured text sequence that can be parsed by a large language model includes: Following the fixed format of "time step, subject, feature, scene prompt", the features of each individual node in the tracking spatiotemporal dynamic interaction diagram of teachers and students at each time step are converted into text entries. All text entries are arranged in chronological order, and features of the Chinese language classroom scene that identify the current teaching segment are embedded in the sequence to obtain a structured text sequence.

[0010] Furthermore, the step of inputting the structured text sequence into a large language model and outputting multimodal interaction feature embeddings includes: Based on the limb orientation and vocal cord features in the structured text sequence, the spatial orientation relationship and emotional state of teachers and students are identified and spatial association labels and emotional labels are output. The coherent behavioral logic of teachers and students within the time step is explored and a trace-following temporal interaction chain is constructed. It also includes multimodal interaction feature embedding that extracts and integrates multidimensional information such as semantic relevance, geometric association, emotional state and temporal logic.

[0011] Furthermore, the implementation of the hierarchical multimodal fusion strategy includes: When constructing the tracking-based spatiotemporal dynamic interaction graph between teachers and students, the visual modal features and the auditory modal features are spliced ​​and fused to obtain the initial node features of the interaction graph. The multimodal interaction features are embedded as cross-modal guidance signals, and the visual modal features are dynamically modulated through a cross-attention mechanism to obtain modulated visual features. The modulated visual features are globally pooled to obtain a visual global representation, which is then weighted and fused with the auditory global representation obtained by globally pooling the auditory modal features to generate a multimodal interaction feature embedding.

[0012] Furthermore, the step of concatenating the unified multimodal interaction representation with the language teaching scenario-specific prompt template, and then inputting it again into the large language model, decodes and outputs the structured semantic tuple of the target field through multi-task instruction design, including: Customized prompt templates are designed for different teaching stages in high school Chinese classes. The prompt templates clearly define the constraints of each field of the structured semantic tuple. The target fields include {subject, action, object, responder, text association, and thought state}. The large language model decodes and outputs the structured semantic tuples that conform to the field constraints according to the instructions of the customized prompt template. The values ​​of the action fields are selected from a preset set of actions related to teaching interaction. The object fields are associated with paragraphs, literary elements or discussion topics in high school Chinese textbooks. The text association fields correspond to specific sentences in the textbook.

[0013] Furthermore, the statistical generation of classroom interaction density and text association metrics based on the continuously output structured semantic tuples includes: Based on the structured semantic tuple sequence output within a continuous time window, quantitative evaluation indicators including classroom interaction density, text relevance, interaction type distribution, and the proportion of thinking states are statistically generated.

[0014] Secondly, the present invention also provides a Chinese language classroom interaction analysis device based on tracking-based interactive modeling, comprising: a multimodal feature data acquisition module, used to acquire the visual modal features of the classroom video stream and the auditory modal features of the classroom audio stream, so as to obtain multimodal feature data; The first data processing module is used to construct a tracking spatiotemporal dynamic interaction graph of teachers and students based on multimodal feature data within a preset time window, and to convert the graph features of the tracking spatiotemporal dynamic interaction graph of teachers and students into a structured text sequence that can be parsed by a large language model. The first data fusion module is used to execute a hierarchical multimodal fusion strategy, inputting the structured text sequence into a large language model and outputting multimodal interaction feature embeddings. The second data fusion module is used to embed the multimodal interaction features as cross-modal guidance signals, dynamically modulate the visual modal features through a cross-attention mechanism, and perform weighted fusion with the auditory modal features to generate a unified multimodal interaction representation. The second data processing module is used to combine the unified multimodal interaction representation with the language teaching scenario-specific prompt template and input it again into the large language model. Through multi-task instruction design, it decodes and outputs the structured semantic tuple of the target field. The analysis module is used to statistically generate quantitative evaluation indicators for classroom interaction density and text association based on the continuously output structured semantic tuples, thereby realizing classroom teaching evaluation.

[0015] Thirdly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described methods for analyzing interactive language classroom interactions based on line-following interactive modeling.

[0016] The method and apparatus for analyzing language classroom interactions based on line-following interactive modeling provided by this invention can produce the following beneficial technical effects compared with the prior art: This invention acquires visual modal features from classroom video streams and auditory modal features from audio streams simultaneously. Based on multimodal feature data within a preset time window, it constructs a spatiotemporal dynamic interaction graph of teachers and students, and transforms the graph features into a structured text sequence that can be parsed by a Large Language Model (LLM). By executing a hierarchical multimodal fusion strategy, the structured text sequence is input into the LLM to output multimodal interaction feature embeddings. These embeddings are then deeply synergistically fused with the original visual and auditory modal features to generate a unified multimodal interaction representation. This allows the high-level semantic information learned by the LLM (such as interaction intentions and logical chains) to serve as guiding signals, inversely optimizing and modulating the original visual and auditory features, achieving deep coupling and complementarity between semantic and perceptual information. Then, the generated unified multimodal interaction representation is concatenated with a prompt template specific to the Chinese language teaching scenario, inputted again into the LLM, and decoded through multi-task instructions to output structured semantic tuples of the target fields. This fully utilizes the semantic generation and logical reasoning capabilities of the LLM to accurately map the fused multimodal information to specific dimensions relevant to Chinese language teaching evaluation, thereby automatically generating semantically rich structured classroom interaction records. Finally, based on the continuous output of structured semantic tuples, quantitative evaluation indicators such as classroom interaction density and text relevance are generated. This transforms the qualitative observation of the classroom teaching process into quantifiable and comparable objective data, which can clearly present core quality dimensions such as interaction frequency and the degree of relevance between interaction and teaching content. This provides a direct and reliable basis for achieving automated, refined, and data-driven classroom teaching evaluation. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is one of the flowcharts of an optional method for analyzing interactive language classroom interactions based on line-following interactive modeling provided by the present invention; Figure 2 This is a schematic diagram illustrating an application scenario of an optional method for analyzing interactive language classroom interactions based on line-following interactive modeling provided by the present invention. Figure 3 This is the second flowchart of an optional method for analyzing interactive language classroom interactions based on line-following interactive modeling provided by the present invention; Figure 4 This is a schematic diagram of an optional overall model architecture for analyzing interactive Chinese language classroom interactions based on line-following interactive modeling, provided by the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0020] It should be noted that, in the description of the embodiments of the present invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0021] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more.

[0022] Figure 1 This is one of the flowcharts illustrating an optional method for analyzing interactive language classroom interactions based on line-following interactive modeling provided by this invention. Figure 1 As shown, including but not limited to the following steps: S102. Obtain the visual modal features of the classroom video stream and the auditory modal features of the classroom audio stream to obtain multimodal feature data; S104. Based on the multimodal feature data within a preset time window, construct a tracking spatiotemporal dynamic interaction graph of teachers and students, and convert the graph features of the tracking spatiotemporal dynamic interaction graph of teachers and students into a structured text sequence that can be parsed by a large language model. S106. Execute a hierarchical multimodal fusion strategy, input the structured text sequence into a large language model, and output multimodal interaction feature embeddings; S108. The multimodal interaction features are embedded as cross-modal guidance signals, the visual modal features are dynamically modulated through a cross-attention mechanism, and weighted and fused with the auditory modal features to generate a unified multimodal interaction representation. S110. After splicing the unified multimodal interaction representation with the Chinese language teaching scenario-specific prompt template, input it again into the large language model, and decode and output the structured semantic tuple of the target field through multi-task instruction design. S112. Based on the continuous output of the structured semantic tuples, statistically generate classroom interaction density and text association quantification evaluation indicators to realize classroom teaching evaluation.

[0023] Figure 2 This is a schematic diagram illustrating an application scenario of an optional method for analyzing interaction in a Chinese language classroom based on line-following interaction modeling, as provided by the present invention. Figure 2 As shown, the implementation scenario of this embodiment is a room equipped with two Seewo cameras. (Collection of teacher information) A standard smart classroom (collecting learner data) and deploying a line-following classroom recorder, along with an overall model built using a line-following interactive modeling method for analyzing Chinese language classroom interactions, is deployed on an edge computing server to achieve real-time interactive analysis of the "Lotus Pond in the Moonlight" interactive classroom.

[0024] Before deploying a large language model, offline training is required to optimize the model parameters: 1) Collect a large amount of recorded videos and synchronized audio data from interactive high school Chinese language classes, focusing on teacher-student interactive interpretations and group reading discussions of the prose "Lotus Pond in the Moonlight." Feature trajectory collection was completed using a tracking-based multimodal perception module. High school Chinese language teaching experts and educational technology experts were invited to manually annotate key interactive segments in the videos according to the {subject, action, object, responder, text association, thought state} format defined in this invention. For example, a scene from this class was annotated as (Teacher, explanation, description of the moonlight in paragraph 3 of "Lotus Pond in the Moonlight," whole class of learners, the phrase "moonlight like flowing water," emphasized).

[0025] 2) The training video data is converted into a structured text sequence in the format of "time step + subject + feature + scene cue", and manually annotated structured semantic tuples are used as training targets. This is used to construct an instruction fine-tuning dataset for supervised fine-tuning of the LLM model.

[0026] 3) The overall model described in this invention is built using the PyTorch deep learning framework. An end-to-end training strategy is adopted, with all modules (except the pre-trained pose measurement sub-network) participating in training together. Simultaneously, the deployment and debugging of the tracking-based classroom interaction recorder are completed, achieving full-process retention and backtracking of classroom interaction data. The optimizer used is AdamW, with an initial learning rate set to 1e-4, and decayed using a cosine annealing strategy. InfoNCE loss is used to optimize feature representation.

[0027] in For feature similarity function, This is a temperature parameter with a value of 0.07.

[0028] In this embodiment, Figure 3 This is the second flowchart of an optional method for analyzing interactive language classroom interactions based on line-following interactive modeling provided by this invention. Figure 3 As shown, this is achieved by deploying dual cameras (such as Seewo cameras) in the classroom. and Simultaneously acquire video and audio streams covering both the teacher and learner areas. For the video stream, input it into the visual modality feature extraction sub-network VisNet for processing, which may have a built-in vocal cord motion feature processing network CocalNet.

[0029] VisNet is used to perform human detection and tracking on video frame sequences, estimating the two-dimensional joint coordinate sequence of each teacher and student in the classroom at each time step. For example, in the "Lotus Pond in the Moonlight" class, the coordinates of the teacher's right wrist pointing to the blackboard and the coordinates of the student's arm joint when raising their hand to answer can be accurately detected.

[0030] Meanwhile, the CocalNet network locates the region of interest in the larynx based on face detection and cervical joints, performs inter-frame difference and optical flow calculation on the region, and quantifies it to obtain three types of features: vocal cord vibration amplitude, frequency stability, and vibration regularity. These features together constitute the vocal cord motion feature vector, which can be used to infer the emotional state when speaking. For example, when a teacher is calmly explaining, the vibration amplitude is stable at 0.7-0.9.

[0031] For synchronously acquired audio streams, the input is processed by the auditory modality feature extraction subnetwork AusNet. Short-time Fourier transform and Mel-frequency cepstral coefficients are used to extract voiceprint features. Speech activity detection and deep learning-based speaker separation and recognition are then performed, ultimately generating speaker state coding features at each time step that identify "who is speaking." This completes the acquisition of visual modality features (including limb orientation and vocal cord movement features) and auditory modality features, resulting in multimodal feature data.

[0032] Next, based on the multimodal feature data within a preset time window (e.g., 5 seconds), a spatiotemporal dynamic interaction graph of teachers and students is constructed. Specifically, the classroom scene within this time window is modeled as a graph structure G=(V, E), where nodes V represent individual teachers and students, and edges E represent the spatiotemporal interaction relationships between individuals. The initial features of each node are formed by splicing and fusing the individual's joint coordinate vector, limb direction vector, vocal cord motion feature vector, and speaker state code through a multilayer perceptron, which achieves early feature fusion.

[0033] Next, the graph features of this interaction graph need to be transformed into input that can be parsed by a Large Language Model (LLM). The overall model can be structured according to a fixed format, where the features of each node at each time step are converted into a text entry, and all entries are arranged chronologically. Simultaneously, features identifying the current teaching segment (e.g., the "text analysis segment of 'Lotus Pond in the Moonlight'") are embedded into the sequence, thus generating a structured text sequence. This sequence transforms complex multimodal graph data into a semantic description that LLM can directly understand.

[0034] Then, a hierarchical multimodal fusion strategy is implemented. On the one hand, the structured text sequence generated above is input into the LLM, which is guided to perform deep learning of interactive features through a carefully designed customized prompt template engineering. For example, the prompt template can ask the LLM to "identify the spatial orientation relationship and emotional state of teachers and students in the classroom based on the body direction and vocal cord features in the sequence." Based on this, the LLM can output "spatial association labels" (such as "Teacher → blackboard, paragraph 3 of 'Lotus Pond in the Moonlight' → learner A") and "emotional labels" (such as "Teacher calmly explains, learner A listens excitedly"). At the same time, the prompt template can also guide the LLM to explore the coherent behavioral logic of teachers and students within the time step, and construct a trace-based temporal interaction chain such as "explanation-question-response-supplement" or "speaking-responding-refuting-consensus". In this deep learning process, a multimodal interactive feature embedding that integrates multi-dimensional information such as semantic relevance, geometric association, emotional state and temporal logic is formed within the LLM.

[0035] On the other hand, while LLM performs feature processing, the overall model implements hierarchical fusion: the initial feature splicing of nodes is completed during the early construction of the interaction graph (early fusion); then, the multimodal interaction features generated by LLM are embedded as "cross-modal guiding signals," and the original visual feature set is dynamically modulated through a cross-attention mechanism, enabling the visual features to be adaptively enhanced or suppressed according to the semantic context (such as "who is calmly explaining"); finally, the modulated visual features are globally pooled to obtain a visual global representation, and the auditory features are globally pooled to obtain an auditory global representation, and the two are weighted and fused to generate the final unified multimodal interaction representation. The fusion weights here are not fixed and can be dynamically generated by LLM according to the current classroom interaction scenario (whether it is "text parsing" or "group discussion"), thereby achieving output layer optimization.

[0036] Finally, the obtained unified multimodal interaction representation is combined with a dedicated prompt template designed for high school Chinese teaching scenarios, input into LLM again, and decoded and output the structured semantic tuple of the target field through multi-task instruction design.

[0037] Specifically, customized prompt templates can be designed for different teaching stages such as "teacher-student interactive text analysis" and "group discussion and exchange." These prompt templates clearly define the constraints of each target field in the structured semantic tuples. These target fields can include {subject, action, object, responder, text association, and thought state}. The LLM decodes and outputs tuples that conform to the constraints. The action field selects values ​​from a pre-defined set of actions related to teaching interaction (such as lecturing, asking questions, pointing, answering, speaking, responding, refuting, and reaching consensus); the object field associates with paragraphs, literary elements, or discussion topics in high school Chinese textbooks (such as "the moonlight description in paragraph 3 of 'Moonlight over the Lotus Pond'"); and the text association field specifies the corresponding sentences in the text (such as "the moonlight, like flowing water, quietly spills onto these leaves and flowers").

[0038] Based on these structured semantic tuple sequences output within a continuous time window, this invention can statistically generate quantitative evaluation indicators including classroom interaction density (e.g., 3 effective interactions within a 5-second window), text relevance (e.g., 90% of interaction events are directly related to the core sentences of the text), distribution of interaction types, and proportion of thinking states, thereby achieving automated and refined classroom teaching evaluation and providing data support for generating personalized teaching ability evaluation reports.

[0039] In an optional embodiment, the visual modal features of the acquired classroom video stream in the Chinese language classroom interaction analysis method based on line-following interaction modeling provided by the present invention include: The visual modal features include the two-dimensional joint coordinates and limb direction unit vectors of individual teachers and students; The video stream is input into the visual modality feature extraction subnetwork for frame-by-frame processing to obtain the two-dimensional key point coordinates of each teacher and student at each time step; Based on the teacher and student's nose tip joint coordinates and neck joint coordinates in the two-dimensional joint coordinates, calculate the limb direction unit vector representing the head orientation, and based on the wrist joint coordinates and shoulder joint coordinates in the two-dimensional joint coordinates, calculate the limb direction unit vector representing the arm orientation.

[0040] In this embodiment, Figure 4 This is a schematic diagram of an optional overall model architecture for analyzing interactive elements in a Chinese language classroom based on line-following interactive modeling, provided by the present invention. Figure 4 This illustrates the complete data flow and module connections from multimodal feature extraction (VisNet, CocalNet, AusNet), interaction graph construction, structured text sequence generation, LLM feature embedding extraction, hierarchical multimodal fusion, to structured semantic tuple decoding. Visual modal features are extracted using the VisNet network. This includes: Step 1.1.1: Input the video frame sequence into the pre-trained VisNet sub-network, and detect individual teacher and learner interactions in the classroom frame by frame. , Output heatmaps of key points, including the teacher's hand gestures when pointing to the passage from "Lotus Pond in Moonlight" on the blackboard, and the learners' hand-raising postures when responding. And coordinate offsets. Among them, the visual modality feature extraction subnetwork can adopt the VisNet subnetwork, for example, adopting a human pose estimation architecture based on high resolution network (HRNet), with an input size of 384 × 288, and outputting J=17 two-dimensional joint heatmaps and corresponding coordinate offsets.

[0041] No. k The two-dimensional coordinates of each key point are decoded from the heatmap using the following formula:

[0042] in, argmax The operation retrieves the location of the maximum heatmap response. The subpixel-level offsets for network regression together constitute high-precision coordinates at heatmap resolution. .

[0043] To improve the robustness of pose estimation, 2D joints are added. loss:

[0044] in These are the actual coordinates of the 2D joints. For predicted coordinates.

[0045] Step 1.1.2: Calculate limb geometric vectors: Using the keypoint coordinate scale recovery formula, ... Mapping to the original image coordinate system yields the final joint coordinates. ,based on Calculate the unit vectors of body orientation for two types of characters with semantic meaning related to interactive elements in a Chinese language classroom, thereby ensuring consistency of posture at different distances:

[0046] in, For the final joint coordinates, For the scale recovery ratio, , This represents the height and width of the heatmap.

[0047] In an optional embodiment, the method for analyzing language classroom interaction based on line-following interaction modeling provided by the present invention, which involves obtaining the auditory modal features of audio streams in the classroom, includes: The auditory modal features include vocal cord motion feature vectors and speaker state coding features; The video stream is input into the CocalNet vocal cord motion feature processing network to perform laryngeal region localization and vocal cord motion feature extraction, resulting in a vocal cord motion feature vector. The audio stream is input into the auditory modality feature extraction subnetwork AusNet for feature extraction, speech activity detection, speaker separation and recognition, to obtain the speaker state coding features at each time step.

[0048] In this embodiment, vocal cord motion features are extracted using the CocaNet network, including: Step 1.2.1: Using the coordinates of the neck joints, the region of interest (PoG) of the larynx is selected. After locating the PoG region of the larynx, it is converted to grayscale, and a motion mask is generated using the inter-frame difference method.

[0049]

[0050] in, For the first t Frame PoG grayscale image, The difference threshold is used (an empirical value of 20-30).

[0051] Step 1.2.2: Quantify the vocal cord motion characteristics, including frequency stability, vibration regularity, and construct feature vectors. The vibration amplitude remained stable at 0.7-0.9 when the teacher calmly explained during the "text analysis session," while the vibration amplitude fluctuated between 0.6-0.8 when the learners debated during the "group discussion session."

[0052] In this embodiment, acoustic features are extracted and speaker processing is performed using the AusNet network, including: Step 1.3.1: Analyze the audio signal First, a short-time Fourier transform is performed to obtain the time spectrum, and then the energy of the Mel filter bank is calculated.

[0053] Where M is the Mel filter bank matrix, for The MFCC coefficients are obtained by performing a discrete cosine transform.

[0054] To simplify computation, speech activity detection can employ a short-time energy-based approach. and zero crossing rate The dual-threshold method is used for preliminary endpoint detection.

[0055] Step 1.3.2: For scenarios where multiple people speak simultaneously, a deep learning-based speech separation method is employed. For example, the idea of ​​deep clustering can be borrowed, where the loss function encourages the embedding of time-frequency units belonging to the same speaker into the speech vector. near:

[0056] in, V It is the embedding vector matrix of all time-frequency points. Y This is the actual one-hot speaker label matrix. Calculate the speaker feature confidence scores:

[0057] in A This is a multi-head attention graph. Ha The number of attention heads. The final output is the active state encoding (one-hot vector) for each speaker in each time slice. .

[0058] In an optional embodiment, the method for analyzing language classroom interaction based on line-following interaction modeling provided by the present invention, which involves converting the graph features of the line-following spatiotemporal dynamic interaction graph between teachers and students into a structured text sequence that can be parsed by a large language model, includes: Following the fixed format of "time step, subject, feature, scene prompt", the features of each individual node in the tracking spatiotemporal dynamic interaction diagram of teachers and students at each time step are converted into text entries. All text entries are arranged in chronological order, and features of the Chinese language classroom scene that identify the current teaching segment are embedded in the sequence to obtain a structured text sequence.

[0059] In this embodiment, all individual data within the current processing time window are modeled as a dynamic graph. G .node V Dimensional unification was achieved through the MLP network, completing early-stage fusion. Spatial edge This relates to the teacher-student interaction in the "text analysis phase (e.g., paragraph 3 of 'Lotus Pond in Moonlight')" and the interaction among group members in the "group discussion phase (is the author's emotion "pure tranquility" or "a complex mix of joy and sorrow")." At this point, spatial attention masking filters out invalid connections.

[0060] in For attention mask, ⊙ represents element-wise product. Temporal edge. It refers to the coherent behavior of learners from listening to raising their hands to respond, and from independent thinking to group presentations.

[0061] Step 2.2: Following a fixed format of "time step + subject + feature + scene cues," transform the spatiotemporal graph features into a structured text sequence understandable by LLM, embedding scene information from the two main stages of the "Lotus Pond in Moonlight" classroom lesson. This data is also retained as core data for the tracking-based classroom interaction recorder, including: 1) Example of a sequence for teacher-student interactive analysis of paragraph 3 of the text "Lotus Pond in Moonlight": { Teacher, coordinates: (120, 350), head orientation: (120, 350), arm orientation: (280, 420), vocal cords (vibration amplitude 0.8, frequency stability 0.9, vibration regularity 0.85), pointing to the blackboard, reading paragraph 5 of "Lotus Pond in the Moonlight", explaining, "the emotional expression in the description of moonlight"; { Learner A, coordinates: (450, 380), head orientation: (450, 380), arm orientation: (460, 320), vocal cords (vibration amplitude 0.3, frequency stability 0.7, vibration regularity 0.6), facing the teacher, listening, “emotional expression in the description of moonlight”; 2) Group discussion and exploration: "Is the author's emotion 'pure tranquility' or 'a complex mix of joy and sorrow'?" Example of the sequence: { Learner A, coordinates: (320, 400), head orientation: (320, 400), arm orientation: (300, 430), vocal cords (vibration amplitude 0.7, frequency stability 0.6, vibration regularity 0.5), facing learner B, discussing, "the effect of lotus leaf fragrance"; { Learner B, coordinates: (350, 405), head orientation: (350, 405), arm orientation: (360, 380), vocal cords (vibration amplitude 0.6, frequency stability 0.7, vibration regularity 0.6), facing learner A, responds, "reflecting a peaceful state of mind".

[0062] In an optional embodiment, the method for analyzing language classroom interaction based on line-following interaction modeling provided by the present invention, which involves inputting the structured text sequence into a large language model and outputting multimodal interaction feature embeddings, includes: Based on the limb orientation and vocal cord features in the structured text sequence, the spatial orientation relationship and emotional state of teachers and students are identified and spatial association labels and emotional labels are output. The coherent behavioral logic of teachers and students within the time step is explored and a trace-following temporal interaction chain is constructed. It also includes multimodal interaction feature embedding that extracts and integrates multidimensional information such as semantic relevance, geometric association, emotional state and temporal logic.

[0063] In this embodiment, the generated structured text sequence is input into the LLM, and the model is guided to complete the deep learning of the interactive features of the two major stages of the "Lotus Pond in the Moonlight" class through a customized prompt template project.

[0064] Step 3.1: Identify the spatial and emotional dual-dimensional interaction states, and output spatial association tags and emotional tags. 1) For the "Teacher-Student Interaction in Text Analysis" segment, input the following prompt template into the LLM: {"Based on the structured text sequence of the high school Chinese text 'Lotus Pond in the Moonlight,' identify the spatial orientation relationship and emotional state between teachers and students. Analyze the spatial association between the teacher's body language and the learner's position and the text area; combine vocal cord movement characteristics to determine the behavioral state (explaining / listening / responding) and emotions (calm / tense / excited) of teachers and students, and output spatial association labels and emotion labels."} Model output: {"Spatial association label: Teacher → Blackboard, paragraph 3 of 'Lotus Pond in the Moonlight' → Learner A (pointing to the blackboard, horizontal angle of 15°); Emotion label: Teacher calmly explains, Learner A listens excitedly, Learner B nervously prepares to add a comment."} 2) For the "Group Discussion and Discussion" segment, input the following prompt template into the LLM: {"Based on the structured text sequence of the group discussion segment of the high school Chinese text 'Lotus Pond in the Moonlight,' identify the spatial orientation relationships and emotional states among group members: analyze the interactive correlation of group members' body pointing and head orientation; combine vocal cord movement characteristics to determine the group members' discussion state (speaking / responding / thinking) and emotions (calm / excited / arguing), and output spatial association labels and emotion labels."} Model output: {"Spatial association labels: Learner A → Learner B (pointing angle 5°), Learner B → Learner A (gaze angle 3°); Emotion labels: Learner A speaks enthusiastically, Learner B responds calmly, Learner C is in a contemplative state."} Step 3.2: Construct a trace-based temporal interaction chain 1) For the "Teacher-Student Interaction in Text Analysis" section, input the following prompt template into the LLM: {"Summarize the timeline for the text analysis of 'Lotus Pond in the Moonlight'"} The structured text sequence is analyzed to uncover the coherent behavioral logic between teachers and students, annotating the complete interaction chain of 'explanation-question-response-supplementation', and outputting a temporal interaction summary. Model output: {" The teacher explains the emotional expression (calm state) of the moonlight description in the third paragraph of "Moonlight over the Lotus Pond". The teacher turns to face the learners and asks, "What sensory descriptions are used in this passage?" (The tone is questioning, and the vocal cords vibrate at an amplitude of 0.9). Learner A looks up at the teacher, raises his hands in preparation to respond (vocal cord vibration amplitude 0.4, tense state). Learner A's speech included visual descriptions, such as "the moonlight was like flowing water" (vocal cord vibration amplitude was 0.6, and the emotion was calm). Learner B adds, "There's also the olfactory description 'the fragrance of lotus leaves' (vocal cord vibration amplitude 0.7, emotional excitement)." 2) For the "Group Discussion and Seminar" session, input the following prompt template into the LLM: {"Summarize the timeline for the 'Lotus Pond Moonlight' group discussion session."} The structured text sequence is analyzed to uncover the coherent behavioral logic of group members, annotating the complete interaction chain of 'speaking-responding-refuting-consensus', and outputting a temporal interaction summary. Model output: {" Learner A suggests that "the fragrance of lotus leaves is to enhance the lively atmosphere" (vocal cord vibration amplitude 0.7, excited state). Learner B countered, "It should reflect a tranquil state of mind" (vocal cord vibration amplitude 0.6, a firm state). Learner C adds, "Considering the context, a tranquil state of mind better fits the overall tone of the text" (vocal cord vibration amplitude 0.5, calm state). The three reached a consensus, confirming that 'the description of smell evokes a tranquil state of mind' (the amplitude of vocal cord vibration remained stable at 0.5-0.6). Step 3.3: Extract the vectors from the last hidden layer of the LLM and construct a multimodal interaction feature embedding. This vector integrates multi-dimensional information such as semantic relevance, geometric association, emotional state, and temporal logic from the two major stages of the "Lotus Pond in the Moonlight" classroom.

[0065] In an optional embodiment, the hierarchical multimodal fusion strategy implemented in the Chinese language classroom interaction analysis method based on line-following interaction modeling provided by the present invention includes: When constructing the tracking-based spatiotemporal dynamic interaction graph between teachers and students, the visual modal features and the auditory modal features are spliced ​​and fused to obtain the initial node features of the interaction graph. The multimodal interaction features are embedded as cross-modal guidance signals, and the visual modal features are dynamically modulated through a cross-attention mechanism to obtain modulated visual features. The modulated visual features are globally pooled to obtain a visual global representation, which is then weighted and fused with the auditory global representation obtained by globally pooling the auditory modal features to generate a multimodal interaction feature embedding.

[0066] In this embodiment, during the LLM feature processing, a hierarchical fusion strategy of "early fusion - intermediate layer modulation - output layer optimization" is implemented to achieve deep collaboration between LLM features and visual, auditory, and vocal cord motion features, completing the full-process fusion of feature tracking and adapting to the multimodal interaction requirements of the two main stages of the "Lotus Pond in the Moonlight" classroom. Audio information is deeply integrated into the visual processing pipeline in this step.

[0067] S4.1 Early Feature Fusion: The node initialization formula in step 2 has been implemented. This is implemented in [the framework / system], providing prior knowledge for cross-modal associations in the model. Feature dimension alignment operation:

[0068] PWConv is a pointwise convolution, which ensures that the dimensions of visual and auditory features are consistent.

[0069] S4.2 Intermediate Layer Cross-Attention Fusion: Embed the multimodal interaction features output from step 3. As a cross-modal guiding signal, the visual feature matrix of all nodes is: The auditory feature matrix for the aligned time period is Through a cross-attention mechanism, audio information is dynamically modulated to visual features, where the cross-attention calculation is as follows:

[0070] Semantic matching score graph optimization for cross-modal association:

[0071] in, For feature concatenation operations, These are the normalized visual features. Embedding audio semantic cues. This operation enables visual features to be adaptively enhanced or suppressed based on audio context, such as "who is speaking" and "whether the tone is interrogative or affirmative."

[0072] S4.3 Output Layer Decision Fusion: This involves fusing the node feature sequences of the final output of the visual stream. and auditory flow feature sequences Perform global average pooling to obtain the global vector. and Introducing adaptive fusion weight calculation:

[0073] The LLM dynamically generates integration weights based on the specific steps of the "Lotus Pond in Moonlight" lesson. Weight It is dynamically generated by the network based on the current context:

[0074] This allows the model to autonomously determine modal weights based on the importance of the scene, in the text parsing stage of this scenario. The value was 0.6 during the group discussion. It is 0.5.

[0075] In an optional embodiment, the method for analyzing interactive language classroom interactions based on line-following interactive modeling provided by the present invention involves concatenating the unified multimodal interactive representation with a language teaching scenario-specific prompt template, and then inputting it again into the large language model. Through multi-task instruction design, the method decodes and outputs structured semantic tuples of the target field, including: Customized prompt templates are designed for different teaching stages in high school Chinese classes. The prompt templates clearly define the constraints of each field of the structured semantic tuple. The target fields include {subject, action, object, responder, text association, and thought state}. The large language model decodes and outputs the structured semantic tuples that conform to the field constraints according to the instructions of the customized prompt template. The values ​​of the action fields are selected from a preset set of actions related to teaching interaction. The object fields are associated with paragraphs, literary elements or discussion topics in high school Chinese textbooks. The text association fields correspond to specific sentences in the textbook.

[0076] Furthermore, the statistical generation of classroom interaction density and text association metrics based on the continuously output structured semantic tuples includes: Based on the structured semantic tuple sequence output within a continuous time window, quantitative evaluation indicators including classroom interaction density, text relevance, interaction type distribution, and the proportion of thinking states are statistically generated.

[0077] In this embodiment, a customized prompt template is designed based on the characteristics of high school Chinese language teaching scenarios. Prompt templates are customized for two classroom segments of "Lotus Pond in the Moonlight": "Teacher-Student Interactive Text Analysis" and "Group Discussion and Discussion." The field constraints of the structured semantic tuple are clearly defined, with the fields fixed as {Subject, Action, Object, Responder, Text Relationship, Mind State}. Standardized instructions are input into the LLM: 1) For the "Teacher-Student Interactive Text Analysis" stage, input a customized prompt template into the LLM: {"Unified Multimodal Interactive Representation Based on the Text Analysis Stage of the High School Chinese Text 'Lotus Pond in the Moonlight'" The decoded output (subject, action, object, responder, text association, and thought state) structured semantic tuples must meet the following requirements: 1. The action field should be selected from 'lecturing / questioning / directing / responding / supplementing'; 2. The object field should be associated with paragraphs or literary elements (such as sensory descriptions and imagery) from the text "Lotus Pond in the Moonlight"; 3. The text association field should clearly correspond to specific sentences in the text; 4. The thought state field should align with the thinking characteristics of Chinese text interpretation and be selected from 'close reading of the text / guided text analysis / expression of viewpoints'. 2) For the "group discussion and seminar" segment, input a customized prompt template into the LLM: {"Multimodal interactive representation based on the group discussion segment of the high school Chinese text 'Lotus Pond in the Moonlight'"} The output should be a structured tuple (subject, action, object, responder, text association, and thought state), satisfying the following conditions: 1. The action field should be selected from 'speak / respond / refute / supplement / reach consensus'; 2. The object field should be associated with the discussion topic of "Lotus Pond in Moonlight" (e.g., the role of olfactory description, emotional expression); 3. The text association field should clearly correspond to relevant sentences in the text; 4. The thought state field should be selected from 'exchange of viewpoints / logical debate / consensus formation'. Step 5.2: LLM decodes and outputs structured semantic tuples according to customized prompt template instructions, accurately presenting the core information of classroom interaction. At the same time, based on the continuous sequence of structured semantic tuples, it statistically quantifies evaluation indicators.

[0078] 1) Output of the "Teacher-Student Interactive Text Analysis" segment: {The teacher explains the description of the moonlight in the third paragraph of "Moonlight over the Lotus Pond," and the whole class learns from it: "The moonlight, like flowing water, quietly pours onto these leaves and flowers," inspiring the teacher to read the text carefully.} {Teacher asks a question about the sensory description type in paragraph 3 of "Lotus Pond in the Moonlight," specifically for learner A: "The leaves and flowers seemed to have been washed in milk," and "a faint fragrance lingered." The teacher calmly guides the text analysis.} {Learner A responds to the visual descriptions in "Moonlight over the Lotus Pond," while the teacher uses phrases like "moonlight like flowing water," "dark shadows," and "graceful figures" to express a tense interpretation of the text.} {Learner B adds: the olfactory description in "Lotus Pond in Moonlight," teacher and learner A, "wisps of fragrance," close reading of the text}.

[0079] 2) Output of the "Group Discussion and Seminar" session: {Learner A speaks on the role of olfactory description in "Lotus Pond by Moonlight," while learner B excitedly exchanges views on "wisps of fragrance."} {Learner B refutes Learner A's view of "lively atmosphere," while Learner A, with "moonlight like flowing water" and "thin blue mist," passionately debates the logic}; {Learner C, supplementing the connection between the olfactory description and the overall tone of the text; Learners A and B, "a quiet night" and "a faint sorrow," calmly supplementing the logic}; {Learners A, B, and C reach a consensus: the olfactory description in "Lotus Pond in Moonlight" evokes a tranquil state of mind; there are no instances of "wisps of fragrance" or "the silence surrounding the lotus pond," thus forming a consensus.}

[0080] Decoding process: For each field of the semantic tuple c (Subject, Action), the decoder independently calculates a classification probability distribution:

[0081] in, The function is defined as: The output is transformed into a probability distribution. Multi-task loss balancing strategy:

[0082] in The loss weights are all set to 1.0.

[0083] Training objective: During model training, optimize by minimizing the sum of cross-entropy losses across all fields.

[0084] in C For a collection of fields, M For batch size, For fields c Number of categories It is a sample m In the field c The true category label (one-hot encoding) on ​​the surface. That is the corresponding predicted probability.

[0085] Step 5.3: Quantify the evaluation indicators based on the structured semantic tuple sequence to provide data support for teaching quality evaluation. At the same time, determine the teaching type of teachers and generate personalized teaching suggestions based on the indicator results.

[0086] 1) Evaluation of the "Teacher-Student Interaction in Text Analysis" segment: 1. Interaction density: 3 effective interactions within a 5-second window, corresponding to 36 interactions / minute (quality judged as "excellent"); 2. Text relevance: 90% of the interactive events are directly related to the core sentences of the text "Lotus Pond in the Moonlight" (quality judged as "excellent"); 3. Interaction type distribution: There are 2 teacher-learner Q&A sessions and 1 learner-learner supplement session, with a balanced interaction type (quality is judged as "good").

[0087] 4. Percentage of thinking state: 40% text analysis guidance, 30% opinion expression, and 30% close reading of the text. The thinking guidance is strong (quality is judged as "excellent").

[0088] 2) Example of evaluation for the "Group Discussion and Seminar" session: 1. Interaction density: 4 effective interactions within a 5-second window, corresponding to 48 interactions / minute (quality judged as "excellent"); 2. Text relevance: 85% of the interactive events are related to the discussion topics and relevant statements in "Lotus Pond in the Moonlight" (quality is judged as "excellent"); 3. Percentage of thinking states: 1 exchange of viewpoints, 1 logical debate, 1 logical supplement, and 1 consensus formation, with complete thinking levels (quality judged as "excellent").

[0089] 4. Percentage of Thinking Processes: 25% exchange of viewpoints, 25% logical debate, 25% textual evidence, 25% consensus formation; complete thinking structure (quality rated "Excellent"). Step 5.4: Based on quantitative indicators and classroom interaction data, determine the teacher's teaching type (in this embodiment, it is determined to be [single interaction type]) and generate personalized teaching suggestions that fit the teaching scenario of "Lotus Pond in the Moonlight"; Table 1 transforms the evaluation indicators (interaction density, type distribution, subject association) into specific data + score mapping + quality judgment, which not only echoes the evaluation results of the two major stages of "teacher-student interaction to analyze text" and "group exchange and discussion", but also provides a quantifiable standard for teaching quality evaluation.

[0090]

[0091] Table 1. Quantitative Evaluation Table of Interactive Classroom Interaction Indicators for High School Chinese Lesson "Moonlight over the Lotus Pond" Table 2, based on the quantitative data of interactive activities in the "Lotus Pond Under the Moonlight" classroom, focuses on the core teaching performance of young teachers in interactive high school Chinese classes. It extracts key indicators from five dimensions: "interaction structure, student initiative, subject suitability, depth of thinking, and quality of basic interaction," to achieve a closed-loop evaluation of "data quantification - quality level - teaching type determination - optimization suggestions."

[0092]

[0093] Table 2. Summary Table of Core Data Assessment and Teaching Type Determination for Young Teachers' Teaching Quality After training, the model parameters are fixed and deployed to the classroom edge server. The system processes the audio and video streams of the "Lotus Pond in Moonlight" class in real time in a pipeline manner: executing steps S1->S2->S3->S4->S5, and outputting a structured sequence of interactive events arranged by timestamps. This sequence can be directly called by upper-level application systems to generate advanced teaching analysis reports such as classroom interaction density heatmaps (e.g., interaction density = number of interaction events / total number of time segments), statistical analysis of teacher-student dialogue rounds, and ST diagrams of teaching behaviors, providing objective and detailed data support for the supervision and quality assessment of interactive classroom teaching in high school Chinese.

[0094] In summary, this invention provides a complete interactive analysis method for high school Chinese classrooms, from low-level multimodal perception to high-level instructional semantic analysis. It is precisely adapted to core scenarios such as text interpretation and group discussions, effectively promoting the development of high school Chinese classroom teaching analysis towards automation, intelligence, and refinement, and providing objective and efficient technical support for teaching quality assessment and optimization.

[0095] This invention also provides a language classroom interaction analysis device based on line-following interactive modeling, comprising: The multimodal feature data acquisition module is used to acquire the visual modal features of the classroom video stream and the auditory modal features of the classroom audio stream to obtain multimodal feature data. The first data processing module is used to construct a tracking spatiotemporal dynamic interaction graph of teachers and students based on multimodal feature data within a preset time window, and to convert the graph features of the tracking spatiotemporal dynamic interaction graph of teachers and students into a structured text sequence that can be parsed by a large language model. The first data fusion module is used to execute a hierarchical multimodal fusion strategy, inputting the structured text sequence into a large language model and outputting multimodal interaction feature embeddings. The second data fusion module is used to embed the multimodal interaction features as cross-modal guidance signals, dynamically modulate the visual modal features through a cross-attention mechanism, and perform weighted fusion with the auditory modal features to generate a unified multimodal interaction representation. The second data processing module is used to combine the unified multimodal interaction representation with the language teaching scenario-specific prompt template and input it again into the large language model. Through multi-task instruction design, it decodes and outputs the structured semantic tuple of the target field. The analysis module is used to statistically generate quantitative evaluation indicators for classroom interaction density and text association based on the continuously output structured semantic tuples, thereby realizing classroom teaching evaluation.

[0096] It should be noted that the language classroom interaction analysis device based on line-following interaction modeling provided in this embodiment of the invention can execute the language classroom interaction analysis method based on line-following interaction modeling described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.

[0097] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to execute the Chinese language classroom interaction analysis method based on line-following interactive modeling provided in the above embodiments.

[0098] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the Chinese language classroom interaction analysis method based on line-following interactive modeling provided in the above embodiments.

[0099] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0100] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for analyzing classroom interaction in Chinese language teaching based on line-following interaction modeling, characterized in that, include: The visual modal features of the classroom video stream and the auditory modal features of the classroom audio stream are obtained to obtain multimodal feature data. Based on multimodal feature data within a preset time window, a tracking spatiotemporal dynamic interaction graph of teachers and students is constructed, and the graph features of the tracking spatiotemporal dynamic interaction graph of teachers and students are transformed into a structured text sequence that can be parsed by a large language model. A hierarchical multimodal fusion strategy is implemented, inputting the structured text sequence into a large language model and outputting multimodal interaction feature embeddings; The multimodal interaction features are embedded as cross-modal guidance signals, and the visual modal features are dynamically modulated through a cross-attention mechanism and weighted and fused with the auditory modal features to generate a unified multimodal interaction representation. After combining the unified multimodal interaction representation with the language teaching scenario-specific prompt template, the data is input into the large language model again. Through multi-task instruction design, the structured semantic tuple of the target field is decoded and output. Based on the continuous output of the structured semantic tuples, a quantitative evaluation index for classroom interaction density and text association is generated to realize classroom teaching evaluation.

2. The method for analyzing language classroom interaction based on line-following interactive modeling according to claim 1, characterized in that, The visual modal features of the acquired classroom video stream include: The visual modal features include the two-dimensional joint coordinates and limb direction unit vectors of individual teachers and students; The video stream is input into the visual modality feature extraction subnetwork for frame-by-frame processing to obtain the two-dimensional key point coordinates of each teacher and student at each time step; Based on the teacher and student's nose tip joint coordinates and neck joint coordinates in the two-dimensional joint coordinates, calculate the limb direction unit vector representing the head orientation, and based on the wrist joint coordinates and shoulder joint coordinates in the two-dimensional joint coordinates, calculate the limb direction unit vector representing the arm orientation.

3. The method for analyzing classroom interaction in Chinese language teaching based on line-following interactive modeling according to claim 2, characterized in that, The auditory modal features of the classroom audio stream include: The auditory modal features include vocal cord motion feature vectors and speaker state coding features; The classroom video stream is input into the vocal cord motion feature processing network to perform laryngeal region localization and vocal cord motion feature extraction, resulting in a vocal cord motion feature vector. The classroom audio stream is input into the auditory modality feature extraction subnetwork for feature extraction, speech activity detection, speaker separation and recognition, to obtain the speaker state coding features at each time step.

4. The method for analyzing language classroom interaction based on line-following interactive modeling according to claim 1, characterized in that, The process of converting the graph features of the tracking-based spatiotemporal dynamic interaction graph between teachers and students into a structured text sequence that can be parsed by a large language model includes: Following the fixed format of "time step, subject, feature, scene prompt", the features of each individual node in the tracking spatiotemporal dynamic interaction diagram of teachers and students at each time step are converted into text entries. All text entries are arranged in chronological order, and features of the Chinese language classroom scene that identify the current teaching segment are embedded in the sequence to obtain a structured text sequence.

5. The method for analyzing classroom interaction in Chinese language teaching based on line-following interactive modeling according to claim 1, characterized in that, The step of inputting the structured text sequence into a large language model and outputting multimodal interaction feature embeddings includes: Based on the limb orientation and vocal cord features in the structured text sequence, the spatial orientation relationship and emotional state of teachers and students are identified and spatial association labels and emotional labels are output. The coherent behavioral logic of teachers and students within the time step is explored and a trace-following temporal interaction chain is constructed. It also includes multimodal interaction feature embedding that extracts and integrates multidimensional information such as semantic relevance, geometric association, emotional state and temporal logic.

6. The method for analyzing language classroom interaction based on line-following interactive modeling according to claim 1, characterized in that, The implementation of the hierarchical multimodal fusion strategy includes: When constructing the tracking-based spatiotemporal dynamic interaction graph between teachers and students, the visual modal features and the auditory modal features are spliced ​​and fused to obtain the initial node features of the interaction graph. The multimodal interaction features are embedded as cross-modal guidance signals, and the visual modal features are dynamically modulated through a cross-attention mechanism to obtain modulated visual features. The modulated visual features are globally pooled to obtain a visual global representation, which is then weighted and fused with the auditory global representation obtained by globally pooling the auditory modal features to generate a multimodal interaction feature embedding.

7. The method for analyzing language classroom interaction based on line-following interactive modeling according to claim 1, characterized in that, The process of concatenating the unified multimodal interaction representation with the language teaching scenario-specific prompt template and then inputting it again into the large language model, through multi-task instruction design, decodes and outputs the structured semantic tuple of the target field, including: Customized prompt templates are designed for different teaching stages in high school Chinese classes. The prompt templates clearly define the constraints of each field of the structured semantic tuple. The target fields include {subject, action, object, responder, text association, and thought state}. The large language model decodes and outputs the structured semantic tuples that conform to the field constraints according to the instructions of the customized prompt template. The values ​​of the action fields are selected from a preset set of actions related to teaching interaction. The object fields are associated with paragraphs, literary elements or discussion topics in high school Chinese textbooks. The text association fields correspond to specific sentences in the textbook.

8. The method for analyzing language classroom interaction based on line-following interactive modeling according to claim 1, characterized in that, The statistical generation of classroom interaction density and text association metrics based on the structured semantic tuples of continuous output includes: Based on the structured semantic tuple sequence output within a continuous time window, quantitative evaluation indicators including classroom interaction density, text relevance, interaction type distribution, and the proportion of thinking states are statistically generated.

9. A language arts classroom interaction analysis device based on line-following interactive modeling, characterized in that, include: The multimodal feature data acquisition module is used to acquire the visual modal features of the classroom video stream and the auditory modal features of the classroom audio stream to obtain multimodal feature data. The first data processing module is used to construct a tracking spatiotemporal dynamic interaction graph of teachers and students based on multimodal feature data within a preset time window, and to convert the graph features of the tracking spatiotemporal dynamic interaction graph of teachers and students into a structured text sequence that can be parsed by a large language model. The first data fusion module is used to execute a hierarchical multimodal fusion strategy, inputting the structured text sequence into a large language model and outputting multimodal interaction feature embeddings. The second data fusion module is used to embed the multimodal interaction features as cross-modal guidance signals, dynamically modulate the visual modal features through a cross-attention mechanism, and perform weighted fusion with the auditory modal features to generate a unified multimodal interaction representation. The second data processing module is used to combine the unified multimodal interaction representation with the language teaching scenario-specific prompt template and input it again into the large language model. Through multi-task instruction design, it decodes and outputs the structured semantic tuple of the target field. The analysis module is used to statistically generate quantitative evaluation indicators for classroom interaction density and text association based on the continuously output structured semantic tuples, thereby realizing classroom teaching evaluation.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the Chinese language classroom interaction analysis method based on line-following interaction modeling as described in any one of claims 1 to 8.