AI-assisted teaching methods and systems using cognitive big data models
By acquiring multimodal teaching data and using cognitive big data models to generate supplementary teaching results, the problem of low classroom teaching efficiency has been solved, realizing the intelligentization and automation of the entire teaching process, and improving teaching efficiency and personalized feedback.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-04-03
AI Technical Summary
Existing educational technology systems have failed to effectively address the problem of low classroom teaching efficiency, insufficient teacher preparation and interaction, inability to meet individualized needs, and delayed feedback.
By acquiring multimodal teaching qualification data, heterogeneous data analysis and feature extraction are performed to generate multimodal features, update the standardized resource library, and use a cognitive big model to generate auxiliary teaching results, including instructions for segmented lectures, equipment control, and classroom interaction.
It has enabled the intelligentization and automation of the entire teaching process, improved classroom teaching efficiency, reduced teachers' mechanical labor, and enhanced personalized teaching and real-time feedback.
Smart Images

Figure CN121213309B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of educational resource technology, and in particular to an AI-assisted teaching method and system that applies a large cognitive model. Background Technology
[0002] With the deepening of educational informatization, artificial intelligence technology is being widely applied to teaching scenarios to improve teaching efficiency and quality. However, current mainstream technical solutions mostly focus on after-school self-study tutoring and static resource management, such as online question banks, recorded course platforms, and learning management systems (LMS). These systems have achieved the digital distribution of educational resources to a certain extent, but they have not penetrated into the core aspects of classroom teaching (such as lesson preparation, teacher-student interaction, and real-time feedback), and therefore cannot effectively solve the core pain point of low classroom teaching efficiency.
[0003] On the one hand, teachers are bogged down in repetitive and mechanical tasks such as lesson preparation and homework correction, which severely squeezes their time for effective instructional design, resulting in low teaching efficiency. On the other hand, the traditional "one-size-fits-all" teaching model cannot meet the individual needs of students, resulting in insufficient classroom interaction and delayed feedback, leading to low efficiency in knowledge transmission. Summary of the Invention
[0004] This invention provides an AI-assisted teaching method and system that applies a large cognitive model, solving the technical problem of low teaching efficiency and improving classroom teaching efficiency.
[0005] In a first aspect, the present invention provides an AI-assisted teaching method using a cognitive big data model. This method includes: acquiring multimodal teaching data during the teaching process, including static and dynamic data. Static data includes courseware, instructional design, blackboard writing, micro-lessons on key points and difficulties, lesson guidance, and lesson evaluation; dynamic data includes classroom audio, student response text, and interactive behavior logs. Based on the multimodal teaching data, the method performs analysis and feature extraction using heterogeneous data parsing and feature extraction techniques to determine multimodal features. Based on the multimodal features, the method performs data clustering and mapping to update the standardized resource library. Based on the updated standardized resource library and teacher input instructions, the method uses a cognitive big data model to generate assisted teaching results. The teacher input instructions include at least one of the following: segmented lecture instructions, device control instructions, classroom interaction instructions, tiered questions and answers, classroom summary instructions, and student answer and comment instructions.
[0006] Secondly, embodiments of the present invention provide an AI-assisted teaching device using a cognitive big data model. This device includes a communication module and a processing module. The communication module acquires multimodal teaching data during the teaching process. This multimodal teaching data includes static and dynamic data. Static data includes courseware, instructional design, blackboard writing, micro-lessons on key points and difficulties, lesson guidance, and lesson evaluation. Dynamic data includes classroom audio, student response text, and interactive behavior logs. The processing module analyzes and extracts features from the multimodal teaching data using heterogeneous data analysis and feature extraction techniques to determine multimodal features. Based on these multimodal features, it performs data clustering and mapping to update the standardized resource library. Based on the updated standardized resource library and teacher input instructions, it generates assisted teaching results using a cognitive big data model. The teacher input instructions include at least one of the following: segmented lecture instructions, device control instructions, classroom interaction instructions, tiered questions and answers, classroom summary instructions, and student answer and comment instructions.
[0007] Thirdly, embodiments of the present invention provide an AI-assisted teaching system for applying a cognitive big model. The system includes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is used to call and run the computer program stored in the memory to perform the steps of the method as described in the first aspect and any possible implementation thereof.
[0008] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of the method as described in the first aspect and any possible implementation thereof.
[0009] This invention provides an AI-assisted teaching method and system that applies a cognitive big data model. The invention analyzes and extracts multimodal teaching data, including courseware, instructional designs, blackboard writing, micro-lessons on key points and difficulties, lesson guidance, lesson evaluation, classroom audio, student response text, and interactive behavior logs, to obtain multimodal features. These features are then used for data clustering and mapping to update a standardized resource library, improving the utilization rate of teaching resources. Subsequently, based on different teacher input instructions, the cognitive big data model and standardized resource library are used to assist teaching, achieving intelligent and automated teaching throughout the entire process. This solves the technical problem of low teaching efficiency and improves classroom teaching efficiency. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart illustrating an AI-assisted teaching method using a large cognitive model, as provided in an embodiment of the present invention.
[0012] Figure 2 This is a schematic diagram of the structure of an AI course-assisted teaching device that applies a large cognitive model, provided in an embodiment of the present invention.
[0013] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0014] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0015] In the description of this invention, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" and "more than one" refer to two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0016] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner to facilitate understanding.
[0017] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the steps or modules listed, but may optionally include other steps or modules not listed, or may optionally include other steps or modules inherent to such process, method, product, or device.
[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0019] like Figure 1 As shown, this embodiment of the invention provides an AI-assisted teaching method that applies a large cognitive model. The method includes steps S101-S104.
[0020] S101. Obtain multimodal teaching data during the teaching process.
[0021] In some embodiments, multimodal teaching data includes static data and dynamic data. Static data includes courseware, instructional design, blackboard writing, micro-lessons on key points and difficulties, lesson guidance, and lesson evaluation. Dynamic data includes classroom audio, student response text, and interactive behavior logs.
[0022] In some embodiments, multimodal teaching data refers to data generated during the teaching process that contains multiple different types of information. This data can be obtained from different sensory channels, such as visual (courseware, blackboard writing, video frames, etc.) and auditory (classroom audio), comprehensively reflecting all aspects of teaching.
[0023] For example, static data refers to data that remains relatively fixed before or during the teaching process, such as courseware and instructional designs, which provide the basic framework and content for teaching.
[0024] For example, dynamic data, such as classroom audio and student response text, generated in real time during the teaching process, reflects the interaction and real-time reactions of students during the teaching process.
[0025] For example, the courseware can be directly extracted from the database storing the courseware by connecting with the teaching management system of the school or educational institution, and the courseware content can be read using the corresponding document parsing library.
[0026] For example, instructional design documents are typically written by teachers in a specific format and stored in a designated location. This file is retrieved via a file reading interface, and then text processing techniques are used to extract key information, such as learning objectives and teaching methods.
[0027] For example, the whiteboard uses a whiteboard capture device in a smart classroom to capture the teacher's writing on the electronic whiteboard in real time and store it as an image or in a specific format. For traditional blackboards, a high-definition camera can be installed to capture images, and then image recognition technology can be used to convert the whiteboard content into editable text.
[0028] For example, the key and difficult points of the micro-lessons are obtained from educational resource platforms or teachers' personal storage spaces. The videos are then preliminarily processed using video processing tools for subsequent analysis.
[0029] For example, similar to obtaining instructional designs, lesson-based learning guidance documents are obtained from the teaching management system or a storage location specified by the teacher, and key information is extracted using text processing technology.
[0030] For example, lesson evaluation data, including student scores and comments, is obtained through a teaching evaluation system. Relevant data is then extracted from the evaluation system's database using database query statements.
[0031] For example, classroom audio systems install multiple microphone arrays in the classroom to capture the voice signals of teachers and students in real time. An audio acquisition card converts the analog voice signals into digital signals, and the digital voice data is then transmitted to a computer for processing via an audio processing interface.
[0032] For example, student response texts are collected in real time via electronic devices used by students. These devices send the response texts to a teaching server via a network, where the server receives and stores this text data using a message queue.
[0033] For example, interaction behavior logs record teacher-student interactions in teaching platforms or classroom management systems, such as student questions, teacher responses, and group discussions. Log recording tools are used to write this interaction information into log files according to a specific format. Then, log parsing programs read and analyze these logs to extract key information about the interaction, such as interaction time, participants, and interaction type.
[0034] S102. Based on multimodal teaching qualification data, analyze and extract features through heterogeneous data analysis and feature extraction techniques to determine multimodal features.
[0035] In some embodiments, heterogeneous data parsing and feature extraction techniques are used to process data from different data sources with different formats and structures, transforming it into feature representations that computers can understand and analyze. By parsing different types of data, key information is extracted to discover patterns and regularities behind the data.
[0036] In some embodiments, multimodal features integrate feature information extracted from multiple modalities, enabling a more comprehensive description of various factors in the teaching process and providing richer evidence for subsequent data analysis and teaching support.
[0037] As one possible implementation, step S102 can be specifically implemented as steps S1021-S1025.
[0038] S1021. Perform audio-visual separation processing on the key and difficult micro-lessons in static data, extract video key frames, and perform noise reduction and speech recognition on the separated audio to determine the text data and video frame data marked with timestamps.
[0039] For example, this embodiment of the invention utilizes a video processing library to decode the micro-lesson video, separating the audio stream and the video stream. For the audio stream, audio processing tools are used for noise reduction to remove background noise and other interference factors. Then, a speech recognition engine is used to convert the processed audio into text, and timestamps are added to the text to correspond with video frames. For the video stream, keyframes are extracted. Keyframe extraction algorithms based on frame difference or optical flow can be used to select representative video frames as a summary of the video content.
[0040] S1022. Analyze the content of courseware, blackboard writing, teaching design and lesson guidance in static data to identify multiple key knowledge data.
[0041] In some embodiments, each knowledge point data includes text, charts, formulas, and knowledge annotations.
[0042] For example, embodiments of the present invention utilize natural language processing technology to perform word segmentation, part-of-speech tagging, and named entity recognition on text content to extract key information. For charts and formulas, image recognition technology is used to identify the chart type and formula structure, and then convert them into a processable text format. Simultaneously, knowledge annotations are added to each extracted key knowledge point, such as the subject to which the knowledge point belongs and its difficulty level.
[0043] S1023. Perform real-time speech recognition and sentiment analysis on classroom speech in dynamic data to determine the key changes in the teacher's tone and the content of the students' responses.
[0044] In some embodiments, the teacher's tone emphasis variation includes the teacher's tone at each moment and the degree of tone variation at each moment.
[0045] For example, the real-time speech recognition in this embodiment of the invention also uses a speech recognition engine to convert speech into text. For sentiment analysis, a deep learning model can be used. This model, pre-trained on a large amount of sentiment-annotated speech data, can determine the speaker's emotional state based on features such as tone, speed, and volume, thereby determining the teacher's tone changes, including the teacher's tone at different times (e.g., calm, excited, confused, etc.) and the degree of tone change (obtained by calculating the difference in tone features between adjacent times). Simultaneously, student responses are extracted from the text obtained from speech recognition.
[0046] S1024. Perform pattern analysis on the interactive behavior logs in the dynamic data to determine the characteristics of the interactive behavior.
[0047] In some embodiments, interactive behavior characteristics include teacher-student interaction frequency, response time, and participation characteristics.
[0048] For example, embodiments of the present invention utilize data mining algorithms to analyze interactive behavior logs. For instance, association rule mining identifies knowledge points or teaching segments with high frequency of teacher-student interaction; sequence pattern mining analyzes the order and time intervals of interactive behaviors to determine interactive behavior characteristics, including teacher-student interaction frequency, response time, and participation characteristics.
[0049] S1025. Based on text data and video frame data, multiple key knowledge data, changes in teacher tone and student responses, and interactive behavior characteristics, generate multimodal features.
[0050] For example, in this embodiment of the invention, the text data and video frame data obtained by the above processing, multiple key knowledge data, changes in the teacher's tone and emphasis, student responses and interactive behavior features are integrated and organized in a specific format to form a comprehensive feature representation as a multimodal feature.
[0051] S103. Based on multimodal features, perform data clustering and mapping, and update the standardized resource library.
[0052] As one possible implementation, step S103 can be specifically implemented as steps S1031-S1035.
[0053] S1031. Vectorize the multimodal features to generate feature vectors in a unified format.
[0054] For example, embodiments of the present invention can perform Z-score standardization on the features of each modality, making their mean 0 and standard deviation 1, thus eliminating the influence of dimensions. For high-dimensional features (such as image features), principal component analysis is used to reduce them to a preset dimension (such as 512 dimensions). The processed feature vectors of each modality are concatenated and then projected onto a unified, low-dimensional semantic vector space through a fully connected neural network layer. The weights of this network layer are learned through training, aiming to make the vector distance between semantically similar content closer in this space. A fixed-dimensional, uniformly formatted semantic feature vector is output, which comprehensively represents the core semantic information of the original multimodal data.
[0055] S1032. Perform semantic clustering on the feature vectors and determine the clustering results.
[0056] In some embodiments, the clustering results include the semantic content of the feature vectors.
[0057] For example, embodiments of the present invention can set two parameters: neighborhood radius and minimum number of points. The algorithm traverses all feature vectors, grouping density-connected vectors into the same cluster. Vectors that cannot be grouped into any high-density cluster are marked as outliers; these may be new, undefined knowledge points or noisy data. The output is the clustering result, i.e., a set of clusters. Each cluster contains several semantically similar feature vectors, and the cluster center vector is automatically generated. The semantic content of each cluster is jointly represented by all vectors within the cluster, and can be interpreted through keyword extraction or reverse lookup of the cluster center vector.
[0058] S1033. Based on the clustering results, calculate the similarity between the feature vector and each knowledge point in the standardized resource library.
[0059] For example, each knowledge point in the standardized resource repository is also represented by a knowledge point vector, which is trained using expert knowledge or massive amounts of data during the initial construction of the resource repository. For each cluster center vector obtained after clustering, its cosine similarity with all knowledge point vectors in the resource repository is calculated. A similarity matrix is output, where each element represents the semantic similarity between a cluster and an existing knowledge point, with a value range of [-1, 1], where the closer to 1, the more similar the cluster.
[0060] S1034. Based on the clustering results and similarity, perform semantic mapping and association matching to determine the matching results.
[0061] For example, for each cluster, find the existing knowledge point with the highest similarity to the cosine. Set a similarity threshold. If the highest similarity is greater than or equal to the threshold, the cluster is considered to have successfully matched this existing knowledge point. If the highest similarity is less than the threshold, the cluster is considered to be unable to match any existing knowledge point and is a "new knowledge point candidate". Output the matching results, a list, recording which existing knowledge point ID each cluster matched with, or was marked as a "new knowledge point".
[0062] For example, semantic mapping assigns specific semantic labels (i.e., maps them to existing knowledge points in a resource repository) to unlabeled clusters obtained through clustering. Association matching refers to the technical operation of establishing connections between new data and existing knowledge entities by calculating indicators such as similarity.
[0063] S1035. Update the standardized resource library based on the matching results.
[0064] For example, step S1035 can be specifically implemented as steps A1-A3.
[0065] A1. If the matching result indicates that the knowledge node does not exist in the standardized resource library, a new knowledge node is created in the standardized resource library, and the association relationship and weight allocation of the teaching content in the standardized resource library are dynamically updated based on the clustering result and the matching result.
[0066] For example, creating a new knowledge node: the system assigns a globally unique UUID as the node ID to this new cluster. Semantic tag generation: using a keyword extraction algorithm, the system extracts the most representative keywords from all the text features contained in the cluster, and uses the combination of these keywords as the initial semantic tag of the new knowledge node (e.g., "the law of refraction of light - experimental verification").
[0067] For example, updating associations and weights: Relationship discovery calculates the cosine similarity between the new node vector and all existing node vectors in the resource library. The top K nodes with the highest similarity are identified. Relationship establishment: Based on similarity, the system automatically establishes semantic associations between the new node and these existing nodes. Relationship types can be "related knowledge points," "belong to," "prerequisite knowledge," etc. For example, the new node "Law of Refraction of Light - Experimental Verification" will automatically establish a "specific application" relationship with the existing node "Refraction of Light." Weight allocation: The strength (weight) of the relationship is directly determined by their similarity scores. Simultaneously, the initial weight of the new node itself is assigned based on the confidence level of its source data (such as the tightness of the cluster and the authority of the data source).
[0068] A2. If the matching result indicates that the knowledge node exists in the standardized resource library, then the knowledge node in the standardized resource library is dynamically updated based on the matching result.
[0069] For example, node vector updates are performed using an exponential moving average algorithm to update the vector representation of existing knowledge nodes. The formula is: Vnew = α * Vold + (1-α) * V_cluster; where Vold is the original vector of the knowledge node, Vcluster is the center vector of the newly matched cluster, and α is a smoothing factor (usually set to 0.8-0.9) to control the weight ratio of old and new information. This allows for the smooth integration of new knowledge and avoids drastic drift in node vectors due to a single update.
[0070] For example, the node content is enriched: multimodal instances contained in the new cluster (such as a new teaching case video or an excellent classroom lecture audio) are associated with the knowledge node as "teaching examples" or "related materials" to enrich its teaching content. The popularity weight of the node is updated, and the increase in weight is proportional to the similarity score of this match and the magnitude of the new data.
[0071] A3. Perform a consistency check on the updated standardized resource library; if the consistency check passes, the update of the standardized resource library is considered successful.
[0072] In some embodiments, consistency verification includes logical verification and statistical verification.
[0073] For example, logical checks include: Circular association check: Using a graph traversal algorithm to detect circular references in the knowledge graph; if found, an alarm is triggered, requiring manual review. Isolated node check: Scanning the graph to find isolated nodes without any incoming or outgoing edges.
[0074] For example, statistical verification: Before and after the update, a snapshot comparison is performed on the overall statistical indicators of the repository (such as the total number of nodes, the total number of relationships, and the average node degree). If any indicator changes abruptly (such as a surge in the number of nodes), the update is rolled back and an anomaly alert is issued.
[0075] For example, all update operations completed in memory (creating nodes, updating vectors, establishing associations, and modifying weights) are formally committed and persisted to storage as a database transaction. Index rebuilding, to ensure efficient subsequent queries, triggers incremental updates or asynchronous rebuilding of the indexes for both the vector database and the graph database.
[0076] Optionally, step S103 can also be implemented as steps S1036-S1038.
[0077] S1036. Based on the changes in the teacher's tone and the characteristics of teacher-student interaction, dynamically calculate the teaching emphasis weight of each knowledge point.
[0078] In some embodiments, teacher-student interaction features include the frequency of questions asked on each knowledge point and the response time.
[0079] For example, embodiments of the present invention can convert the emphasis of tone into numerical indicators: for instance, the system assigns a higher base score (e.g., 1.2) to an "emphasis" tone and a base score (e.g., 0.8) to a "calm" tone. The degree of tone variation (gradient) itself acts as a multiplier factor; the more drastic the change, the larger the multiplier, indicating a higher level of teacher emotional investment and a more important knowledge point. Teacher-student interaction characteristics are also converted into numerical indicators: the frequency of questioning is directly used as a bonus; the higher the frequency, the more points are accumulated. Response time, on the other hand, is used as a deduction; a shorter average student response time indicates a simpler question or a more familiar knowledge point, potentially indicating lower importance; a longer response time may mean a more complex question, an important knowledge point, or student confusion.
[0080] For example, the system aggregates the aforementioned quantitative scores from all related classroom segments for each knowledge point. An initial "teaching focus weight" is calculated using a weighted summation method. The system assigns different empirical weights to tone data and interaction data to reflect their varying importance.
[0081] S1037. Based on the analysis results of students' responses, dynamically calculate the weight of learning difficulties for each knowledge point.
[0082] For example, the system focuses on students' incorrect or low-similarity responses. For a given knowledge point, the lower the average similarity of students' answers, or the higher the frequency of incorrect answers, the greater the initial assessment of the learning difficulty of that knowledge point. Simultaneously, the system analyzes the types of incorrect answers. If the errors are concentrated and common (e.g., most students make mistakes on the same formula), this is a stronger indication that the point represents a core difficulty than scattered and random errors.
[0083] For example, based on the above analysis, a "learning difficulty weight" is calculated for each knowledge point. This weight is positively correlated with the error rate and error concentration. That is, the higher the error rate of a knowledge point and the more consistent the wrong answers are, the greater its "learning difficulty weight" value.
[0084] S1038. Based on the weights of teaching focus and learning difficulty, the comprehensive weights of each knowledge point in the standardized resource library are corrected to obtain the corrected comprehensive weights of each knowledge point.
[0085] In some embodiments, the overall weight is used to characterize the importance of each knowledge point.
[0086] For example, the system uses a multi-factor fusion model to integrate the two weights. The basic logic of the model is that the ultimate importance of a knowledge point depends on both whether the teacher emphasizes it (teaching emphasis weight) and whether students generally find it difficult (learning difficulty weight).
[0087] For example, for knowledge points with high weighting for both teaching focus and learning difficulty, their overall weight will be significantly increased. This indicates that these are "core difficulties" that teachers value but students find difficult to grasp, requiring more resources in subsequent teaching. For knowledge points with high weighting for teaching focus but low weighting for learning difficulty, their overall weight will be moderately increased or maintained. This indicates that these are "core foundational" knowledge points that teachers value and students have already mastered relatively well. For knowledge points with low weighting for teaching focus but high weighting for learning difficulty, their overall weight will be significantly increased. This is a very important signal, indicating that this is a "hidden difficulty" that teachers may overlook in their teaching, requiring special prompting from the system. For knowledge points with low weighting in both areas, their overall weight will be decreased or maintained.
[0088] For example, the revised overall weight will be written back to the attribute fields of each knowledge point in the standardized resource library in real time. This weight value will become the key decision-making basis for all subsequent intelligent functions. High-weight knowledge points and related explanation materials for difficult points will be prioritized for teachers. During test paper generation, a higher proportion of high-weight knowledge points, especially questions with high difficulty weights, will be extracted. Path planning: In learning path recommendation, more learning time and practice opportunities will be allocated to high-weight knowledge points. Learning analysis: A "distribution map of key points and difficulties" for the class will be presented to teachers, intuitively showing which important knowledge points students have not mastered well.
[0089] It should be noted that this invention introduces a dynamic, data-driven weight adjustment mechanism, which transforms the standardized resource base from a static collection of knowledge into an "intelligent entity" capable of self-awareness of teaching focus and learning difficulties. By quantitatively analyzing real-time feedback data (tone, interaction, and responses) from both the teaching and learning sides in the classroom, the system can accurately and automatically identify the key points and difficulties in the teaching process and dynamically adjust the value judgments of the knowledge system accordingly. This greatly improves the accuracy and automation level of personalized teaching.
[0090] S104. Based on the updated standardized resource library and teacher input instructions, use the cognitive big model to generate auxiliary teaching results.
[0091] In some embodiments, teacher input instructions include at least one of the following: segmented lecture instructions, device control instructions, classroom interaction instructions, tiered questions and answers, classroom summary instructions, and student response and comment instructions.
[0092] For example, when the teacher inputs a segmented lecture instruction, step S104 can be specifically implemented as steps B11-B14.
[0093] B11. Analyze the segmented lecture instructions input by the teacher to determine the teaching topic and target knowledge points.
[0094] For example, the system receives a teacher's natural language instruction, such as "Prepare a lesson to explain Newton's First Law." An intent recognition model is used to determine that the instruction belongs to the "segmented lecture" type.
[0095] For example, named entity recognition technology is then used to extract key entities from the instruction text. These include core themes (such as "Newton's First Law") and limiting conditions (such as "45 minutes" and "for first-year high school students").
[0096] For example, knowledge point mapping involves matching the extracted core topics with the knowledge graph in a standardized resource repository to precisely locate one or more target knowledge point nodes. For instance, "Newton's First Law" can be mapped to the knowledge point node with the ID "Physics-Mechanics-NFL" in the resource repository, and all its attributes and associated content can be retrieved.
[0097] B12. Retrieve teaching content related to the target knowledge points from the updated standardized resource library.
[0098] In some embodiments, the teaching content includes concept explanations, case studies, and common student misconceptions.
[0099] For example, the system uses the target knowledge point as the core and performs a traversal search within its first- or second-order neighborhood of the knowledge graph. This is not a simple keyword matching, but a semantic retrieval based on vector similarity.
[0100] For example, multimodal content retrieval: the retrieved content is multimodal and highly structured, including authoritative textual definitions of the knowledge point, relevant formulas, and embedded micro-lecture video clips. Case materials include positive examples, negative examples, and real-life application examples (such as images, news clips, and animated demonstrations).
[0101] B13. Based on the teaching content, utilize the cognitive big model to generate structured lecture content.
[0102] In some embodiments, the structured lecture content includes multiple time segments, each time segment corresponds to a core sub-knowledge point, and each time segment has preset interactive trigger points; the interactive trigger points include instructions to start a question session, a group discussion session, or an experimental demonstration session, and the interactive trigger points are used to automatically trigger interactive activities during the lecture.
[0103] For example, the system constructs a highly structured prompt word, incorporating the retrieved teaching content as context. This prompt word explicitly instructs the large model to act as a "senior teaching expert" and requires it to output in a time-segment format.
[0104] For example, intelligent insertion of interactive points: The large model will intelligently preset interactive trigger points at appropriate locations based on the content of sub-knowledge points. For instance, after explaining a concept, it will automatically insert "[Question: Please give an example of inertia in life]"; before proceeding with formula derivation, it will insert "[Group discussion: Predict the experimental results]".
[0105] B14. Based on the teaching progress logic, sort the generated structured lecture content and adapt it with the PPT courseware to generate auxiliary teaching results corresponding to the segmented lecture instructions.
[0106] In some embodiments, supplementary teaching outcomes include structured PowerPoint presentations, segmented speech scripts, and audio-visual content.
[0107] For example, logical sorting and optimization: The generated multiple content segments are processed by a teaching logic verification module. This module ensures that the content order conforms to teaching principles such as "from simple to complex" and "from concrete to abstract." For example, it ensures that "introducing concepts" precedes "deriving formulas." Automatic adaptation with PPT courseware: The system automatically associates the sorted structured text content with the teacher's pre-uploaded PPT courseware. By analyzing the text and image content of each PPT page, the system intelligently matches each lecture segment to the most relevant PPT page number. Finally, the system packages and generates an executable "lecture package," which is the aforementioned auxiliary teaching result.
[0108] The system includes: Structured PPT presentations: Original PPT slides are marked with indicators to specify which segment of the presentation is triggered when a slide is turned; Segmented speech scripts: A detailed, time-stamped script clearly outlines what the teacher should say in each segment; and Voice playback: The system can use text-to-speech technology to convert the speech script into fluent AI voice for classroom presentation assistance or to provide demonstrations for teachers.
[0109] It should be noted that this invention provides a highly automated, intelligent, and deeply structured lesson preparation process. This invention organizes scattered multimodal teaching content into structured knowledge that can be accessed by a large model. It creatively requires the large model to generate content according to "time segments" and intelligently inserts "interactive trigger points," directly integrating teaching strategies into the lecture notes. This achieves automatic adaptation between the generated lecture notes and PPT courseware, forming a synchronized audio-visual "lecture package" that can be directly used in the classroom, greatly improving lesson preparation efficiency and teaching quality.
[0110] For example, when the teacher inputs a device control command, step S104 can be specifically implemented as steps B21-B24.
[0111] B21. Analyze the equipment control commands input by the teacher to determine the target teaching equipment and control actions to be controlled.
[0112] In some embodiments, the target teaching equipment includes a smart screen, microphone, sound system, or lighting equipment.
[0113] For example, the system receives device control commands input by the teacher through a specific communication interface. These commands may be transmitted in text form, such as "Turn on the smart screen" or "Adjust the speaker volume to 70%." The system first performs preliminary parsing of the received command, identifying the types of key information it may contain, and determining that it is a device control command.
[0114] For example, embodiments of the present invention determine the target teaching device to be controlled through keyword matching. For instance, if the instruction contains keywords such as "smart screen" or "screen," the target teaching device is determined to be a smart screen; if it contains words such as "microphone" or "microphone," the target device is a microphone. The system maintains a device keyword database internally to accurately identify different devices.
[0115] For example, the control action determination in this embodiment of the invention is also based on keyword matching and semantic understanding. For instance, "on" and "off" correspond to the on and off operations of the device; "adjust parameters" may involve adjusting parameters such as volume and brightness; and "mode switching" corresponds to the switching between different working modes of the device, such as switching between normal mode and surround sound mode of an audio system.
[0116] B22. Obtain the device control protocol template from the updated standardized resource library, and encapsulate the control instructions based on the STRClass protocol and control actions.
[0117] In some embodiments, control commands include turning the device on, turning it off, adjusting parameters, or switching modes.
[0118] For example, the system retrieves device control protocol templates by storing control protocol templates for various teaching devices in an updated standardized resource library. Based on the identified target teaching device, the system searches the resource library for the corresponding device control protocol template.
[0119] For example, the STRClass protocol is a specific protocol for device control that defines rules for instruction encoding, data transmission formats, and other related aspects. The system encodes control actions according to the requirements of the STRClass protocol. For instance, the "on" action is encoded as a specific binary or hexadecimal value, and parameter adjustment values are converted according to the format specified in the protocol.
[0120] For example, control command encapsulation involves combining the device control protocol template and coded control action information into a complete control command. During encapsulation, it's ensured that the command conforms to the device's communication requirements, including command length and checksum settings. For instance, a checksum is added to the end of the command to guarantee the accuracy of command transmission.
[0121] B23. Send the encapsulated control commands to the target teaching equipment to control the target teaching equipment to perform control actions.
[0122] For example, in this embodiment of the invention, a suitable communication method is selected to send control commands based on the communication interface and characteristics of the target teaching device. Common communication methods include wired network communication, wireless network communication, and serial port communication. For instance, for teaching devices that support network connectivity, control commands are sent to the device's IP address and designated port via a network socket; for devices connected via a serial port, the commands are sent byte-by-byte using a serial communication library.
[0123] For example, the system sends the encapsulated control commands to the target teaching device according to the selected communication method. Considering the possibility of data loss or errors during communication, a command retry mechanism is set up. If no response is received from the device within a specified time or an erroneous response is received, the system will resend the control commands. The number of retries can be configured according to the actual situation, generally set to 3-5 times.
[0124] B24. Receive action feedback from the target teaching equipment and generate auxiliary teaching results based on control commands and action feedback.
[0125] For example, after receiving a control command and executing the corresponding action, the target teaching device sends action feedback information to the system. The system receives this feedback information through the same communication interface. The feedback information may include the device's action execution status (success or failure), current device parameter values, etc. For instance, a smart screen will display "Successfully turned on" information after being turned on, and a speaker will display the current volume value after adjusting the volume.
[0126] For example, the system parses the received action feedback information to extract key information, such as the action execution result and device status parameters. By matching this information with a predefined feedback format, the system ensures correct understanding of the feedback content. For instance, it parses the status code from the feedback information to determine whether the device action was successfully executed.
[0127] For example, based on control commands and action feedback information, the system generates auxiliary teaching results. These results can be presented in various forms, such as text prompts or graphical interface displays. For instance, if the control command is to turn on the smart screen and is executed successfully, the auxiliary teaching result could be a message displayed on the teacher's interface stating "Smart screen successfully turned on." If the control action fails, the auxiliary teaching result could display the error reason, such as "Smart screen failed to turn on, please check device connection," helping the teacher understand the device control status for subsequent operations.
[0128] For example, when the teacher inputs a classroom interaction instruction, step S104 can be specifically implemented as steps B31-B37.
[0129] B31. Analyze the classroom interaction instructions input by the teacher to determine the interaction type and the target knowledge points.
[0130] For example, teachers can input instructions via voice or text, such as "Now we'll have a quick Q&A session about photosynthesis" or "Organize a group discussion about Newton's laws." The system first uses a semantic understanding model to parse the instructions. Intent and Entity Recognition: The model identifies the core intent of the instruction as "initiating classroom interaction" and extracts two key entities: Interaction Type: such as "quick Q&A," "group discussion," "quiz," "debate," etc. Target Knowledge Point: such as "photosynthesis" and "Newton's laws." The system can understand different ways of expressing knowledge points (such as "Newton's First Law" or "Law of Inertia"). Knowledge Point Mapping: The system precisely matches the identified knowledge point names with knowledge nodes in a standardized resource library to confirm their uniqueness and validity.
[0131] B32. Extract core terms and concepts related to the target knowledge points from the updated standardized resource library.
[0132] For example, the system uses the target knowledge point node as the root and traverses its first-order neighborhood within the knowledge graph. Multimodal content extraction: All relevant core terms and concepts are extracted from connected nodes and edges. This information includes not only keywords and definitions in text form, but also structured content such as relevant formulas, chart labels, and experiment names.
[0133] B33. Use the TF-IDF algorithm to filter out key terms from core terms.
[0134] For example, the system virtually aggregates all text content related to the knowledge point (such as textbook paragraphs, lesson plan instructions, and historical lecture notes) into a single "document." TF-IDF calculation is performed: The core idea of the TF-IDF algorithm is that the higher the frequency of a word in the current knowledge point document (high TF) and the lower its frequency in other knowledge point documents (high IDF), the more representative and important it is of that knowledge point. Key terms are selected: The system calculates the TF-IDF values of all candidate terms and sorts them from highest to lowest score. The top-ranked terms (e.g., 10-15) are then selected as key terms.
[0135] B34. Based on key terms and concepts, utilize a cognitive big model to generate multi-level interactive questions, as well as standard answers and scoring criteria for multi-level interactive questions.
[0136] In some embodiments, multi-level interactive questions include L1 memory questions, L2 application questions, and L3 critical thinking questions.
[0137] For example, the system constructs a sophisticated prompt, the core of which is a large-scale instruction model acting as a "rigorous question-generating expert." The prompt explicitly injects key terms and mandates that the large-scale model generate questions at three levels: L1 Memory-based: This requires generating questions testing basic concepts and factual recall (e.g., "What are the raw materials for photosynthesis?"). L2 Application-based: This requires generating questions testing knowledge application and simple reasoning (e.g., "How will the rate of photosynthesis change if the carbon dioxide concentration increases?"). L3 Critical Thinking-based: This requires generating higher-order thinking questions testing analysis, evaluation, and creativity (e.g., "Please design an experiment to verify the effect of temperature on photosynthesis"). Generating Answers and Scoring Criteria: For each generated question, the large-scale model must simultaneously generate a standard answer and a scoring criterion. The scoring criterion is not a simple "right / wrong" answer, but rather includes key scoring points. For example, for L3 level questions, the scoring criterion lists several key steps that must be included in the experimental design, assigning a corresponding score to each step.
[0138] B35. Develop student response strategies based on interaction types and multi-level interactive questions.
[0139] In some embodiments, student response strategies include individual responses, group discussions, or a buzzer-style response mode.
[0140] For example, the system maintains a strategy mapping table to match different interaction types with the most suitable response strategy. For instance, "Quick Q&A" matches an individual answer strategy, and the system quickly and randomly selects students. "Group discussion" matches a group discussion strategy, and the system automatically groups students based on seating or ability, assigning different discussion sub-topics. "Bubble answer" matches a buzzer-style mode, and the system activates a buzzer system to create a competitive atmosphere. Strategy execution instruction generation: Based on the selected strategy, specific, executable instructions are generated. For example, for individual answering, the instruction is generated: "Student ID 25, please answer"; for group discussion, the instruction is generated: "Front row group discusses question A, back row group discusses question B, timer 5 minutes."
[0141] B36. Accept students' real-time responses and, based on the standard answer and scoring criteria, use the BERT+BiLSTM model to analyze the students' real-time responses, calculate the semantic similarity between the students' real-time responses and the standard answer, and generate personalized feedback comments.
[0142] For example, students can answer via microphone or text input on the terminal. Voice responses are first converted to text through speech recognition. Deep semantic analysis: The BERT+BiLSTM model is used to perform deep semantic encoding and similarity calculation on the student's response text and the standard answer. This model can understand semantically similar but differently expressed answers such as "forces act in pairs" and "the interactions between objects are mutual," and give them high scores. Generating personalized feedback: The model not only outputs a similarity score, but also combines key points from the scoring criteria to generate a personalized verbal feedback. For example: "Correct answer! You accurately mentioned the key concept of 'action and reaction forces.'" or "The answer is close, but you missed the point of 'equal magnitudes,' think again?"
[0143] B37. Based on students' real-time responses, standard answers and scoring criteria, as well as semantic similarity and personalized feedback, generate supplementary teaching results.
[0144] For example, the system synthesizes all the above information into a final supplementary teaching result. This result is a structured data package containing: the original question and standard answer; the student's response; a semantic similarity score (which can be used as a scoring reference); generated personalized commentary (which can be read aloud by the teacher or displayed on the screen); and a record of this interaction, used to update student profiles and the weighting of knowledge point difficulties.
[0145] It should be noted that this invention provides a highly automated, intelligent, and strategically designed classroom interaction generation and execution system. Precise content anchoring: Key terms are accurately extracted from the knowledge base using the TF-IDF algorithm, ensuring that generated questions do not deviate from the teaching focus. Structured hierarchical generation: Through carefully designed prompts, the instruction model generates high-quality question sets with scoring criteria according to cognitive levels, replacing the teacher's repetitive mechanical work. Intelligent evaluation and feedback: Advanced NLP models (BERT+BiLSTM) are used for deep semantic evaluation, providing immediate and personalized verbal feedback, realizing high-value tasks that previously could only be performed by teachers. Complete closed-loop process: From parsing instructions to generating final feedback, a complete interactive teaching closed loop is formed, greatly enriching classroom interaction methods and improving interaction efficiency and teaching effectiveness.
[0146] Optionally, the AI course-assisted teaching method using the big cognitive model provided in this embodiment of the invention further includes steps S201-S204.
[0147] S201. During the execution of segmented lecture instructions, student feedback data is collected in real time.
[0148] In some embodiments, student feedback data includes facial expression attention analysis data, real-time answer accuracy, and interaction engagement.
[0149] For example, the system collects data in real-time and non-intrusively through an IoT sensor array deployed in the classroom. Using high-definition smart cameras deployed at the front of the classroom, it analyzes students' facial orientation, eye movement frequency, and eyelid opening and closing in real time using computer vision models to calculate a classroom attention concentration index. The system can distinguish whether a student is looking at the blackboard, taking notes, or daydreaming. During interactive classroom activities (such as quizzes and voting), the system collects answers through student terminals, grades them instantly, and calculates the class-wide accuracy rate for the current question. System logs record each student's behavior during interactive activities, including the number of times they raise their hand, the number of times they are called upon to answer, and the frequency of messages sent in the discussion area, quantifying their participation through this behavioral data.
[0150] S202. Based on student feedback data, calculate the student comprehension index for the current teaching segment.
[0151] For example, the comprehension index is not a single data point, but a multi-dimensional comprehensive indicator that integrates behavior, performance, and emotion. The system establishes a calculation model that assigns different weights to different types of data. Real-time answer accuracy usually has the highest weight because it directly reflects the mastery of knowledge points. Facial expression attention analysis data serves as an important auxiliary indicator; high attention is generally positively correlated with high comprehension, but situations of "daydreaming due to not understanding" must be excluded. Interactive participation reflects initiative; active participation usually means better follow-up and understanding. The model ultimately outputs a quantitative student comprehension index (e.g., a value between 0 and 100), which represents the overall level of understanding of the knowledge points being taught by the entire class.
[0152] S203. If the comprehension index is lower than the preset threshold, the teaching path adjustment strategy will be automatically triggered.
[0153] In some embodiments, instructional path adjustment strategies include invoking a cognitive big model to generate supplementary cases or analogical explanations in real time, lowering the difficulty level of questions, or generating a "re-explain" instruction and providing feedback to the teacher.
[0154] For example, the system in this embodiment of the invention presets a comprehension threshold (e.g., 60 points). This threshold can be configured by the teacher or system administrator according to the course difficulty and grade level. When the comprehension index calculated by the system in real time remains below this threshold (e.g., for more than 30 seconds), it automatically determines that the current teaching effect has not met expectations, immediately triggers the interruption of the current preset teaching process, and initiates a dynamic adjustment mechanism.
[0155] For example, Strategy 1: Generate supplementary cases or analogies. The system immediately uses the currently taught knowledge points and signals of low comprehension as prompts to invoke the cognitive big model. The big model receives the instruction to "re-explain the current concept using more relatable real-life examples or analogies." The system then inserts the new explanations and cases generated by the big model in real time into the current classroom teaching through voice broadcast or text prompts, as an immediate remedial measure.
[0156] For example, Strategy Two: Lowering the difficulty level of questions. If low comprehension is detected during the interactive Q&A session, the system will automatically adjust the generation strategy for subsequent questions. For instance, if the original plan was to ask a Level 3 critical thinking question, the system will dynamically lower it to a Level 2 application question, or even a Level 1 memory question, to help students build confidence and solidify their foundation.
[0157] For example, Strategy 3: Generating a "re-explain" instruction and providing feedback to the teacher is the most direct intervention strategy. The system will display a prominent prompt on the teacher's control terminal (such as a tablet) and generate a clear voice or text suggestion, such as: "We have detected that the student does not fully understand the concept of 'inertia magnitude.' We suggest you explain it again in a different way." At the same time, the system may also recommend backup explanation materials (such as pictures or short videos) prepared for the teacher to assist in the re-explanation.
[0158] It should be noted that this invention provides a closed-loop feedback teaching system based on real-time data perception. It breaks through the bottleneck of high latency in traditional teaching feedback, achieving second-level monitoring and intervention in the teaching process. Intelligent: It is no longer a simple data display, but rather obtains deep indicators (comprehension index) through fusion calculation, and automatically triggers intelligent adjustment strategies based on this. Diverse: It provides multi-level and diverse adjustment strategies, from automatic machine supplementation to assisting teachers in decision-making, forming a flexible and adaptive teaching closed loop. Enhanced: It transforms the teacher's role from an "isolated speaker" to an "enhanced teacher" with an "AI teaching assistant" providing real-time data support and strategy suggestions, thereby significantly improving the flexibility and efficiency of classroom teaching.
[0159] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0160] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.
[0161] Figure 2 This diagram illustrates the structure of an AI course auxiliary teaching device using a large cognitive model, according to an embodiment of the present invention. The auxiliary teaching device 300 includes a communication module 301 and a processing module 302.
[0162] The communication module 301 is used to acquire multimodal teaching data during the teaching process. The multimodal teaching data includes static data and dynamic data. The static data includes courseware, teaching design, blackboard writing, micro-lessons on key points and difficulties, lesson guidance and lesson evaluation. The dynamic data includes classroom audio, student response text and interactive behavior logs.
[0163] The processing module 302 is used to analyze and extract features from multimodal teacher qualification data using heterogeneous data analysis and feature extraction techniques to determine multimodal features; based on the multimodal features, it performs data clustering and mapping to update the standardized resource library; based on the updated standardized resource library and teacher input instructions, it uses a cognitive big model to generate auxiliary teaching results. The teacher input instructions include at least one of the following: segmented lecture instructions, equipment control instructions, classroom interaction instructions, tiered questions and answers, classroom summary instructions, and student answer and comment instructions.
[0164] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 400 includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, it implements the steps in the above-described method embodiments. Alternatively, when the processor 401 executes the computer program 403, it implements the functions of each module / unit in the above-described device embodiments.
[0165] For example, the computer program 403 may be divided into one or more modules / units, which are stored in the memory 402 and executed by the processor 401 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program 403 in the electronic device 400.
[0166] The processor 401 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0167] The memory 402 can be an internal storage unit of the electronic device 400, such as a hard disk or memory of the electronic device 400. The memory 402 can also be an external storage device of the electronic device 400, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD) card, flash card, etc., equipped on the electronic device 400. Furthermore, the memory 402 can include both internal and external storage units of the electronic device 400. The memory 402 is used to store the computer program and other programs and data required by the terminal. The memory 402 can also be used to temporarily store data that has been output or will be output.
[0168] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for AI-assisted teaching using a large cognitive model, characterized in that, include: Acquire multimodal teaching data during the teaching process. The multimodal teaching data includes static data and dynamic data. The static data includes courseware, instructional design, blackboard writing, micro-lessons on key points and difficulties, lesson guidance, and lesson evaluation. The dynamic data includes classroom audio, student response text, and interactive behavior logs. Based on the aforementioned multimodal teaching qualification data, heterogeneous data parsing and feature extraction techniques are used for analysis and feature extraction to determine multimodal features. These include: real-time speech recognition and sentiment analysis of classroom speech in the dynamic data to determine changes in teacher tone emphasis and student responses; the changes in teacher tone emphasis include the teacher's tone at various times and the degree of tone change at each time; pattern analysis is performed on the interactive behavior logs in the dynamic data to determine interactive behavior features, including teacher-student interaction frequency, response time, and participation characteristics; and multimodal features are generated based on the changes in teacher tone emphasis, student responses, and interactive behavior features. Based on the aforementioned multimodal features, data clustering and mapping are performed to update the standardized resource library, including: dynamically calculating the teaching focus weight of each knowledge point based on changes in teacher tone and teacher-student interaction features; the teacher-student interaction features include the questioning frequency and response time for each knowledge point; dynamically calculating the learning difficulty weight of each knowledge point based on the analysis results of student responses; correcting the comprehensive weight of each knowledge point in the standardized resource library based on the teaching focus weight and learning difficulty weight, obtaining the corrected comprehensive weight of each knowledge point, which is used to characterize the importance of each knowledge point; and writing the corrected comprehensive weight of each knowledge point back to the attribute field of each knowledge point in the standardized resource library in real time as the basis for subsequent intelligent decision-making. Based on the updated standardized resource library and teacher input instructions, the cognitive big model is used to generate auxiliary teaching results. The teacher input instructions include at least one of the following: segmented lecture instructions, equipment control instructions, classroom interaction instructions, tiered questions and answers, classroom summary instructions, and student answer comments instructions.
2. The AI-assisted teaching method using a large cognitive model as described in claim 1, characterized in that, The step of analyzing and extracting features from the multimodal teacher qualification data using heterogeneous data analysis and feature extraction techniques to determine multimodal features also includes: The key and difficult micro-lessons in the static data are subjected to audio-visual separation processing, video key frames are extracted, and the separated audio is subjected to noise reduction processing and speech recognition to determine the text data and video frame data marked with timestamps. The courseware, blackboard writing, teaching design, and lesson guidance in the static data are analyzed to identify multiple key knowledge data. Each key knowledge data includes text, charts, formulas, and knowledge annotations. Based on the text data and video frame data, the multiple key knowledge data, the changes in the teacher's tone and the students' responses, and the interactive behavior features, multimodal features are generated.
3. The AI-assisted teaching method using a large cognitive model as described in claim 1, characterized in that, The step of performing data clustering and mapping based on the multimodal features and updating the standardized resource library also includes: The multimodal features are vectorized to generate feature vectors in a unified format; Semantic clustering is performed on the feature vectors to determine the clustering results, wherein the clustering results include the semantic content of the feature vectors; Based on the clustering results, the similarity between the feature vector and each knowledge point in the standardized resource library is calculated; Based on the clustering results and the similarity, semantic mapping and association matching are performed to determine the matching results; Based on the matching results, the standardized resource library is updated.
4. The AI-assisted teaching method using a large cognitive model according to claim 3, characterized in that, Updating the standardized resource library based on the matching results includes: If the matching result is that the knowledge point is not in the standardized resource library, a new knowledge point is created in the standardized resource library, and the teaching content association and weight allocation in the standardized resource library are dynamically updated according to the clustering result and the matching result. If the matching result indicates that the knowledge point exists in the standardized resource library, then the knowledge point in the standardized resource library is dynamically updated based on the matching result. Perform a consistency check on the updated standardized resource library; if the consistency check passes, the update of the standardized resource library is considered successful.
5. The AI-assisted teaching method using a large cognitive model according to claim 1, characterized in that, The teacher input instructions include segmented lecture instructions; Accordingly, the generation of auxiliary teaching results based on the updated standardized resource library and teacher input instructions, using a cognitive big data model, includes: Analyze the segmented lecture instructions input by the teacher to determine the teaching topic and target knowledge points; Retrieve teaching content associated with the target knowledge point from the updated standardized resource library. The teaching content includes concept explanations, case materials, and common student misconceptions. Based on the teaching content, a structured lecture content is generated using a cognitive big model. The structured lecture content includes multiple time segments, each time segment corresponds to a core sub-knowledge point, and each time segment has preset interactive trigger points. The interactive trigger points include instructions to start a question-and-answer session, a group discussion session, or an experimental demonstration session. The interactive trigger points are used to automatically trigger interactive activities during the lecture. Based on the teaching progress logic, the generated structured lecture content is sorted and adapted to the PPT courseware to generate auxiliary teaching results corresponding to the segmented lecture instructions. The auxiliary teaching results include structured PPT courseware, segmented speech text, and voice broadcast content.
6. The AI-assisted teaching method using a large cognitive model according to claim 5, characterized in that, The method further includes: During the execution of segmented lecture instructions, student feedback data is collected in real time, including facial expression attention analysis data, real-time answer accuracy rate, and interactive participation. Based on the student feedback data, calculate the student comprehension index for the current teaching segment; If the comprehension index is lower than a preset threshold, a teaching path adjustment strategy will be automatically triggered. The teaching path adjustment strategy includes calling the cognitive big model to generate supplementary cases or analogies in real time, lowering the difficulty level of the questions, or generating a "re-explain" instruction and feeding it back to the teacher.
7. The AI-assisted teaching method using a large cognitive model according to claim 1, characterized in that, The teacher input instructions include device control instructions; Accordingly, the generation of auxiliary teaching results based on the updated standardized resource library and teacher input instructions, using a cognitive big data model, includes: The system analyzes the device control commands input by the teacher to determine the target teaching equipment and control actions to be controlled. The target teaching equipment includes a smart screen, microphone, audio equipment, or lighting equipment. Obtain the device control protocol template from the updated standardized resource library, and encapsulate it into control instructions based on the STRClass protocol and the control actions. The control instructions include device turn-on, turn-off, parameter adjustment, or mode switching. The encapsulated control command is sent to the target teaching device to control the target teaching device to perform the control action; The system receives action feedback from the target teaching device and generates the auxiliary teaching results based on the control commands and the action feedback.
8. The AI-assisted teaching method using a large cognitive model according to claim 1, characterized in that, The teacher input instructions include classroom interaction instructions; Accordingly, the generation of auxiliary teaching results based on the updated standardized resource library and teacher input instructions, using a cognitive big data model, includes: Analyze the classroom interaction instructions input by the teacher to determine the type of interaction and the target knowledge points for the interaction; Extract core terms and concepts related to the target knowledge points from the updated standardized resource library; Key terms were obtained by filtering the core terms using the TF-IDF algorithm; Based on the aforementioned key terms and concepts, a cognitive big model is used to generate multi-level interactive questions, as well as standard answers and scoring criteria for these multi-level interactive questions; the multi-level interactive questions include L1 memory-based questions, L2 application-based questions, and L3 critical thinking-based questions; Based on the interaction type and the multi-level interactive questions, student response strategies are formulated, including individual answers, group discussions, or buzzer-style responses. The system receives real-time responses from students and analyzes these responses using a BERT+BiLSTM model based on the standard answer and scoring criteria. It calculates the semantic similarity between the student's real-time response and the standard answer and generates personalized feedback comments. Based on the student's real-time response, the standard answer and scoring criteria, as well as the semantic similarity and personalized feedback, the auxiliary teaching results are generated.
9. An AI-assisted teaching system that applies a large cognitive model, characterized in that, The auxiliary teaching system includes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor is used to call and run the computer program stored in the memory to perform the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Visual analysis method and system for correlation between multi-modal emotion of teacher and behavior of student
CN115641537A
Teaching resource generation method and device, equipment and storage medium
CN117975967A