Teaching method and device based on digital object, storage medium and electronic equipment
By using a multimodal large language model to perform semantic understanding and structured chain reasoning on the teaching content, the problems of insufficient understanding of teaching content and unreasonable generation process in the digital human teaching system are solved, and a logically clear and structurally complete intelligent teaching process is realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-24
AI Technical Summary
Existing digital human teaching systems lack a deep understanding and structured representation of teaching content, cannot automatically identify teaching knowledge points and their prerequisite relationships, the teaching sequence does not conform to cognitive laws, and the generation process lacks interpretability, making it difficult to meet the needs of large-scale application of intelligent education systems.
A multimodal large language model is introduced to achieve unified semantic understanding of teaching content. Combined with a structured chain reasoning mechanism, teaching knowledge points and their prerequisite dependencies are explicitly modeled to generate teaching scripts that are logically clear and structurally complete.
It has enabled the intelligent transformation of the digital human teaching process, improved the logic, systematicness and rationality of teaching, conformed to the laws of cognition, and provided a reliable technical foundation for intelligent teaching systems.
Smart Images

Figure CN121920537A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a teaching method, apparatus, storage medium, and electronic device based on digital objects. Background Technology
[0002] With the rapid development of digital human technology, multimodal large language models (MLLM), and intelligent education systems, digital human teaching is gradually being applied to various teaching scenarios such as online education, smart classrooms, and corporate training. However, in practical applications, the automatic understanding of teaching content, the reasoning ability of teaching logic, and the interpretability and controllability of the teaching process have become key technical bottlenecks restricting the further intelligentization and large-scale application of digital human teaching systems. Summary of the Invention
[0003] The purpose of this application is to provide a teaching method, apparatus, storage medium, and electronic device based on digital objects, so as to improve the broadcasting effect of digital objects.
[0004] In a first aspect, embodiments of this application provide a teaching method based on digital objects, including: Obtain the teaching content to be explained; The teaching content is semantically understood based on a multimodal large language model to determine the teaching semantic information corresponding to the teaching content; Using the teaching semantic information as reasoning constraints, chain reasoning is performed on the teaching content according to the multimodal large language model to sequentially determine each knowledge point contained in the teaching content, the dependency relationship between each knowledge point, and the arrangement order of each knowledge point. Based on each of the knowledge points and the dependencies between them, a teaching script that conforms to the stated order is generated. Control the digital objects used for teaching to broadcast the teaching script.
[0005] Secondly, embodiments of this application also provide a teaching device based on digital objects, comprising: The acquisition module is used to acquire the teaching content to be explained; The semantic processing module is used to perform semantic understanding of the teaching content based on a multimodal large language model, and determine the teaching semantic information corresponding to the teaching content; The reasoning module is used to use the teaching semantic information as reasoning constraints, perform chain reasoning on the teaching content according to the multimodal large language model, and sequentially determine each knowledge point contained in the teaching content, the dependency relationship between each knowledge point, and the arrangement order of each knowledge point. The generation module is used to generate teaching materials that conform to the arranged order based on each of the knowledge points and the dependencies between them. The control module is used to control the digital objects used for teaching to broadcast the teaching script.
[0006] Thirdly, embodiments of this application also provide a computer storage medium storing computer-executable instructions for use in any of the above-described digital object-based teaching methods.
[0007] Fourthly, embodiments of this application also provide an electronic device, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the above-described digital object-based teaching methods.
[0008] In the solution provided by the first aspect of this application, a multimodal large language model is introduced to perform unified semantic understanding of the teaching content. Combined with a structured chain reasoning mechanism, explicit modeling of teaching knowledge points and their prerequisite dependencies is performed. Based on this, a reasonable explanation order is reasoned, generating a logically clear, structurally complete digital human teaching output with a good explanation order. This transforms the digital human teaching process from simple content broadcasting into an intelligent teaching process with teaching logical reasoning capabilities. By constructing a clear teaching knowledge structure and reasoning a reasonable explanation order based on it, the logicality, systematicity, and rationality of the digital human teaching output are significantly improved. This achieves a digital human teaching generation mechanism that conforms to cognitive laws, providing a reliable technical foundation for the engineering deployment and practical application of intelligent teaching systems.
[0009] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1A flowchart of a digital object-based teaching method provided in an embodiment of this application is shown; Figure 2 A flowchart of another digital object-based teaching method provided in an embodiment of this application is shown; Figure 3 This illustration shows a schematic diagram of the instructional semantic information generation process provided in an embodiment of this application; Figure 4 This paper illustrates an overall flowchart of a digital object-based teaching method provided in an embodiment of this application. Figure 5 This illustration shows a schematic diagram of the structure of a digital object-based teaching device provided in an embodiment of this application; Figure 6 A schematic diagram of the structure of an electronic device for performing a digital object-based teaching method, provided in an embodiment of this application, is shown. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0014] With the widespread application of multimodal large language model (MLLM) and digital human technology in online education, smart classrooms, corporate training, and intelligent teaching systems, how to enable digital humans to produce teaching outputs that are clearly structured, logically sound, and in line with cognitive laws during the teaching process has become a core technical bottleneck in the implementation of current intelligent education systems.
[0015] Although existing systems can achieve voice broadcasting and visual display of teaching content, most existing digital human teaching systems rely on manually written lecture notes, template-based script generation, or simple reading of textbooks and PPT (PowerPoint presentation) content. They lack a deep understanding of the teaching content itself and the ability to model the teaching structure. They still have significant shortcomings in terms of understanding teaching content, logical reasoning in teaching, and interpretability of the teaching process. They are unable to ensure the rationality, consistency, and interpretability of teaching logic in complex teaching scenarios, and cannot meet the requirements of real teaching scenarios for systemicity, stability, and teaching quality.
[0016] Specifically, the relevant solutions mainly have the following technical problems.
[0017] (1) The generation of teaching content relies on manual labor, and its automation and generalization capabilities are insufficient: In related technologies, digital human-led instruction typically requires pre-written scripts or simple retelling of teaching materials. This approach is not only costly to produce and inefficient in updating, but also ill-suited to the demands of frequent course adjustments or large-scale course deployments. Because the system cannot automatically understand the teaching content and generate a structure that aligns with the learning objectives, digital humans often function merely as content delivery tools, lacking the teaching organization capabilities expected of intelligent teachers. This severely limits their practicality and scalability in educational settings. This statically content-driven teaching method makes the teaching process highly dependent on manual design, with low automation, and difficult to adapt to the actual needs of frequent course updates or expanded teaching scale.
[0018] (2) Lack of reasoning ability regarding teaching knowledge points and their prerequisite relationships; the order of instruction does not conform to cognitive laws: The teaching process is essentially a logical reasoning process that follows cognitive laws, and different knowledge points often have clear prerequisite dependencies. However, existing digital human teaching systems typically explain concepts according to the order of textbook chapters or pages, failing to automatically identify key knowledge points within the teaching content or infer the sequential dependencies between them. In this situation, the system is prone to problems such as a lack of necessary background information on core concepts and an unreasonable order of explanation, affecting learning outcomes and failing to meet the requirements of logic and systematic approach in real-world teaching scenarios.
[0019] (3) The teaching process lacks structured modeling, resulting in low overall interpretability and controllability: Existing digital human teaching systems mostly employ end-to-end text generation methods, whose internal generation logic often remains a black box, making it difficult to explain the basis for the generated content. The system cannot clearly explain the source of the knowledge points corresponding to a particular topic, nor can it trace the reasons for the order of explanations, thus hindering teaching quality assessment, teaching strategy adjustment, and subsequent personalized teaching expansion. In application scenarios involving teaching standards, quality supervision, or integration with intelligent teaching systems, this lack of interpretability presents significant limitations. The lack of interpretability in generation methods makes teaching quality assessment, teaching strategy optimization, and subsequent personalized teaching expansion extremely difficult, posing a clear limitation in application scenarios requiring standardized teaching management or integration with intelligent teaching systems.
[0020] Overall, existing digital human teaching technologies generally suffer from the following common problems: First, they lack a deep understanding and structured representation of the teaching content, relying on manual processes or simple generation; second, they lack the reasoning ability for teaching knowledge points and their prerequisite relationships, making it difficult to align the explanation sequence with cognitive patterns; and third, they lack interpretable teaching modeling mechanisms, making it difficult to analyze, evaluate, and control the teaching process. These problems not only limit the teaching effectiveness of digital human teaching systems but also significantly restrict their large-scale application and long-term development in the field of intelligent education.
[0021] To address some or all of the aforementioned problems, this application provides a teaching method based on digital objects. By introducing a multimodal large language model to achieve unified semantic understanding of the teaching content, and combining it with a structured reasoning mechanism, the teaching knowledge points and their prerequisite dependencies are explicitly modeled. Based on this, a reasonable explanation order is reasoned to generate a digital human teaching output that is logically clear, structurally complete, and has a good explanation order. This transforms the digital human teaching process from simple content broadcasting into an intelligent teaching process with teaching logic reasoning capabilities.
[0022] This application provides a teaching method based on digital objects. This method can be applied to terminal devices, such as teaching equipment; or it can be applied to a digital teaching system within the device. See also... Figure 1 As shown, the method includes steps 101 to 105.
[0023] Step 101: Obtain the teaching content to be explained.
[0024] Before using digital objects (such as digital humans) for digital teaching, the teaching content to be taught must first be acquired and used as input for subsequent processing. This teaching content can specifically include one or more of the following: textbook text, teaching slides (PowerPoint presentations), or teaching videos.
[0025] Step 102: Perform semantic understanding of the teaching content based on the multimodal large language model to determine the teaching semantic information corresponding to the teaching content.
[0026] In this embodiment, the original teaching content includes multimodal data, which may include text-formatted content (e.g., teaching text, teaching courseware), audio-formatted content (e.g., audio content in teaching recordings, audio content in teaching videos), and image-formatted content (e.g., images in teaching courseware, frames in teaching videos), etc. This embodiment utilizes a Multimodal Large Language Model (MLLM) to perform semantic understanding of the teaching content, enabling accurate semantic understanding of the multimodal teaching content and obtaining more accurate and comprehensive teaching semantic information.
[0027] The teaching semantic information describes what is taught, what the objectives are, and which knowledge areas are involved, providing a clear semantic foundation for subsequent logical reasoning in teaching. Specifically, this teaching semantic information may include the teaching topic, teaching objectives, and the scope of knowledge covered.
[0028] The teaching theme can be summarized in one sentence, using standard terminology from the textbook as much as possible. Teaching objectives should be written as action sentences that can be taught, such as "understand… / master… / be able to derive… / be able to apply…", etc. There can be multiple teaching objectives, which can be presented in a list format. The scope of knowledge covered can be represented by labels, such as "calculus - derivatives / limits / tangents and secants", etc. Generally, multiple labels need to be identified, for example, no fewer than five.
[0029] The multimodal large language model can adopt a pre-trained model that supports joint modeling of text, image and speech, such as a multimodal large language model based on the Transformer architecture. This embodiment does not limit this.
[0030] Step 103: Using teaching semantic information as reasoning constraints, perform chain reasoning on the teaching content based on the multimodal large language model to determine the knowledge points included in the teaching content, the dependencies between the knowledge points, and the order of the knowledge points.
[0031] In this embodiment, the teaching semantic information generated by the multimodal large language model, such as the teaching topic, teaching objectives, and knowledge coverage, can be used as reasoning constraints or reasoning boundaries for subsequent processing. Specifically, chain-of-thought can be performed on the teaching content based on the multimodal large language model to generate structured information related to knowledge points. This structured information specifically includes the various knowledge points contained in the teaching content, the dependencies between the various knowledge points, and the order in which the various knowledge points are arranged.
[0032] This requires defining the processing order of the multimodal large language model: first, determining the knowledge points included in the teaching content; then, determining the dependencies between these knowledge points; and finally, determining the order in which they are arranged. This chain-like reasoning approach allows for the decomposition of teaching reasoning tasks, enabling the multimodal large language model to output reasoning results step by step according to a fixed reasoning structure, thereby achieving a controllable and interpretable teaching logical reasoning process.
[0033] In this embodiment, a knowledge point refers to the smallest unit of knowledge extracted from the teaching content that has independent teaching significance. It is used to represent the core concepts, definitions, principles, or skill points in the course. In this embodiment, the knowledge point can serve as the basic unit for teaching logical reasoning and knowledge structure modeling, and is the core object for generating teaching materials.
[0034] Teaching sequence reasoning refers to the reasoning process of arranging and determining the logical order of explanation of teaching knowledge points based on identifying their prerequisites and dependencies. This order is not a fixed template and can be dynamically generated by a multimodal large language model based on the specific teaching content.
[0035] Step 104: Generate a teaching script that conforms to the order of the various knowledge points and their dependencies.
[0036] In this embodiment, after determining each knowledge point and the dependencies between them, the teaching content can be combined to generate the corresponding knowledge point lecture scripts. The lecture scripts are then sorted according to the order of the knowledge points to generate teaching lecture scripts that can explain each knowledge point in sequence.
[0037] The teaching materials can be generated based on a large language model. This large language model can be the same as or different from the multimodal large language model mentioned above. For example, the large language model can be a dedicated text generation model (text-only LLM) to improve generation efficiency and stability.
[0038] Specifically, each knowledge point, the dependencies between knowledge points, and the order of knowledge points can be represented in a structured form such as a list, adjacency list, or matrix. This structured representation can be input into the large language model to instruct the large language model to generate teaching materials.
[0039] When generating teaching materials, it is necessary to explicitly tell the large language model that the materials must be generated in this order. For example, the prompt words should contain content related to the order constraint. Furthermore, the content summary and evidence fragments corresponding to each knowledge point should be used as the basis for generation. Based on this, the teaching materials for each knowledge point can be generated, which may include explanatory text, examples, and related images.
[0040] Step 105: Control the digital object used for teaching to broadcast the teaching script.
[0041] After generating the teaching script, the corresponding digital object can be controlled to broadcast the script, disseminating the content corresponding to each knowledge point to students and other audiences. This digital object is a virtual object used for teaching; for example, it could be a digital human or a digital animal, but this embodiment does not limit its use.
[0042] The digital object-based teaching method provided in this application introduces a multimodal large language model to achieve unified semantic understanding of teaching content. Combined with a structured chain reasoning mechanism, it explicitly models teaching knowledge points and their pre-dependent relationships, and on this basis, infers a reasonable explanation sequence. This generates logically clear, structurally complete digital human teaching output with a well-ordered explanation, transforming the digital human teaching process from simple content recitation into an intelligent teaching process with teaching logical reasoning capabilities. By constructing a clear teaching knowledge structure and reasoning a reasonable explanation sequence based on it, the logicality, systematicity, and rationality of the digital object teaching output are significantly improved. This achieves a digital human teaching generation mechanism that conforms to cognitive laws, providing a reliable technical foundation for the engineering deployment and practical application of intelligent teaching systems.
[0043] This application provides another teaching method based on digital objects, which can be applied to terminal devices, such as teaching equipment; or, it can also be applied to a digital teaching system within the device. See also Figure 2 As shown, the method includes steps 201 to 206.
[0044] Step 201: Obtain the teaching content to be explained.
[0045] For details, please refer to Figure 1 The descriptions related to step 101 in the illustrated embodiments are not repeated here.
[0046] Step 202: Perform semantic understanding of the teaching content based on the multimodal large language model to determine the teaching semantic information corresponding to the teaching content.
[0047] For details, please refer to Figure 1 The descriptions related to step 102 in the illustrated embodiments are not repeated here.
[0048] In some optional implementations, step 202 above, "performing semantic understanding of the teaching content based on the multimodal large language model and determining the teaching semantic information corresponding to the teaching content", may include steps A1 to A5.
[0049] Step A1: Perform semantic understanding on the text content in the teaching materials to generate structured text features with a hierarchical structure.
[0050] Step A2: Perform speech recognition on the audio content in the teaching materials to determine the speech-text features.
[0051] Step A3: Perform semantic understanding on the image content in the teaching materials to determine the corresponding visual semantic features; and extract key images from the image content.
[0052] Multimodal teaching content includes various formats, such as text, audio, and images. In this embodiment, the multimodal teaching content is pre-processed and converted into structured features (i.e., multimodal teaching features) that are easy for multimodal large language models to recognize, thus facilitating semantic understanding of the teaching content by the multimodal large language model.
[0053] Figure 3 This diagram illustrates a process for generating instructional semantic information. For example... Figure 3 As shown, for textual content such as textbooks and teaching materials in teaching content, the system (e.g., a digital teaching system) uses a pre-trained language model to parse the content and extract the structure. This pre-trained language model can be a Transformer-based model, such as BERT, RoBERTa, ERNIE, or a large language model with fine-tuned instructions based on it. Through the ability of this type of pre-trained language model to understand the semantic context of the text, the system can identify chapter titles, section titles, and their hierarchical relationships in the text content (e.g., textbooks or teaching materials), and divide the main text into paragraphs and extract key teaching sentences, thereby forming a text representation with a clear structural hierarchy, i.e., structured text features.
[0054] For audio content in teaching materials, such as original teaching recordings or audio tracks extracted from teaching videos, an automatic speech recognition model can be used to transcribe the audio content into text, achieving speech recognition of the audio content. This speech recognition model can include an end-to-end deep neural network-based speech recognition model used to convert continuous speech signals into corresponding text. After speech recognition, the corresponding speech-text features can be extracted.
[0055] For image content in teaching materials, such as images in teaching courseware or frames in teaching videos, image information extraction models can be used to perform semantic understanding on the image content, such as video frames, thereby extracting visual semantic features that can represent visual characteristics. This image information extraction model can include optical character recognition models for text extraction, and image understanding models or large language models for image semantic understanding, to extract teaching-related textual and visual semantic information from the image content, forming visual semantic features.
[0056] Simultaneously, it can extract key frames from image content, such as performing keyframe extraction on video footage to obtain a set of keyframes associated with timestamps / frame numbers. Furthermore, as... Figure 3 As shown, a visual encoding model can be further used to encode the explicit key images into visual embeddings, thereby obtaining the corresponding key image encoding features, which can then be used as the visual input of a multimodal large language model.
[0057] Step A4: Generate multimodal teaching features based on at least one of structured text features, speech text features, visual semantic features, and key images.
[0058] After generating structured text features, speech text features, visual semantic features, and key images (such as key image encoding features), these features can be fused, for example, by splicing them together, to generate structured features that can be subsequently input into a large language model, i.e., multimodal teaching features.
[0059] Through the above processing, the system uniformly converts teaching content from different modalities such as text, speech, and images into an input representation that can be processed by a multimodal large language model, providing basic data for subsequent understanding of teaching content and logical reasoning in teaching.
[0060] Step A5: Input the multimodal teaching features into the multimodal large language model to obtain teaching semantic information, which includes teaching topics, teaching objectives, and knowledge coverage.
[0061] In this embodiment, multimodal teaching features are input into a multimodal large language model, which performs unified semantic understanding processing on the teaching content to automatically identify the teaching topics, teaching objectives, and knowledge coverage involved in the current teaching content, forming teaching semantic information. This teaching semantic information can eliminate modal differences, making the teaching content consistent at the semantic level and providing a standard input for teaching reasoning and structural modeling.
[0062] Specifically, the system can generate corresponding prompts using preset structured prompt templates to limit the output format of the multimodal large language model, ensuring that its output includes at least a structured result containing the teaching topic, a set of teaching objectives, and a set of knowledge coverage. Optionally, the output structured teaching semantic information can further include evidence source identifiers corresponding to each field, such as chapter numbers, page numbers, frame numbers, timestamps, or text fragment positions corresponding to the teaching objectives, to achieve traceable semantic recognition.
[0063] Optionally, the various features obtained from preprocessing (structured text features, speech text features, visual semantic features, key images, etc.) can be aligned and fused. The fusion can be guided by prompts such as words. Alternatively, multiple features can be aligned and fused first to obtain aligned and fused multimodal teaching features. Afterward, the multimodal language model can focus primarily on semantic understanding tasks, minimizing the problem of redundant or erroneous content in the teaching semantic information output by the multimodal language model due to high task complexity.
[0064] Specifically, step A4 above, "generating multimodal teaching features based on at least one of structured text features, speech text features, visual semantic features, and key images," may include steps A41 to A42.
[0065] Step A41: Align structured text features, speech text features, visual semantic features, and key images.
[0066] In this embodiment, appropriate alignment methods can be adopted for different features. Specifically, alignment can be based on timestamps, structural hierarchy, or semantic similarity.
[0067] Specifically, timestamp-based alignment involves labeling the start and end times of each text segment in the speech-to-text features obtained from speech transcription, and labeling the corresponding timestamps for visual semantic features and key images. This allows for the establishment of associations between speech-to-text features, visual semantic features, and key images according to time windows, thereby achieving alignment.
[0068] Alignment based on structural hierarchy: Based on chapter titles, section titles, page numbers, and page layout, a mapping is established between text paragraphs in the structured text features and visual semantic features determined based on courseware pages or video images to achieve alignment.
[0069] Semantic similarity-based alignment: Text segments in structured text features and speech text features (and may also include text content in visual semantic features) are vectorized and encoded respectively, similarity is calculated, and matching is performed based on similarity to establish cross-modal association.
[0070] Step A42: Perform feature fusion based on the aligned structured text features, speech text features, visual semantic features, and key images to obtain multimodal teaching features that integrate multiple features.
[0071] Based on the alignment results, further fusion can be performed.
[0072] Specifically, the aligned features can be concatenated to obtain concatenated multimodal teaching features. Alternatively, the aligned features can be projected onto a unified feature space, followed by weighted or gated fusion to obtain a fused semantic vector, which can then be used as the multimodal teaching feature. Alternatively, for each aligned content, a cross-modal semantic summary can be generated, and all summaries can be globally aggregated to form a unified semantic representation for instructional reasoning.
[0073] By leveraging the aligned and fused multimodal teaching features, the multimodal large language model can perform high-level semantic understanding of teaching content. It can more accurately identify the teaching topics, teaching objectives, and knowledge coverage involved in the teaching content, and output the corresponding semantic representation results, providing a clear semantic foundation for subsequent teaching logical reasoning.
[0074] Step 203: Using teaching semantic information as reasoning constraints, perform chain reasoning on the teaching content based on the multimodal large language model to determine the knowledge points included in the teaching content, the dependencies between the knowledge points, and the order of the knowledge points.
[0075] For details, please refer to Figure 1 The descriptions related to step 103 in the illustrated embodiments are not repeated here.
[0076] Optionally, step 203, "using teaching semantic information as reasoning constraints, performing chain reasoning on the teaching content based on the multimodal large language model, and sequentially determining the knowledge points included in the teaching content, the dependencies between the knowledge points, and the order of the knowledge points," may include steps B1 to B3.
[0077] Step B1: Using teaching semantic information as reasoning constraints, a multimodal large language model is used to identify knowledge points in the teaching content and determine the various knowledge points contained in the teaching content.
[0078] Step B2: Based on the identified knowledge points, use a multimodal large language model to infer dependencies and determine the dependencies between the knowledge points.
[0079] Step B3: Based on the identified knowledge points and the dependencies between them, use a multimodal large language model to sort the knowledge points and determine their order.
[0080] In this embodiment, after determining the teaching semantic information, it is used as a high-level semantic constraint for chain reasoning. This allows for further structured chain-of-thought teaching reasoning based on a multimodal large language model, enabling explicit teaching logic analysis of the teaching content.
[0081] Specifically, the first step is to guide a multimodal large language model to perform a task of identifying teaching knowledge points. Based on the obtained teaching semantic information, the multimodal large language model performs fine-grained semantic analysis on the teaching content, breaking it down into several knowledge points with clear teaching significance, forming a set of knowledge points. Each knowledge point can be output in a structured text format, including information such as the knowledge point name, knowledge point type, and corresponding teaching content segment.
[0082] Subsequently, the multimodal large language model is further guided to perform a prerequisite dependency inference task on the identified set of knowledge points. This inference process specifically analyzes the dependencies between various knowledge points in terms of concept definition, formula derivation, logical relationships, or cognitive order to determine which knowledge points need to be explained before others, i.e., identifying which knowledge points are prerequisites for others. The logical inference, performed using the multimodal large language model's knowledge association and reasoning capabilities, outputs the dependencies between knowledge points, which can be specifically represented as a structured relationship of "prerequisite knowledge point—subsequent knowledge point."
[0083] After inferring the prerequisite dependencies, the multimodal language model is further guided to integrate the importance of each knowledge point and its prerequisite dependencies, and to rank the knowledge points to generate an order that conforms to the cognitive patterns of teaching. This order can then be used as the order of explanation for subsequent lessons. The importance of each knowledge point can be scored by the multimodal language model under the constraints of a prompt template, and the score can be output in a structured manner (e.g., 0-1 or 1-5, which can be output synchronously in step B1). The scoring criteria can also be output, such as the evidence fragments, page numbers, frame numbers, and timestamps corresponding to the knowledge points.
[0084] In this embodiment, the teaching content analysis is broken down into multiple sub-tasks such as knowledge point identification, prerequisite dependency inference, and order generation, instead of being generated directly end-to-end. This not only ensures the accuracy and reliability of the execution results of each sub-task, but also facilitates verification and error correction, such as performing acyclicity or coverage checks, resulting in higher stability.
[0085] It should be noted that when performing structured chain-of-thought instructional reasoning, the content input to the multimodal large language model is not limited to instructional semantic information such as instructional topics, instructional objectives, and knowledge coverage. It can also include evidence-based instructional content (such as original instructional content or multimodal instructional features generated in step A4 above), and chain-like reasoning is achieved through inference template constraints. That is, the content input to the multimodal large language model at this time can specifically include: instructional semantic information + evidence-based instructional content + inference template constraints.
[0086] Specifically, during instructional reasoning, the multimodal large language model is instructed to use the teaching topic, list of teaching objectives, and knowledge coverage as reasoning boundaries and task constraints to limit the scope of knowledge to be covered in this reasoning and avoid off-topic output. Simultaneously, the teaching content serves as reasoning evidence, enabling the model to identify knowledge points, infer prerequisite dependencies, and generate the order of explanation based on traceable evidence. This teaching content can specifically include chapter-level textbooks, courseware text blocks and key sentences, video-to-speech transcription clips, keyframes or courseware page images, and alignment indices between the aforementioned text fragments and keyframes.
[0087] In addition, predefined teaching reasoning prompt templates or structured outputs will be injected into the multimodal large language model as reasoning template constraints, so that the multimodal large language model can output intermediate reasoning results step by step according to fixed fields (including knowledge point sets, dependency edges and sorting sequences, and can provide evidence sources and optional confidence levels), thereby realizing a controllable, interpretable and convenient teaching logic reasoning process for subsequent lecture note generation.
[0088] For example, this inference template constraint can require the model to output in three structured steps: ① knowledge point identification (output concepts and provide evidence); ② dependency inference (output dependencies and provide reason / evidence); ③ order generation (output order and provide transition logic), with the output uniformly in JSON format (optionally with confidence).
[0089] Optionally, consistency checks can be performed on any of the above inference results, such as acyclicity, coverage, and conflict detection, and local re-inference or regeneration can be triggered when the check fails to avoid large model output drift.
[0090] In this embodiment, by injecting predefined teaching reasoning prompt templates into the multimodal large language model, the model outputs intermediate reasoning results step by step according to a fixed reasoning structure, thereby achieving a controllable and interpretable teaching logical reasoning process. Furthermore, each knowledge point, along with its related dependencies and order, can be output in a structured form, such as a hierarchical or sequential structure, providing clear and directly usable input for subsequent teaching knowledge structure modeling and digital human teaching script generation.
[0091] Step 204: Treat each knowledge point as a node, establish directed edges between the corresponding nodes based on the dependencies between the knowledge points, and generate a teaching knowledge graph.
[0092] In this embodiment, based on the above-mentioned structured teaching reasoning results, a teaching knowledge graph related to knowledge points can be constructed to explicitly represent the overall knowledge structure and explanation logic of the teaching content.
[0093] Specifically, the teaching knowledge point identification results obtained from chain reasoning and the dependencies between knowledge points are used as input data. These data are standardized, mapping each knowledge point to a node in a knowledge graph. Each node is assigned corresponding attribute information, which may include the knowledge point name, category, corresponding teaching content summary, and importance level. This attribute information is simultaneously output during chain reasoning by the multimodal large language model. Furthermore, based on the derived dependencies, directed edges can be established between knowledge point nodes to represent the sequential explanation or dependency relationships between knowledge points, with the direction of the edges indicating the teaching order constraint.
[0094] The teaching knowledge graph can be stored and represented using graph data structures, such as graph databases or in-memory graph structures. A graph construction algorithm transforms the "knowledge point-dependency relationship" into a directed acyclic graph structure to ensure that there are no logical conflicts in the teaching sequence.
[0095] In this embodiment, the teaching content is transformed from raw text or multimodal information into a well-structured and hierarchical knowledge graph, making the teaching knowledge structure explicitly interpretable and providing a unified structured basis for subsequent lecture note generation and teaching process control. Furthermore, the teaching knowledge graph can explicitly store node attributes and dependencies, facilitating consistency verification, evidence tracing, and subsequent expansion (such as personalized rearrangement, adding or deleting nodes, and tiered instruction).
[0096] Step 205: Generate teaching materials that conform to the order of the teaching knowledge graph.
[0097] In this embodiment, the teaching knowledge graph can represent each knowledge point and the dependencies between them. Therefore, teaching materials that conform to the arrangement order can be generated based on the teaching knowledge graph. Specifically, step 104 or step 205 above may include the following steps C1 to C3.
[0098] Step C1: Identify the core knowledge points within each knowledge point.
[0099] Step C2: Based on the dependencies between various knowledge points, determine the prerequisite knowledge points that the core knowledge points depend on, and determine the subsequent knowledge points that depend on the core knowledge points.
[0100] Step C3: Generate the lecture notes for each knowledge point in the order of prerequisite knowledge points, core knowledge points, and subsequent knowledge points.
[0101] In this embodiment, it is possible to identify which of the knowledge points are the most important core knowledge points. For example, core knowledge points can be selected when performing chain reasoning in a multimodal large language model or after constructing a teaching knowledge graph.
[0102] For example, the system can first use the teaching objectives / knowledge scope output in step 2 as constraints to filter knowledge points strongly related to the teaching objectives. Then, it can assign weights based on the structural information of the teaching content (such as chapter titles, key points, and whether the knowledge point is located in the main chapter rather than the extension section). Furthermore, it can rely on the importance of relevant nodes in the teaching knowledge graph (the importance of a node can be determined based on its out-degree / in-degree, betweenness centrality, and the number of subsequent knowledge points it covers), and combine this with the frequency and emphasis of the knowledge point in the teaching content to form an importance score. Finally, based on the importance scores of each knowledge point, one or more core knowledge points that are relatively important are selected. It is understood that this score can also be output by a large model under the constraints of a prompt template; this embodiment does not limit this.
[0103] Furthermore, before explaining a core knowledge point, it may be necessary to first explain other knowledge points, i.e., the prerequisite knowledge points upon which the core knowledge point depends; similarly, after explaining the core knowledge point, other knowledge points that depend on it can be further explained, i.e., the subsequent knowledge points. When generating teaching materials, the teaching materials can be generated sequentially according to the order of prerequisite knowledge points, core knowledge points, and then subsequent knowledge points, ultimately generating a teaching material containing the teaching materials for each knowledge point. This allows subsequent digital objects to be explained sequentially in the order of prerequisite knowledge points → core knowledge points → subsequent knowledge points.
[0104] Taking a teaching knowledge graph as an example, after the teaching knowledge point graph is constructed, a step-by-step teaching script can be automatically generated based on the teaching knowledge graph. This teaching script can be generated by a large language model under the constraints of the knowledge graph to ensure that the content strictly follows the established teaching structure and order.
[0105] Specifically, using the teaching knowledge graph as a structured constraint input, the knowledge point nodes and their directed dependencies in the teaching knowledge graph are transformed into explanation order instructions corresponding to each knowledge point. These explanation order instructions and the teaching content of the corresponding knowledge points are then input into a large language model (such as a generative language model based on the Transformer architecture) to generate the knowledge point lecture notes corresponding to each knowledge point in sequence.
[0106] The large language model first locates the starting node (corresponding to the starting knowledge point) in the teaching knowledge graph. This starting node is the entry or root node of the lecture flow, used to generate course introduction content, such as the topic, objectives, scope, and prerequisite overview. Teaching introduction content is then generated for this starting node to explain the course topic and learning objectives.
[0107] Furthermore, starting from the core knowledge points, the system sequentially generates explanations of prerequisite knowledge points, core knowledge points, and subsequent knowledge points according to the dependencies in the knowledge graph. Finally, a teaching summary is generated for the terminal node of the entire teaching knowledge graph. The generated step-by-step teaching materials are output in a structured format, such as organized by time sequence or step numbering, ultimately providing directly executable teaching materials for digital human broadcast control in the order of "Import (starting node) → Prerequisite / Prerequisite Knowledge Points → Core Knowledge Points → Subsequent Knowledge Points → Summary (terminating node)".
[0108] Optionally, the method may further include: generating transitional explanatory text between adjacent knowledge points; and adding the transitional explanatory text to the teaching materials at the position corresponding to the adjacent knowledge points.
[0109] In this embodiment, for any two adjacent knowledge points, transitional explanatory text can be inserted between them. This transitional explanatory text can be automatically generated based on the generated structured explanation order and the information of adjacent knowledge points.
[0110] For example, the system can input the attribute information (name, type, summary, etc.) of two adjacent knowledge points, their dependencies, and corresponding evidence fragments into a large language model. It then requires the output of a transitional sentence in the prompt template, such as reviewing the previous knowledge point, pointing out the connection, and finally introducing the next knowledge point, thus forming a transitional explanatory text. This transitional explanatory text connects and explains adjacent knowledge points; its content is determined by the knowledge point summary and dependency constraints, and the model is responsible for language organization and logical connection. Furthermore, after generation, it can verify whether the next knowledge point is correctly referenced and whether new concepts are introduced out of bounds, thus achieving text verification.
[0111] Similarly, consistency checks can be performed on the generated lecture content for each knowledge point. For example, checks can be performed to see if any knowledge points are missing, if the order is out of order, or if any knowledge points outside the scope are introduced. If any of these checks fail, a partial rewrite or regeneration is triggered to ensure that the output lectures strictly follow the established teaching structure and explanation order.
[0112] Step 206: Control the digital object used for teaching to broadcast the teaching script.
[0113] For details, please refer to Figure 1 The description related to step 105 in the illustrated embodiment will not be repeated here.
[0114] Optionally, step 206, "Controlling the digital object used for teaching to broadcast the teaching script," may include steps D1 to D2.
[0115] Step D1: Convert the text content in the teaching script into a speech signal and control the digital object to broadcast the speech signal.
[0116] Step D1: At the switching position between various knowledge points, trigger the control command of the digital object. The control command is used to control at least one of the digital object's actions, expressions, and camera angles.
[0117] In this embodiment, the text content of the teaching script can be input into the speech synthesis module, converting the text into a natural and fluent speech signal. Simultaneously, based on the structural information in the teaching script, the actions, facial expressions, and teaching flow transitions of digital objects can be synchronously controlled. Specifically, when switching knowledge points, corresponding control signals are triggered to control the actions, expressions, or camera angles of the digital object. For example, the explanation posture or emphasis of the digital object can be adjusted, forming a closed loop of content comprehension and reasoning → structured script → broadcast control execution, thereby ensuring that the teaching process is controllable, explainable, and traceable.
[0118] Figure 4 An overall flowchart of this digital object-based teaching method is shown. Figure 4 As shown, the teaching content is first obtained, and the teaching content is understood according to MLLM. Then, structured CoT reasoning is performed, which includes identifying knowledge points, inferring dependencies, and determining the order of arrangement (explanation order). Then, a teaching knowledge graph of knowledge points is constructed, and teaching scripts for digital human to teach step by step are generated based on the knowledge graph, which can then drive digital human to teach output.
[0119] Compared to related solutions that rely on manually written lecture notes, templated script generation, and sequential reading of textbooks or PowerPoint presentations, the digital object-based teaching method provided in this embodiment first utilizes a multimodal large language model to perform unified semantic understanding of multi-source teaching content, including textbook texts, teaching courseware, and teaching videos. This automatically parses the teaching theme, teaching objectives, and teaching scope, providing a reliable semantic foundation for subsequent teaching reasoning. Through this mechanism, the system can break free from dependence on manual lecture notes and achieve automated understanding and processing of teaching content.
[0120] Based on an understanding of the teaching content, a structured chain-of-thought reasoning mechanism is introduced to explicitly model the teaching process. Through reasoning steps such as knowledge point identification, inference of prior relationships, and determination of the explanation order, a logical explanation that conforms to the cognitive laws of teaching can be gradually constructed. This makes the teaching process form a clear and traceable logical structure, and the teaching output is no longer black-box text, but a result generated based on a clear teaching reasoning path. The teaching process has stability and consistency, thus significantly improving the interpretability and rationality of digital human teaching and effectively avoiding the problems of arbitrary explanation order and logical confusion in traditional systems.
[0121] Furthermore, through a structured reasoning mechanism, the system can dynamically generate a delivery sequence that conforms to cognitive principles based on the teaching content itself, without relying on manually set fixed templates or rules. This mechanism enables the digital human teaching system to adapt to different course content, teaching difficulty, and teaching scenarios, avoiding the need to rewrite lecture notes whenever the content changes, and significantly improving the automation level and engineering usability of the teaching system.
[0122] Based on the results of structured reasoning, a knowledge graph of teaching knowledge points is further constructed to explicitly describe the overall structure of the teaching content. This knowledge graph provides an interpretable structural basis for the generation of digital human teaching scripts, making the teaching process traceable and analyzable, and supporting subsequent personalized and adaptive teaching extensions. Furthermore, the generation of digital human teaching scripts strictly follows the knowledge point nodes and their prerequisite relationships, ensuring the consistency and integrity of the teaching content in its overall structure, effectively avoiding disordered generation and logical deviations. This knowledge structure-constrained generation method enables the digital human teaching process to maintain stable output even in complex teaching scenarios, significantly improving the controllability and consistency of teaching quality and compensating for the shortcomings of traditional generative methods in teaching scenarios.
[0123] This embodiment deeply integrates multimodal teaching content understanding, structured chain-of-thought reasoning, and teaching knowledge structure modeling, achieving a comprehensive upgrade of digital human teaching technology across multiple levels, including teaching logic, content generation methods, and system interpretability. The structured teaching reasoning mechanism ensures logical clarity and interpretability in the teaching process; the automated teaching structure construction mechanism significantly reduces manual costs and enhances system adaptability; and the knowledge structure-constrained generation method guarantees the stability and quality control of teaching output. Without altering the digital human ontology structure or language model structure, it can enhance capabilities on the existing system, significantly improving the intelligence level and engineering implementation capabilities of the digital human teaching system. It possesses good engineering compatibility and has broad educational application prospects and significant industrial value.
[0124] The above describes in detail the teaching method based on digital objects. This method can also be implemented using a corresponding device, the structure and function of which will be described in detail below.
[0125] Based on the same inventive concept, this application also provides a teaching device based on digital objects, see [link to relevant documentation]. Figure 5 As shown, the device includes: Module 501 is used to acquire the teaching content to be explained; The semantic processing module 502 is used to perform semantic understanding on the teaching content based on the multimodal large language model, and determine the teaching semantic information corresponding to the teaching content; The reasoning module 503 is used to use the teaching semantic information as reasoning constraints, perform chain reasoning on the teaching content according to the multimodal large language model, and sequentially determine each knowledge point contained in the teaching content, the dependency relationship between each knowledge point, and the arrangement order of each knowledge point. The generation module 504 is used to generate teaching materials that conform to the arrangement order based on each of the knowledge points and the dependencies between the knowledge points. The control module 505 is used to control the digital objects used for teaching to broadcast the teaching script.
[0126] In some optional implementations, the step of performing semantic understanding of the teaching content based on a multimodal large language model to determine the teaching semantic information corresponding to the teaching content includes: Semantic understanding is performed on the text content in the teaching materials to generate structured text features with a hierarchical structure. Speech recognition is performed on the audio content in the teaching materials to determine the speech-text features; The image content in the teaching content is semantically understood to determine the corresponding visual semantic features; and key images are extracted from the image content. Multimodal teaching features are generated based on at least one of the structured text features, the speech text features, the visual semantic features, and the key images; Multimodal teaching features are input into a multimodal large language model to obtain teaching semantic information, which includes teaching topics, teaching objectives, and knowledge coverage.
[0127] In some optional implementations, generating multimodal teaching features based on at least one of the structured text features, the speech text features, the visual semantic features, and the key images includes: Align the structured text features, the speech text features, the visual semantic features, and the key images; Based on the aligned structured text features, speech text features, visual semantic features, and key images, feature fusion is performed to obtain multimodal teaching features that integrate multiple features.
[0128] In some optional implementations, the step of using the teaching semantic information as reasoning constraints, performing chain-like reasoning on the teaching content based on a multimodal large language model, and sequentially determining each knowledge point included in the teaching content, the dependencies between each knowledge point, and the order of the knowledge points includes: Using the teaching semantic information as reasoning constraints, a multimodal large language model is used to identify knowledge points in the teaching content, thereby determining the various knowledge points contained in the teaching content. Based on the identified knowledge points, a multimodal large language model is used to infer dependencies and determine the dependencies between the knowledge points. Based on the identified knowledge points and the dependencies between them, a multimodal large language model is used to sort the knowledge points and determine their order.
[0129] In some optional implementations, generating teaching materials that conform to the stated order based on each of the knowledge points and the dependencies between them includes: Each knowledge point is treated as a node, and directed edges are established between the corresponding nodes based on the dependencies between the knowledge points to generate a teaching knowledge graph. Based on the teaching knowledge graph, generate teaching materials that conform to the stated order.
[0130] In some optional implementations, generating teaching materials that conform to the stated order based on each of the knowledge points and the dependencies between them includes: Identify the core knowledge points among each of the aforementioned knowledge points; Based on the dependencies between the various knowledge points, determine the prerequisite knowledge points that the core knowledge point depends on, and determine the subsequent knowledge points that depend on the core knowledge point. The lecture notes for each knowledge point are generated sequentially according to the order of prerequisite knowledge points, core knowledge points, and subsequent knowledge points.
[0131] In some alternative implementations, controlling the digital object used for teaching to broadcast the teaching materials includes: Convert the text content in the teaching script into a speech signal, and control the digital object to broadcast the speech signal; At the switching position between the various knowledge points, a control command for the digital object is triggered, which is used to control at least one of the digital object's actions, expressions, and camera angles.
[0132] This application also provides a computer storage medium storing computer-executable instructions, which include a program for executing the above-described teaching method based on digital objects. The computer-executable instructions can execute the methods in any of the above-described method embodiments.
[0133] The computer storage medium can be any available medium or data storage device that a computer can access, including but not limited to magnetic storage (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical storage (e.g., CD, DVD, BD, HVD), and semiconductor storage (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).
[0134] Figure 6 A structural block diagram of an electronic device according to another embodiment of this application is shown. The electronic device 1100 may be a host server with computing capabilities, a personal computer (PC), or a portable computer or terminal, etc. The specific embodiments of this application do not limit the specific implementation of the electronic device.
[0135] The electronic device 1100 includes at least one processor 1110, a communications interface 1120, a memory array 1130, and a bus 1140. The processor 1110, the communications interface 1120, and the memory 1130 communicate with each other via the bus 1140.
[0136] The communication interface 1120 is used to communicate with network elements, including, for example, virtual machine management centers and shared storage.
[0137] Processor 1110 is used to execute programs. Processor 1110 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0138] Memory 1130 is used for executable instructions. Memory 1130 may include high-speed RAM and may also include non-volatile memory, such as at least one disk storage device. Memory 1130 may also be a memory array. Memory 1130 may also be divided into blocks, and the blocks may be combined into virtual volumes according to certain rules. The instructions stored in memory 1130 can be executed by processor 1110 to enable processor 1110 to perform the digital object-based teaching method in any of the above method embodiments.
[0139] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A teaching method based on digital objects, characterized in that, include: Obtain the teaching content to be explained; The teaching content is semantically understood based on a multimodal large language model to determine the teaching semantic information corresponding to the teaching content; Using the teaching semantic information as reasoning constraints, chain reasoning is performed on the teaching content according to the multimodal large language model to sequentially determine each knowledge point contained in the teaching content, the dependency relationship between each knowledge point, and the arrangement order of each knowledge point. Based on each of the knowledge points and the dependencies between them, a teaching script that conforms to the stated order is generated. Control the digital objects used for teaching to broadcast the teaching script.
2. The method according to claim 1, characterized in that, The step of performing semantic understanding of the teaching content based on a multimodal large language model to determine the corresponding teaching semantic information includes: Semantic understanding is performed on the text content in the teaching materials to generate structured text features with a hierarchical structure. Speech recognition is performed on the audio content in the teaching materials to determine the speech-text features; The image content in the teaching content is semantically understood to determine the corresponding visual semantic features; and key images are extracted from the image content. Multimodal teaching features are generated based on at least one of the structured text features, the speech text features, the visual semantic features, and the key images; Multimodal teaching features are input into a multimodal large language model to obtain teaching semantic information, which includes teaching topics, teaching objectives, and knowledge coverage.
3. The method according to claim 2, characterized in that, The step of generating multimodal teaching features based on at least one of the structured text features, the speech text features, the visual semantic features, and the key images includes: Align the structured text features, the speech text features, the visual semantic features, and the key images; Based on the aligned structured text features, speech text features, visual semantic features, and key images, feature fusion is performed to obtain multimodal teaching features that integrate multiple features.
4. The method according to claim 1, characterized in that, The step of using the teaching semantic information as reasoning constraints, performing chain-like reasoning on the teaching content based on a multimodal large language model, and sequentially determining each knowledge point included in the teaching content, the dependencies between each knowledge point, and the order of the knowledge points includes: Using the teaching semantic information as reasoning constraints, a multimodal large language model is used to identify knowledge points in the teaching content, thereby determining the various knowledge points contained in the teaching content. Based on the identified knowledge points, a multimodal large language model is used to infer dependencies and determine the dependencies between the knowledge points. Based on the identified knowledge points and the dependencies between them, a multimodal large language model is used to sort the knowledge points and determine their order.
5. The method according to claim 1, characterized in that, The step of generating teaching materials that conform to the stated order based on each of the knowledge points and the dependencies between them includes: Each knowledge point is treated as a node, and directed edges are established between the corresponding nodes based on the dependencies between the knowledge points to generate a teaching knowledge graph. Based on the teaching knowledge graph, generate teaching materials that conform to the stated order.
6. The method according to claim 1 or 5, characterized in that, The step of generating teaching materials that conform to the stated order based on each of the knowledge points and the dependencies between them includes: Identify the core knowledge points among each of the aforementioned knowledge points; Based on the dependencies between the various knowledge points, determine the prerequisite knowledge points that the core knowledge point depends on, and determine the subsequent knowledge points that depend on the core knowledge point. The lecture notes for each knowledge point are generated sequentially according to the order of prerequisite knowledge points, core knowledge points, and subsequent knowledge points.
7. The method according to claim 1, characterized in that, The control for the digital object used for teaching to broadcast the teaching script includes: Convert the text content in the teaching script into a speech signal, and control the digital object to broadcast the speech signal; At the switching position between the various knowledge points, a control command for the digital object is triggered, which is used to control at least one of the digital object's actions, expressions, and camera angles.
8. A teaching device based on digital objects, characterized in that, include: The acquisition module is used to acquire the teaching content to be explained; The semantic processing module is used to perform semantic understanding of the teaching content based on a multimodal large language model, and determine the teaching semantic information corresponding to the teaching content; The reasoning module is used to use the teaching semantic information as reasoning constraints, perform chain reasoning on the teaching content according to the multimodal large language model, and sequentially determine each knowledge point contained in the teaching content, the dependency relationship between each knowledge point, and the arrangement order of each knowledge point. The generation module is used to generate teaching materials that conform to the arranged order based on each of the knowledge points and the dependencies between them. The control module is used to control the digital objects used for teaching to broadcast the teaching script.
9. A computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions for executing the digital object-based teaching method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the digital object-based teaching method according to any one of claims 1 to 7.