Intelligent Interactive Methods and Systems for Film and Television Teaching Based on Large Models

By using a large-scale model-based intelligent interactive method for film and television teaching, we have achieved in-depth analysis and personalized interaction of film and television teaching files, generated extended knowledge indexes and exercise indexes, solved the problem of insufficient resource utilization in the existing system, and improved teaching effectiveness and student participation.

CN120690068BActive Publication Date: 2026-05-05GUANGZHOU ACADEMY OF FINE ARTS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU ACADEMY OF FINE ARTS
Filing Date
2025-07-31
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing film and television teaching systems fail to fully utilize the in-depth information in video resources for knowledge expansion, instant questioning, or practice exercises, thus limiting students' learning interest and teaching effectiveness.

Method used

The method adopts a big model-based intelligent interactive approach for film and television teaching. It integrates teaching file management, playback settings, and handout export controls through the main interface. The big model is used to automatically parse film and television teaching files, generate subject tags and knowledge point tags, and link them with big data knowledge base and exercise bank. Personalized teaching interaction trigger points are set, and teaching handouts are automatically generated in the end.

Benefits of technology

It enables efficient management and utilization of various formats of video teaching files, enhances the interactivity and relevance of the teaching process, and improves the efficiency of handout production and teaching quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120690068B_ABST
    Figure CN120690068B_ABST
Patent Text Reader

Abstract

This application provides a method and system for intelligent interactive teaching of film and television based on a large model, belonging to the field of film and television teaching technology. This method integrates multiple functions into a single main interface, including teaching file management, playback settings, and handout export controls, enabling efficient management and utilization of film and television teaching files in various formats. The method can automatically parse video content using a large model and automatically generate subject tags and knowledge point tags, which are then linked with a big data knowledge base and exercise bank to form extended knowledge and exercise indexes, facilitating the organization and expansion of educational resources. Furthermore, it supports setting personalized teaching interaction trigger points based on these indexes, enhancing the interactivity and relevance of the teaching process. Finally, teachers can automatically generate and export teaching handouts based on all the above information, effectively improving the efficiency of handout production and the quality of teaching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of film and television teaching technology, and in particular to intelligent interactive methods and systems for film and television teaching based on large models. Background Technology

[0002] With the development of information technology and the advancement of educational philosophies, traditional classroom teaching models are gradually shifting towards digitalization and intelligence. This is particularly true in film and television education, where teaching through video resources has become a common practice. However, how to effectively utilize these rich video resources and combine them with interactive teaching to improve teaching effectiveness and student engagement has become an important research direction. Existing interactive models fail to fully utilize the in-depth information in video content for knowledge expansion, instant questioning, or practice exercises, thus limiting students' learning interest and outcomes. Summary of the Invention

[0003] This application provides a method and system for intelligent interactive teaching of film and television based on a large model, in order to solve one or more technical problems existing in the prior art, and at least provide a beneficial option or create conditions.

[0004] On the one hand, this application provides an intelligent interactive method for film and television teaching based on a large model, including the following steps:

[0005] Displays a smart interactive main interface for film and television teaching based on a large model; wherein, the main interface includes teaching file management controls, playback settings controls, and teaching handout export controls;

[0006] In response to the trigger command of the teaching file management control, the teaching file management sub-interface is displayed, various formats of film and television teaching files are uploaded, and the video screen, audio signal and subtitle text of the film and television teaching files are automatically parsed using the big data model. Subject tags and knowledge point tags are automatically generated according to the parsed content of the film and television teaching files, and they are associated with the big data knowledge base and big data exercise base to generate extended knowledge index and extended exercise index.

[0007] In response to a trigger command to the playback settings control, a playback settings sub-interface is displayed, and teaching interaction trigger points are set according to the extended knowledge index and the extended exercise index;

[0008] In response to the trigger command of the teaching handout export control, the teaching handout export sub-interface is displayed, and the teaching handout is automatically generated and exported based on the parsed content of the film and television teaching file, the extended knowledge index, the extended exercise index and the information of the teaching interaction trigger point.

[0009] Furthermore, the teaching file management sub-interface includes a file upload control, a one-click parsing control, a tag generation control, and a teaching association control;

[0010] The sub-interface for displaying teaching files allows uploading video and audio teaching files in various formats. It then uses the large-scale model to automatically parse the video footage, audio signals, and subtitles of these files. Based on the parsed content, it automatically generates subject tags and knowledge point tags, which are then linked to a large-scale knowledge base and a large-scale exercise database to generate extended knowledge indexes and extended exercise indexes. This process includes the following steps:

[0011] In response to the trigger command of the file upload control, a multi-format file selection window is displayed to receive the video and teaching files uploaded by the user;

[0012] In response to the trigger command of the one-click parsing control, a file parsing window is displayed, and the large model is invoked to automatically parse the video footage, audio signal, and subtitle text of the film and television teaching file, generating and displaying the parsed content;

[0013] In response to a trigger command on the tag generation control, a tag generation window is displayed, and based on the parsed content, the large model is invoked to generate the subject tags and the knowledge point tags;

[0014] In response to the trigger command of the teaching association control, the teaching association window is displayed, and based on the subject tag and the knowledge point tag, it is associated with the big data knowledge base and the big data exercise base to generate the extended knowledge index and the extended exercise index.

[0015] Furthermore, the subject tags and knowledge point tags are displayed in the tag generation window, allowing users to manually adjust them, and the adjusted subject tags and knowledge point tags are associated with and stored with the parsed content.

[0016] Furthermore, the teaching association window includes a label display area, a knowledge base matching control, a question bank matching control, and an index generation control;

[0017] The display teaching association window, based on the subject tags and knowledge point tags, associates them with the big data knowledge base and the big data exercise base to generate the extended knowledge index and the extended exercise index, including the following steps:

[0018] In the label display area, select the desired video teaching file, and its corresponding subject label and knowledge point label will be displayed;

[0019] In response to the trigger command of the knowledge base matching control, based on the subject tag and the knowledge point tag, related extended knowledge is retrieved from the big data knowledge base, and a mapping relationship between the extended knowledge and the timeline or content location of the film and television teaching file is established as the first mapping relationship;

[0020] In response to the trigger command of the exercise library matching control, based on the knowledge point tags, extended exercises containing the same or derived knowledge points are retrieved from the big data exercise library and associated with the exercise attributes. A mapping relationship between the extended exercises and the timeline or content location of the film and television teaching files is established as a second mapping relationship. The exercise attributes include the subject tags, the knowledge point tags, the difficulty tags, and the usage scenario tags.

[0021] In response to a trigger command to the index generation control, the subject tags, knowledge point tags, extended knowledge, extended exercises and their corresponding exercise attributes, the first mapping relationship and the second mapping relationship are integrated to generate the extended knowledge index and the extended exercise index.

[0022] Furthermore, the step of calling the large model to automatically parse the video footage, audio signals, and subtitle text of the film and television teaching file, and generating and displaying the parsed content, includes the following steps:

[0023] Keyframe sampling is performed on the video footage, and visual features are extracted through the visual processing module of the large model.

[0024] Endpoint detection and noise reduction are performed on the speech signal, and the speech-to-text is extracted by the speech processing module of the large model.

[0025] The subtitle text is segmented and entity recognized. Combined with the speech-to-text, the semantic features of the speech text are extracted using the text processing module of the large model.

[0026] The cross-modal fusion module of the large model integrates the visual features and the semantic features of the speech and text to generate and display structured parsed content, including timestamp-associated keywords and their knowledge point descriptions, key image annotations, and speech summaries.

[0027] Furthermore, the step of generating the subject tags and knowledge point tags based on the parsed content by calling the large model includes the following steps:

[0028] Based on the parsed content, the subject labels are output through the subject classification module of the large model;

[0029] The educational knowledge graph module of the large model is used to semantically match the keywords in the parsed content with the preset educational knowledge graph to generate preliminary tags.

[0030] The clustering and filtering module of the large model performs semantic clustering and redundancy filtering on the initial labels, retains the initial labels with a confidence level higher than 0.85 as the knowledge point labels, and supplements the hierarchical relationship between the knowledge point labels.

[0031] Furthermore, the playback settings sub-interface includes an interactive content marking control, an interactive configuration selection control, and an index association control;

[0032] The display playback settings sub-interface sets teaching interaction trigger points based on the extended knowledge index and the extended exercise index, including the following steps:

[0033] In response to the trigger command of the interactive content marking control, the interactive content marking window is displayed. Based on the extended knowledge index and the extended exercise index, the mapping relationship between the extended exercises and extended knowledge and the timeline or content location of the film and television teaching file is extracted, and the teaching interaction trigger point is automatically marked on the timeline or corresponding content location of the film and television teaching file.

[0034] The teaching interaction trigger points include timeline markers and content positioning floating markers in the film and television teaching files;

[0035] In response to a trigger command on the interactive configuration selection control, an interactive configuration window is displayed to receive the interaction type and trigger mode of the teaching interactive trigger point selected by the user;

[0036] The interactive types include knowledge point expansion, instant questioning, and exercise practice; the triggering modes include automatic triggering and manual triggering; automatic triggering means that the video teaching file is automatically triggered when it reaches the timeline marker point; manual triggering means that the user manually clicks the content positioning floating marker point.

[0037] In response to a trigger command to the index association control, the extended exercises and the extended knowledge are bound to the corresponding teaching interaction trigger points according to the extended knowledge index and the extended exercise index.

[0038] Furthermore, the main interface also includes a multi-screen split-screen control;

[0039] In response to the trigger command of the multi-screen split-screen control, a multi-screen split-screen sub-interface is displayed, which synchronously displays the screen of the specified video teaching file and the whiteboard. Users can drag the dividing line to adjust the split-screen ratio and support cross-screen interaction.

[0040] Furthermore, the teaching handout export sub-interface includes an export format selection control, an export content filtering control, an export layout setting control, an export path selection control, and a batch export control;

[0041] The sub-interface for exporting teaching materials automatically generates and exports teaching materials based on the parsed content of the video teaching file, the extended knowledge index, the extended exercise index, and the information of the teaching interaction trigger points, including the following steps:

[0042] In response to a trigger command on the export format selection control, a format selection window is displayed to receive the handout export format selected by the user;

[0043] In response to the trigger command of the exported content filtering control, a content selection list is displayed, and the content options selected by the user are received; the content selection list includes parsed content options, subject tag and knowledge point tag options, extended knowledge options, and extended exercise options;

[0044] In response to a trigger command on the exported layout settings control, a layout settings window is displayed to receive layout parameters set by the user;

[0045] In response to a trigger command on the export path selection control, a path selection window is displayed to receive the storage path selected by the user;

[0046] In response to the trigger command of the batch export control, a batch export window is displayed. Based on the video teaching files specified by the user, the document generation engine is invoked to generate the teaching handouts according to the handout export format, the content options, and the layout parameters, and the handouts are exported in batches according to the storage path.

[0047] On the other hand, this application provides a large-model-based intelligent interactive system for film and television teaching, which is used to execute the aforementioned large-model-based intelligent interactive method for film and television teaching.

[0048] The beneficial effects of this application are as follows: This application provides an intelligent interactive method for film and television teaching based on a large model. By integrating multiple functions into a single main interface, including teaching file management, playback settings, and handout export controls, it achieves efficient management and utilization of film and television teaching files in various formats. This method can automatically parse video content using a large model and automatically generate subject tags and knowledge point tags, which are then linked with a big data knowledge base and exercise bank to form extended knowledge and exercise indexes, facilitating the organization and expansion of educational resources. Furthermore, it supports setting personalized teaching interaction trigger points based on these indexes, enhancing the interactivity and relevance of the teaching process. Finally, teachers can automatically generate and export teaching handouts based on all the above information, effectively improving the efficiency of handout production and teaching quality. This application also provides a corresponding system; the beneficial effects of the system are similar to those of the method and will not be elaborated upon here.

[0049] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the description, claims and drawings. Attached Figure Description

[0050] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation on the technical solutions of the present invention.

[0051] Figure 1 This is a flowchart of the intelligent interactive method for film and television teaching based on a large model provided in this application;

[0052] Figure 2 This is a schematic diagram of the intelligent interactive main interface for film and television teaching based on a large model provided in this application;

[0053] Figure 3 This is a schematic diagram of the teaching document management sub-interface provided in this application;

[0054] Figure 4 This is a schematic diagram of the playback settings sub-interface provided in this application;

[0055] Figure 5 This is a schematic diagram of the teaching handout export sub-interface provided in this application;

[0056] Figure 6 This is a schematic diagram of the label generation window provided in this application;

[0057] Figure 7 This is a schematic diagram of the teaching-related window provided in this application;

[0058] Figure 8 This is a structural diagram of the large model provided in this application;

[0059] Figure 9 This is a structural diagram of the film and television teaching intelligent interactive system based on a large model provided in this application. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0061] The present application will be further described below with reference to the accompanying drawings and specific embodiments. The described embodiments should not be considered as limitations on the present application, and all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of the present application.

[0062] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0064] With the development of information technology, the education field is gradually shifting towards digitalization and intelligence. Particularly in film and television education, video, as an intuitive teaching medium, is widely used in various courses. Through film and television materials, teachers can more vividly demonstrate complex concepts, and students can gain richer learning experiences. However, how to efficiently manage these film and television resources and combine them with interactive teaching to improve teaching effectiveness and student participation has become an important research topic.

[0065] Existing systems allow users to upload, store, and play instructional videos. However, most of these systems lack support for multiple formats and suffer from inconvenience in classifying and retrieving video resources. Furthermore, the lack of effective tag recognition makes finding and using video resources difficult. Some platforms have begun to explore using artificial intelligence to analyze video content, but most are limited to basic speech recognition or simple text extraction, failing to deeply understand the complex information within the videos. While some online learning platforms offer interactive features such as Q&A and discussion forums, support for customized learning paths and interactive points based on specific video content is limited. Moreover, traditionally, teachers need to manually organize the content of instructional videos and combine it with relevant knowledge systems to create lecture notes—a time-consuming and inefficient process.

[0066] Based on technical challenges, this application proposes a method and system for intelligent interactive film and television teaching based on a large-scale model. This method first achieves efficient management and operation of various formats of film and television teaching files through a main interface integrating teaching file management controls, playback settings controls, and teaching handout export controls. It then automatically parses the content of uploaded film and television teaching files using a large-scale model, including video footage, audio, and text, and automatically generates subject tags and knowledge point tags. These tags are then linked with a big data knowledge base and exercise bank to generate extended knowledge indexes and extended exercise indexes. Furthermore, personalized teaching interaction trigger points are set based on these indexes, supporting various forms of intelligent interactive activities such as knowledge point expansion, instant questioning, and exercise practice. Finally, teaching handouts containing rich content are automatically generated and exported based on all relevant information, greatly improving the efficiency of teaching resource organization, personalized learning experience, and overall teaching effectiveness. This method not only simplifies the management process of film and television education resources but also effectively promotes the development of interactive teaching.

[0067] First, the intelligent interactive method for film and television teaching based on a large model provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0068] Reference Figures 1 to 5 The implementation process of the intelligent interactive method for film and television teaching based on a large model provided in this application includes, but is not limited to, the following steps.

[0069] Step S110: Display the intelligent interactive main interface 100 for film and television teaching based on a large model.

[0070] Among them, reference Figure 2 The main interface 100 includes a teaching file management control 101, a playback settings control 102, and a teaching handout export control 103.

[0071] In step S110, the main interface 100 provides users with an intuitive and unified operation entry point. By constructing an integrated main interface 100, teachers can quickly access the system's main functional modules, including uploading and managing teaching videos, setting playback interaction logic, and generating teaching materials. The design of the main interface 100 not only improves the convenience of user operation but also provides a good interactive foundation for subsequent intelligent processes. The teaching file management control 101 guides users into the resource management stage, the playback settings control 102 is used to configure teaching interaction logic, and the teaching materials export control 103 supports the output of teaching results. This structured interface design significantly improves the system's usability and functionality.

[0072] In step S120, in response to the trigger command of the teaching file management control 101, the teaching file management sub-interface 200 is displayed, various formats of film and television teaching files are uploaded, and the video screen, audio signal and subtitle text of the film and television teaching files are automatically parsed using a large model. Subject tags and knowledge point tags are automatically generated according to the parsed content of the film and television teaching files, and they are associated with the big data knowledge base and big data exercise base to generate extended knowledge index and extended exercise index.

[0073] In step S120, in-depth analysis and semantic understanding of the teaching video content are performed. By introducing a large model (such as a multimodal AI model), the system can simultaneously process multi-source information such as images, audio, and subtitles in the video, overcoming the limitations of traditional single-modal analysis. Based on this, the system can automatically identify the subject categories involved in the video (such as Chinese, mathematics, and English) and extract key knowledge points, thereby generating subject tags and knowledge point tags. These tags not only provide a structured representation of the teaching content but also provide data support for subsequent knowledge expansion and interactive settings. Furthermore, the system associates these knowledge points with the background big data knowledge base and exercise database to generate extended knowledge indexes and extended exercise indexes, providing rich content sources and personalized recommendation criteria for subsequent teaching interactions. This step realizes the transformation from "video" to "knowledge," which is the foundation for achieving intelligent teaching.

[0074] In step S130, in response to the trigger command of the playback settings control 102, the playback settings sub-interface 300 is displayed, and the teaching interaction trigger point is set according to the extended knowledge index and the extended exercise index.

[0075] In step S130, after completing the analysis of the teaching videos and the establishment of the knowledge index, the focus shifts to the dynamic interactive design during the teaching process. The system allows teachers or learners to set specific time points as triggers for interactive learning during video playback, such as inserting related questions, knowledge expansion hints, or practice question pop-ups when explaining a key knowledge point. These trigger points are based on the extended knowledge index and exercise index generated in the previous step, ensuring a high degree of matching and accurate delivery of interactive content to the teaching videos. In this way, students can participate in Q&A sessions, consolidate knowledge, and expand their thinking while watching the videos, greatly enhancing the immersion and efficiency of learning. Furthermore, this interactive mechanism also supports the construction of personalized learning paths, allowing students at different levels to choose different interactive content according to their needs, truly realizing the teaching philosophy of individualized instruction.

[0076] In step S140, in response to the trigger command of the teaching handout export control 103, the teaching handout export sub-interface 400 is displayed, and the teaching handout is automatically generated and exported based on the parsed content of the film and television teaching file, the extended knowledge index, the extended exercise index and the information of the teaching interaction trigger points.

[0077] In step S140, by integrating multi-dimensional data such as video content, knowledge point tags, extended knowledge index, exercise index, and interactive trigger points, the system can automatically generate a well-structured, complete, and richly illustrated teaching handout. The handout not only includes a summary of the original video's teaching content but also corresponding knowledge point summaries, links to extended resources, and pre-set interactive questions or exercises, making it convenient for teachers to use directly or share with students as after-class review material. Furthermore, the system supports exporting handouts in various formats (such as PDF and Word), improving flexibility and compatibility. This function greatly enhances teachers' work efficiency and ensures the consistency and integrity of teaching content, serving as a crucial bridge connecting teaching preparation and implementation.

[0078] In some embodiments of this application, reference is made to Figure 3 The teaching file management sub-interface 200 includes a file upload control 201, a one-click parsing control 202, a tag generation control 203, and a teaching association control 204. In step S120, in the teaching file management sub-interface 200, various formats of film and television teaching files are uploaded, and the video images, audio signals, and subtitle texts of the film and television teaching files are automatically parsed using a large model. Subject tags and knowledge point tags are automatically generated based on the parsed content of the film and television teaching files, and these are associated with the big data knowledge base and the big data exercise base to generate extended knowledge indexes and extended exercise indexes. The implementation process includes, but is not limited to, the following steps.

[0079] In step S210, in response to the trigger command of the file upload control 201, a multi-format file selection window is displayed to receive the video and teaching files uploaded by the user.

[0080] In step S210, a flexible and convenient video upload portal is provided to the user. Responding to the user's operation on the file upload control 201, the system pops up a multi-format file selection window, supporting the import of various mainstream video formats including MP4, AVI, MKV, and MOV. This ensures that teachers can easily upload various teaching video resources without prior format conversion. This step not only improves the system's compatibility and usability but also lays the foundation for subsequent automatic parsing and content processing. Furthermore, this step also has a certain file verification function, capable of identifying the integrity and parsability of video files, preventing subsequent process interruptions due to file corruption or unsupported formats, thereby improving user experience and system stability.

[0081] In step S220, in response to the trigger command of the one-click parsing control 202, the file parsing window is displayed, and the large model is called to automatically parse the video screen, audio signal and subtitle text of the film and television teaching file, and generate and display the parsed content.

[0082] In step S220, the video content is structured, transforming the original video into understandable content. Responding to the user's click on the one-click parsing control 202, the system enters the file parsing window and invokes a content extraction engine based on a large model (such as a multimodal AI model) to perform a comprehensive analysis of the uploaded video. The large model can simultaneously process multimodal data such as visual information (e.g., people, scenes, charts), audio content (e.g., narration, background noise), and subtitle text, extracting key teaching information. The system then structures this information and presents it to the user in a visual manner in the parsing window, such as keyframes on the timeline, speech-to-text summaries, and image recognition results. This process not only provides a data foundation for subsequent tag generation but also enhances teachers' understanding and control of the video content, enabling them to design lessons more accurately.

[0083] In step S230, in response to the trigger command of the tag generation control 203, the tag generation window 500 is displayed, and subject tags and knowledge point tags are generated by calling the large model based on the parsed content.

[0084] In step S230, after the automatic parsing of the video content is completed, the extracted information is further subjected to semantic analysis and classification to generate a tag system with pedagogical significance. By triggering the tag generation control 203, the system calls a large model to perform deep semantic understanding of the parsed content, automatically identifying the subject areas involved in the video, such as physics, Chinese language, and biology, as well as specific knowledge points, such as "universal gravitation," "Pythagorean theorem," "redox reaction," and "classical Chinese text appreciation," and generating corresponding subject tags and knowledge point tags. These tags not only help teachers quickly understand the theme and pedagogical value of the video content but also provide data support for subsequent resource classification, retrieval, and teaching interaction. The tag generation window 500 also supports users to manually edit or supplement tags, balancing the flexibility of automation and manual intervention. This step achieves a leap from content understanding to knowledge organization and is an important part of building an intelligent teaching system.

[0085] In step S240, in response to the trigger command of the teaching association control 204, the teaching association window 600 is displayed, and based on the subject tags and knowledge point tags, it is associated with the big data knowledge base and the big data exercise base to generate extended knowledge index and extended exercise index.

[0086] Step S240 elevates the system from content recognition to knowledge expansion, providing core support for intelligent interactive teaching. Responding to user actions on the teaching-related control 204, the system enters the teaching-related window 600. Based on generated subject and knowledge point tags, it intelligently matches these tags with the background big data knowledge base (such as a teaching knowledge point database and subject knowledge graph) and big data exercise database (such as past exam questions and mock exams). This automatically retrieves knowledge expansion points and practice questions related to the current video content and generates corresponding extended knowledge indexes and extended exercise indexes. These indexes not only serve as data sources for setting subsequent interactive teaching trigger points but also provide teachers with rich teaching resource recommendations, helping them build a more complete and systematic teaching content system. Furthermore, this step supports personalized matching logic, recommending knowledge points and exercises of different difficulty levels based on students' learning levels, interests, and preferences, thereby realizing the teaching philosophy of individualized instruction.

[0087] In summary, the four controls and their corresponding steps in the teaching file management sub-interface 200 constitute a complete process from video uploading to knowledge expansion. Through multimodal parsing driven by a large model, semantic tag generation, and knowledge association mechanisms, the system achieves in-depth mining and structured processing of teaching video content, providing solid data support for subsequent functions such as teaching interaction settings and handout generation. This process not only significantly improves the management efficiency and usability of teaching resources but also provides technical assurance for the realization of intelligent and personalized teaching.

[0088] In some embodiments of this application, in step S230, subject tags and knowledge point tags are displayed in the tag generation window 500, allowing users to manually adjust them, and the adjusted subject tags and knowledge point tags are associated with and stored with the parsed content.

[0089] This step introduces a manual intervention mechanism on top of automatic tag generation, aiming to improve the accuracy, applicability, and teaching practicality of the system-generated tags. After completing the content analysis of the film and television teaching files, the system will automatically generate preliminary subject tags and knowledge point tags based on the understanding ability of the large model, such as "function graphs", "usage of classical Chinese function words", and "principle of photosynthesis", and display them to the user in a clear and editable manner in the tag generation window 500.

[0090] The significance of this design lies in the fact that, although large models possess powerful semantic understanding capabilities, they may still exhibit recognition biases when faced with highly specialized, vaguely expressed, or interdisciplinary teaching content. Therefore, by providing a visual tag editing interface, teachers can manually adjust, supplement, or delete system-generated tags based on their own teaching experience and course needs. This includes correcting erroneous tags, adding missing knowledge points, merging similar tags, etc., thereby ensuring that the final tag system better reflects the actual teaching scenario.

[0091] Furthermore, the system has the ability to associate and store user-modified tags with the original video analysis content. This means that all manually optimized tag information will be recorded by the system and structurally linked with corresponding video clips, knowledge point summaries, and other content. This not only provides an accurate data foundation for subsequent functions such as knowledge expansion, interactive settings, and lecture note generation, but also supports the construction of a personalized teaching resource library. At the same time, this human-machine collaborative tag generation mechanism also helps train and optimize the performance of large models in specific teaching scenarios, gradually improving the system's intelligence level.

[0092] Therefore, this step, by combining the automated processing capabilities of artificial intelligence with the subjective judgment of users, achieves a unity of accuracy, flexibility, and controllability in the process of generating teaching tags. It is an important link in promoting the system's evolution from content recognition to knowledge understanding, and lays a solid foundation for the efficient operation of subsequent teaching functions.

[0093] In some embodiments of this application, reference is made to Figure 6 Taking the teaching video on "DNA double helix structure" as an example, in response to the teacher's trigger command to the tag generation control 203, the system displays the tag generation window 500. Based on the analysis content of the "DNA double helix structure" teaching video, including the dynamic demonstration of the helix structure, the transcribed text of the "base complementary pairing" audio explanation, and keywords such as "Watson-Crick discovery process" in the subtitles, the system calls the large model to automatically generate subject tags as "Biology" and knowledge point tags, such as "DNA double helix structure", "base complementary pairing principle", "deoxynucleotide composition", and "Watson and Crick experiment". The tag generation window 500 displays the above tags in a list on the left and a real-time preview of the corresponding video segment thumbnail on the right. The teacher clicks the edit button next to the "base complementary pairing principle" tag and refines it to "base complementary pairing principle (AT, CG)" in the pop-up input box. After the adjustment is completed, the system automatically associates the modified tags with the analysis content corresponding to the video and stores them in the database, and updates the tag mapping relationship in the multimodal teaching association index.

[0094] In some embodiments of this application, reference is made to Figure 7The teaching association window 600 includes a label display area 601, a knowledge base matching control 602, a question bank matching control 603, and an index generation control 604. In step S240, in the teaching association window 600, based on the generated subject labels and knowledge point labels, they are associated with the big data knowledge base and the big data question bank to generate extended knowledge indexes and extended question indexes. The implementation process includes, but is not limited to, the following steps.

[0095] In step S310, in the label display area 601, select the desired video teaching file and display its corresponding subject label and knowledge point label.

[0096] Step S310 provides an intuitive interface and basic data support for the entire teaching resource association process. By displaying the uploaded and parsed video teaching files along with their automatically generated subject and knowledge point tags in the tag display area 601, users can clearly understand the summary of the teaching content of each video. This not only facilitates teachers in quickly locating the required teaching materials but also lays the foundation for subsequent knowledge base matching and question bank matching operations. This step ensures that all pending teaching resource information is transparent and visible, enabling users to perform targeted operations based on actual teaching needs.

[0097] Step S320: In response to the trigger command of the knowledge base matching control 602, based on the subject tags and knowledge point tags, retrieve the associated extended knowledge from the big data knowledge base, and establish a mapping relationship between the extended knowledge and the timeline or content location of the film and television teaching documents as the first mapping relationship.

[0098] In step S320, the system uses subject tags and knowledge point tags as search criteria to automatically retrieve extended knowledge closely related to the current teaching video content from a large-scale knowledge base. This extended knowledge may include, but is not limited to, relevant theoretical background, application examples, and cutting-edge research progress, greatly enriching the depth and breadth of the original teaching materials. Simultaneously, the system establishes a precise mapping relationship (i.e., the first mapping relationship) between the found extended knowledge and the original video's timeline or specific content points, ensuring that corresponding supplementary materials can be accurately inserted or linked during playback. This approach not only enhances the relevance and coherence of the teaching content but also improves students' ability to acquire comprehensive knowledge, promoting deeper learning.

[0099] Step S330: In response to the trigger command of the exercise library matching control 603, based on the knowledge point tags, retrieve extended exercises containing the same or derived knowledge points from the big data exercise library and associate the exercise attributes, and establish a mapping relationship between the extended exercises and the timeline or content location of the film and television teaching files as a second mapping relationship.

[0100] The exercise attributes include subject tags, knowledge point tags, difficulty tags, and usage scenario tags.

[0101] In step S330, interactive elements, namely targeted practice questions, are added to the teaching video through an intelligent matching mechanism. When the user triggers the question bank matching control 603, the system searches the big data question bank based on knowledge point tags to find questions directly related to or derived from the current teaching video content. In addition to the questions themselves, the system also records detailed attributes of each question, such as the subject, specific knowledge point, difficulty level, and applicable teaching context. Then, the system links these selected questions to their corresponding positions in the video (timeline or content point), forming dynamic interactive nodes. This approach not only helps students check their learning outcomes in real time but also provides personalized practice suggestions based on different students' levels, further enhancing the effectiveness and interest of teaching.

[0102] Step S340: In response to the trigger command of the index generation control 604, integrate subject tags, knowledge point tags, extended knowledge, extended exercises and their corresponding exercise attributes, first mapping relationship and second mapping relationship to generate extended knowledge index and extended exercise index.

[0103] In step S340, all relevant information collected in the preceding steps is systematically organized to generate a structured index file. This includes, but is not limited to, subject tags, knowledge point tags, extended knowledge obtained from the big data knowledge base, and extended exercises and their attributes selected from the big data exercise bank. By creating extended knowledge indexes and extended exercise indexes, the system can store these valuable teaching resources in an orderly and easily accessible manner, enabling both teachers and students to quickly find the information they need. Furthermore, these index files provide strong support for subsequent teaching activities, such as customizing personalized learning paths and designing interactive teaching programs, thereby maximizing the utilization of educational resources and achieving the goal of personalized education.

[0104] In some embodiments of this application, taking the teaching video on "DNA double helix structure" as an example, in the teaching association window 600, the left label display area 601 displays the "Biology" subject label and knowledge labels "DNA double helix structure", "base complementary pairing principle", "deoxynucleotide composition", and "Watson and Crick experiment".

[0105] After the teacher clicks the knowledge base matching control 602, the system retrieves and generates cards of extended knowledge such as "The Construction Process of the Watson-Crick DNA Model," "DNA X-ray Diffraction Pattern (Franklin Data)," and "Deoxynucleotide Molecular Structure" from the big data knowledge base based on tags and displays them in the extended knowledge area 605. It also automatically establishes a timeline mapping relationship with the video. Clicking the corresponding card displays the details of the extended knowledge.

[0106] The teacher then triggers the exercise bank matching control 603. The system selects and generates three suitable exercises from the big data exercise bank based on the "base complementary pairing principle" label and displays them in the extended exercise area 606: multiple choice questions (testing AT / CG pairing rules), fill-in-the-blank questions (marking the complementary strands corresponding to DNA single-strand sequences), and experimental analysis questions. The system also automatically associates the "difficulty: medium" and "use scenario: after-class consolidation" attributes and automatically establishes a timeline mapping relationship with the video.

[0107] After clicking the index generation control 604, the system integrates the above tags, extended knowledge, extended exercises (including answer analysis paths) and their mapping relationship with the video timeline, generates extended knowledge index and extended exercise index, and stores them in the database to support subsequent teaching interaction triggers.

[0108] In some embodiments of this application, reference is made to Figure 8 The large model includes a visual processing module 810, a speech processing module 820, a text processing module 830, a cross-modal fusion module 840, a subject classification module 850, an educational knowledge graph module 860, and a clustering filtering module 870.

[0109] The visual processing module 810 is responsible for keyframe sampling of video footage in film and television teaching files and extracting visual features from them. By identifying and analyzing key images in the video, such as people, scenes, or charts, this module provides basic data support for subsequent semantic understanding and tag generation. This process not only helps in understanding the core theme of the video content but also helps teachers grasp the key points of the teaching materials more intuitively.

[0110] The speech processing module 820 performs endpoint detection and noise reduction on the speech signals in the video teaching files, and converts the processed speech into text. This is crucial for accurately extracting the explanatory information in the video. It enables the system to understand and utilize the speech content to enrich the description of teaching resources, and also provides a basis for subsequent knowledge point expansion and exercise matching.

[0111] The text processing module 830 is responsible for word segmentation and entity recognition of the subtitle text, and further extracts semantic features of the speech text by combining the speech-to-text transcription. Through comprehensive analysis of the subtitle and speech text, this module can capture detailed information conveyed in the video, including specific concepts and terms, which form the basis for generating structured parsing content.

[0112] The cross-modal fusion module 840 is used to fuse visual and semantic features extracted from video footage and audio text to generate structured parsing content that includes timestamp-associated keywords and their knowledge point descriptions, key image annotations, and audio summaries. This integration of multimodal information helps to comprehensively understand video content and improve the efficiency and quality of educational resource organization.

[0113] The subject classification module 850 is used to output subject tags based on the parsed content. By analyzing the semantic information of the video content, it can determine the main subject areas involved in the video, such as physics, chemistry, and history, providing important identifiers for the classification and retrieval of educational resources.

[0114] The Educational Knowledge Graph module 860 is used to semantically match keywords in the parsed content with a pre-defined educational knowledge graph, thereby generating preliminary tags. An educational knowledge graph is a structured form of knowledge representation that reflects the relationships between different knowledge points. By matching with the knowledge graph, this module helps the system identify specific knowledge points involved in the video, laying the foundation for further instructional interaction design.

[0115] The clustering and filtering module 870 is responsible for semantic clustering and redundancy filtering of the initial tags, retaining the initial tags with a confidence level higher than 0.85 as the final knowledge point tags, and supplementing the hierarchical relationships between the knowledge point tags. This process ensures the accuracy and logic of the tagging system, avoids interference from duplicate and irrelevant tags, and improves the effectiveness and relevance of teaching resource management.

[0116] In some embodiments of this application, the process of calling a large model to automatically parse the video images, audio signals, and subtitle text of the film and television teaching files, and generating and displaying the parsed content in step S220 includes, but is not limited to, the following steps.

[0117] Step S410: Keyframe sampling is performed on the video footage, and visual features are extracted through the visual processing module 810 of the large model.

[0118] In step S410, keyframe sampling is performed on the video footage, and visual features are extracted through the large-scale model's visual processing module 810. This process aims to identify the most representative image frames from the film and television teaching materials and extract visual information such as scenes, characters, and objects. Through the analysis of these keyframes, the system can understand the basic structure and theme of the video content, providing important visual data support for subsequent knowledge point extraction and instructional interaction design.

[0119] Step S420: Endpoint detection and noise reduction are performed on the speech signal, and the speech-to-text is extracted by the speech processing module 820 of the large model.

[0120] In step S420, endpoint detection and noise reduction are performed on the speech signal, and the speech-to-text is extracted by the speech processing module 820 of the large model. This step first preprocesses the audio part of the video to remove background noise and accurately segment the speech segments, and then converts them into text form. This not only improves the accuracy of speech recognition, but also transforms the explanation content into text data that can be further analyzed and utilized, facilitating the generation of subject tags and knowledge point tags.

[0121] Step S430: The subtitle text is segmented and entity recognized. Combined with the speech-to-text transcription, the semantic features of the speech text are extracted using the text processing module 830 of the large model.

[0122] In step S430, the subtitle text is segmented and entity recognized. Combined with the speech-to-text transcription, the text processing module 830 of the large model extracts the semantic features of the speech-to-text. This step first analyzes the subtitles in the video, identifying key terms and concepts through natural language processing technology. Simultaneously, it combines the previously extracted speech-to-text information to capture the specific knowledge content conveyed in the video. This process helps the system to more comprehensively understand the teaching information in the video, laying the foundation for generating detailed analytical content.

[0123] Step S440: Through the cross-modal fusion module 840 of the large model, visual features and speech-text semantic features are fused to generate and display structured parsed content.

[0124] The analysis includes keywords associated with timestamps and their knowledge point descriptions, key image annotations, and audio summaries.

[0125] In step S440, the cross-modal fusion module 840 of the large model fuses visual features and semantic features of speech and text to generate and display structured parsed content. In this stage, the system integrates the visual and textual information acquired in the first two steps to create a structured document containing timestamp-linked keywords and their knowledge point descriptions, key image annotations, and speech summaries. This integration of multimodal information not only makes the video content easier to understand and use but also provides teachers with rich materials for personalized instructional design and resource management. The final generated parsed content can serve as an important basis for organizing educational resources, constructing extended knowledge indexes, and creating personalized lecture notes.

[0126] In some embodiments of this application, the process of generating subject tags and knowledge point tags by calling the large model based on the parsed content in step S230 includes, but is not limited to, the following steps.

[0127] Step S510: Based on the parsed content, output subject labels through the subject classification module 850 of the large model.

[0128] In step S510, based on the parsed content, subject tags are output through the subject classification module 850 of the large model. The core function of this step is to perform macro-level subject classification of film and television teaching content. Based on the parsed content, the subject classification module 850 can identify the main subject areas to which the video content belongs (such as mathematics, physics, and Chinese), thereby providing a basic classification basis for subsequent knowledge point extraction, tag system construction, and teaching resource organization. This step helps the system quickly locate the teaching position of the video and improve the efficiency of resource retrieval and recommendation.

[0129] Step S520: Through the educational knowledge graph module 860 of the large model, the keywords in the parsed content are semantically matched with the preset educational knowledge graph to generate preliminary tags.

[0130] In step S520, the educational knowledge graph module 860 of the large model semantically matches the keywords in the parsed content with the preset educational knowledge graph to generate preliminary tags. This step aims to establish a connection between the video content and the existing structured educational knowledge system. The educational knowledge graph module 860 utilizes the semantic expression capabilities of nodes and relationships in the knowledge graph to perform semantic understanding and matching of keywords extracted from the video, identifying knowledge points corresponding to the teaching syllabus or curriculum standards. This step not only achieves accurate matching between video content and standard knowledge points but also provides semantic support for subsequent knowledge point indexing, extended recommendations, and interactive teaching.

[0131] In step S530, the clustering and filtering module 870 of the large model performs semantic clustering and redundancy filtering on the preliminary labels, retains the preliminary labels with a confidence level higher than 0.85 as knowledge point labels, and supplements the hierarchical relationship between the knowledge point labels.

[0132] In step S530, the clustering and filtering module 870 of the large model performs semantic clustering and redundancy filtering on the initial tags, retaining the initial tags with a confidence level higher than 0.85 as knowledge point tags, and supplementing the hierarchical relationships between knowledge point tags. This step is a crucial part of tag system optimization and structuring. The clustering and filtering module 870 performs semantic similarity analysis on the initially generated tags, merging semantically similar tags and removing duplicate, low-relevance, or low-confidence tags, thereby ensuring that the final output knowledge point tags have high accuracy and conciseness. At the same time, based on the structure of the knowledge graph, this module also supplements the knowledge point tags with their hierarchical relationships in the knowledge system (such as parent knowledge points, child knowledge points, parallel knowledge points, etc.), providing structured support for subsequent teaching content organization and personalized learning path recommendation.

[0133] In some embodiments of this application, reference is made to Figure 4 The playback settings sub-interface 300 includes an interactive content marking control 301, an interactive configuration selection control 302, and an index association control 303. In step S130, the process of setting the teaching interaction trigger points in the playback settings sub-interface 300 based on the extended knowledge index and the extended exercise index includes, but is not limited to, the following steps.

[0134] In step S610, in response to the trigger command of the interactive content marking control 301, the interactive content marking window is displayed. Based on the extended knowledge index and the extended exercise index, the mapping relationship between the extended exercises and extended knowledge and the timeline or content location of the film and television teaching file is extracted, and the teaching interaction trigger point is automatically marked on the timeline or corresponding content location of the film and television teaching file.

[0135] Among them, the teaching interaction trigger points include timeline markers and content positioning floating markers in film and television teaching documents.

[0136] In step S610, teachers are provided with an intuitive and convenient way to seamlessly integrate extended knowledge and exercises into the teaching videos. By clicking the interactive content marker control 301, the system will pop up an interactive content marker window, displaying information based on the previously generated extended knowledge index and extended exercise index. This information is used to determine which knowledge points or exercises should appear as interactive elements during video playback, and corresponding markers will be automatically added to the video's timeline or specific content locations. These markers are of two types: timeline markers (triggered when the video plays to a specific point) and content positioning floating markers (allowing users to manually select triggers while watching the video). This design not only enhances the interactivity of the teaching videos but also ensures that learners can obtain additional learning resources or test their understanding at key knowledge points.

[0137] In step S620, in response to the trigger command of the interactive configuration selection control 302, the interactive configuration window is displayed to receive the interaction type and trigger mode of the teaching interactive trigger point selected by the user.

[0138] The interactive types include knowledge point expansion, instant Q&A, and practice exercises. Triggering modes include automatic and manual triggering. Automatic triggering occurs when the video / educational file reaches a timeline marker. Manual triggering occurs when the user manually clicks on a floating marker within the content.

[0139] In step S620, teachers are allowed to flexibly configure interaction methods according to specific teaching objectives and student needs. By triggering the interaction configuration selection control 302, the system opens the interaction configuration window, providing options for various interaction types, such as knowledge point expansion (providing additional relevant background knowledge), immediate questioning (directly posing questions to students to test their immediate comprehension), and practice exercises (arranging relevant exercises for students to solve). Furthermore, teachers can choose a trigger mode—automatic trigger (interaction automatically starts when the video playback reaches a preset timeline marker) or manual trigger (interaction is triggered by learners actively clicking on the content's location floating marker). This flexibility allows the teaching process to be adjusted according to the actual situation, supporting both self-directed learning and guided instruction, thus improving teaching efficiency and effectiveness.

[0140] In step S630, in response to the trigger command of the index association control 303, the extended exercises and extended knowledge are bound to the corresponding teaching interaction trigger points according to the extended knowledge index and the extended exercise index.

[0141] In step S630, all prepared interactive elements are integrated to form a complete interactive teaching plan. By activating the index association control 303, the system accurately binds the selected extended exercises and extended knowledge to the previously set teaching interaction trigger points based on the information provided by the extended knowledge index and extended exercise index. This means that when the video plays to a specific marker point, the system can accurately display the corresponding knowledge point extension materials, ask immediate questions, or provide exercises, thereby achieving seamless integration between teaching content and interactive elements. This process not only simplifies teaching preparation but also ensures the effectiveness and relevance of the interactive sessions, helping to improve student participation and learning outcomes. In this way, the system realizes a transformation from simple knowledge transmission to an interactive and exploratory learning model.

[0142] In some embodiments of this application, the main interface 100 also includes a multi-screen split-screen control 104. In response to a trigger command on the multi-screen split-screen control 104, a multi-screen split-screen sub-interface is displayed, simultaneously showing the screen of the specified video / teaching file and a whiteboard. Users can drag the dividing line to adjust the split-screen ratio, supporting cross-screen interaction.

[0143] This step significantly enhances the flexibility and interactivity of the teaching process by introducing multi-screen split-screen functionality. When a user triggers the multi-screen split-screen control 104, the system displays a multi-screen split-screen sub-interface where the user can simultaneously view the specified video teaching file and a virtual whiteboard. This design allows teachers to supplement explanations, draw charts, or write formulas on the whiteboard while explaining video content, thus providing a richer teaching experience. Furthermore, users can freely adjust the ratio of the two screen areas by dragging the dividing line to suit different teaching needs and personal preferences. More importantly, this function supports cross-screen interaction, meaning that teachers can directly link their markings on the whiteboard to the video content, or allow students to participate in whiteboard operations, achieving two-way interaction. This not only promotes effective information delivery but also stimulates student participation, making the learning process more vivid and engaging. In this way, the multi-screen split-screen control 104 provides a powerful tool for modern education, improving both teaching efficiency and enriching teaching methods.

[0144] In some embodiments of this application, reference is made to Figure 5 The teaching materials export sub-interface 400 includes an export format selection control 401, an export content filtering control 402, an export layout setting control 403, an export path selection control 404, and a batch export control 405. In step S140, the process of automatically generating and exporting teaching materials in the teaching materials export sub-interface 400 based on the parsed content of the film and television teaching files, the extended knowledge index, the extended exercise index, and the information of the teaching interaction trigger points includes, but is not limited to, the following steps.

[0145] In step S710, in response to the trigger command of the export format selection control 401, the format selection window is displayed to receive the handout export format selected by the user.

[0146] In step S710, a flexible selection mechanism is provided for users, allowing them to choose the most suitable export format for teaching materials based on their actual needs. By clicking the export format selection control 401, the system will pop up a format selection window, listing various available document formats (such as PDF, Word, etc.) for users to choose from. Different formats are suitable for different scenarios and purposes. For example, PDF format is suitable for ensuring document layout consistency and security, while Word format is easier for further editing and modification. This design not only enhances the system's versatility and adaptability but also enables teachers to prepare and distribute teaching materials more efficiently.

[0147] In step S720, in response to the trigger command of the export content filtering control 402, the content check list is displayed and the content options selected by the user are received.

[0148] The content selection list includes options for parsing content, subject tags and knowledge point tags, extended knowledge, and extended exercises.

[0149] In step S720, users are given precise control over the specific content included in the final generated teaching materials. By activating the export content filtering control 402, the system displays a content selection list, which lists all available content types, such as the analysis content of the original video, automatically generated subject tags and knowledge point tags, extended knowledge obtained from a big data knowledge base, and relevant practice questions selected from a question bank. Teachers can freely select the content items to be included in the teaching materials based on the course focus and their personal preferences, thereby customizing a comprehensive and targeted teaching resource. This step not only improves the flexibility of teaching material creation but also ensures the high relevance and practicality of the teaching content.

[0150] In step S730, in response to the trigger command of the export layout settings control 403, the layout settings window is displayed to receive the layout parameters set by the user.

[0151] In step S730, to meet the personalized needs of different users, this step provides detailed layout settings. When the user clicks the export layout settings control 403, the system opens a layout settings window, allowing the user to adjust various layout parameters, such as font size, line spacing, margins, and heading styles. These settings are crucial for ensuring the professional appearance and readability of the handouts, especially when dealing with complex or technical content. By providing such fine-grained control, the system helps teachers create visually appealing and easy-to-understand teaching handouts, thereby improving students' learning experience and efficiency.

[0152] In step S740, in response to the trigger command of the export path selection control 404, the path selection window is displayed to receive the storage path selected by the user.

[0153] In step S740, after completing the content selection and layout settings for the lecture notes, the generated file is saved. By triggering the export path selection control 404, the system will pop up a path selection window, guiding the user to specify a suitable storage location. This step may seem simple, but it is indispensable for managing and organizing a large number of teaching resources. It ensures that each set of lecture notes is properly preserved and easily retrieved and used later. In addition, a clear file management system helps improve work efficiency and reduce the time wasted due to not being able to find the required materials.

[0154] In step S750, in response to the trigger command of the batch export control 405, the batch export window is displayed. Based on the video teaching files specified by the user, the document generation engine is called to generate teaching handouts according to the handout export format, content options and layout parameters, and the handouts are exported in batches according to the storage path.

[0155] In step S750, the automated generation and output of teaching materials are implemented. Once the user activates the batch export control 405, the system will enter the batch export window and, based on all previously set parameters (including the selected export format, specific content options, and personalized layout settings), call the document generation engine to create teaching materials. Afterwards, these files are automatically saved according to the user-specified storage path. This process greatly simplifies the workload of material creation, especially when generating materials for multiple video files, significantly improving work efficiency. At the same time, this also means that teachers can focus on the teaching itself, rather than tedious technical operations.

[0156] In summary, the intelligent interactive method for film and television teaching based on a large model provided in this application has the following technical effects.

[0157] This method achieves efficient management and operation of various video and audio teaching files through an integrated main interface and diverse controls. It automatically parses video content using a large model and generates subject and knowledge point tags, linking them with a big data knowledge base and question bank to generate extended knowledge and exercise indexes, greatly enriching teaching resources. The system also supports personalized interactive settings, allowing teachers to set interactive triggers as needed, including knowledge point expansion, instant questioning, and exercise practice, enhancing interactivity and participation in the learning process. Furthermore, this method provides a complete teaching handout export solution, covering export format selection, content filtering, and layout settings, significantly improving handout production efficiency.

[0158] Furthermore, the introduced multi-screen split-screen function allows for the simultaneous display of video teaching files and a whiteboard, and supports users in adjusting the split-screen ratio and cross-screen interaction, enhancing the flexibility and real-time interactivity of classroom presentations. Overall, this approach not only simplifies the teaching resource management process and improves the efficiency of educational resource utilization, but also promotes the improvement of educational quality through intelligent and personalized instructional design, providing strong technical support for modern education and contributing to a more efficient and engaging teaching experience.

[0159] Secondly, refer to Figure 9 This application provides a large-model-based intelligent interactive system for film and television teaching, used to execute the aforementioned large-model-based intelligent interactive method for film and television teaching. The system includes a main interface module 910, a teaching file management module 920, a playback settings module 930, and a teaching handout export module 940.

[0160] The main interface module 910 is used to display the intelligent interactive main interface for film and television teaching based on a large model; the main interface includes teaching file management controls, playback settings controls, and teaching handout export controls.

[0161] The teaching file management module 920 is used to respond to the trigger command of the teaching file management control, display the teaching file management sub-interface, upload film and television teaching files in various formats, and use a large model to automatically parse the video screen, audio signal and subtitle text of the film and television teaching files. Based on the parsed content of the film and television teaching files, it automatically generates subject tags and knowledge point tags, and associates them with the big data knowledge base and big data exercise base to generate extended knowledge index and extended exercise index.

[0162] The playback settings module 930 is used to respond to the trigger command of the playback settings control, display the playback settings sub-interface, and set the teaching interaction trigger points according to the extended knowledge index and the extended exercise index.

[0163] The teaching handout export module 940 is used to respond to the trigger command of the teaching handout export control, display the teaching handout export sub-interface, and automatically generate and export teaching handouts based on the parsed content of the film and television teaching files, the extended knowledge index, the extended exercise index, and the information of the teaching interaction trigger points.

[0164] Similarly, the technical effects of the system embodiments provided in this application are the same as those of the method embodiments described above, and will not be repeated here.

[0165] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this application are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.

[0166] Furthermore, although this application is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding this application. Rather, considering the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of ordinary skill of an engineer. Therefore, those skilled in the art can implement the application set forth in the claims using ordinary skill. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of this application, which is determined by the full scope of the appended claims and their equivalents.

[0167] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several programs to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0168] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequential list of executable programs for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, a program execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can retrieve and execute a program from or in conjunction with such a program execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit a program for use by or in conjunction with a program execution system, apparatus, or device.

[0169] More specific examples (a non-exhaustive list) of computer-readable media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or, if necessary, processing in a suitable manner, and then stored in computer memory.

[0170] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable program execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0171] In the foregoing description of this specification, the reference to terms such as "one embodiment / implementation," "another embodiment / implementation," or "certain embodiments / implementations," etc., indicates that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in an embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0172] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

[0173] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.

Claims

1. A film and television teaching intelligent interactive method based on a large model, characterized in that, Includes the following steps: Displays a smart interactive main interface for film and television teaching based on a large model; wherein, the main interface includes teaching file management controls, playback settings controls, and teaching handout export controls; In response to the trigger command of the teaching file management control, the teaching file management sub-interface is displayed, various formats of film and television teaching files are uploaded, and the video screen, audio signal and subtitle text of the film and television teaching files are automatically parsed using the big data model. Subject tags and knowledge point tags are automatically generated according to the parsed content of the film and television teaching files, and they are associated with the big data knowledge base and big data exercise base to generate extended knowledge index and extended exercise index. In response to a trigger command on the playback settings control, a playback settings sub-interface is displayed, which includes an interactive content marking control, an interactive configuration selection control, and an index association control. In the playback settings sub-interface, based on the extended knowledge index and the extended exercise index, the teaching interaction trigger points are set, including the following steps: In response to the trigger command of the interactive content marking control, the interactive content marking window is displayed. Based on the extended knowledge index and the extended exercise index, the mapping relationship between the extended exercises and extended knowledge and the timeline or content location of the film and television teaching file is extracted, and the teaching interaction trigger point is automatically marked on the timeline or corresponding content location of the film and television teaching file. The teaching interaction trigger points include timeline markers and content positioning floating markers in the film and television teaching files; In response to a trigger command on the interactive configuration selection control, an interactive configuration window is displayed to receive the interaction type and trigger mode of the teaching interactive trigger point selected by the user; The interactive types include knowledge point expansion, instant questioning, and exercise practice; the triggering modes include automatic triggering and manual triggering; automatic triggering means that the video teaching file is automatically triggered when it reaches the timeline marker point; manual triggering means that the user manually clicks the content positioning floating marker point. In response to the trigger command of the index association control, the extended exercises and the extended knowledge are bound to the corresponding teaching interaction trigger points according to the extended knowledge index and the extended exercise index; In response to the trigger command of the teaching handout export control, the teaching handout export sub-interface is displayed, and the teaching handout is automatically generated and exported based on the parsed content of the film and television teaching file, the extended knowledge index, the extended exercise index and the information of the teaching interaction trigger point.

2. The intelligent interactive method for film and television teaching based on a large model according to claim 1, characterized in that, The teaching file management sub-interface includes a file upload control, a one-click parsing control, a tag generation control, and a teaching association control; The sub-interface for displaying teaching files allows uploading video and audio teaching files in various formats. It then uses the large-scale model to automatically parse the video footage, audio signals, and subtitles of these files. Based on the parsed content, it automatically generates subject tags and knowledge point tags, which are then linked to a large-scale knowledge base and a large-scale exercise database to generate extended knowledge indexes and extended exercise indexes. This process includes the following steps: In response to the trigger command of the file upload control, a multi-format file selection window is displayed to receive the video and teaching files uploaded by the user; In response to the trigger command of the one-click parsing control, a file parsing window is displayed, and the large model is invoked to automatically parse the video footage, audio signal, and subtitle text of the film and television teaching file, generating and displaying the parsed content; In response to a trigger command on the tag generation control, a tag generation window is displayed, and based on the parsed content, the large model is invoked to generate the subject tags and the knowledge point tags; In response to the trigger command of the teaching association control, the teaching association window is displayed, and based on the subject tag and the knowledge point tag, it is associated with the big data knowledge base and the big data exercise base to generate the extended knowledge index and the extended exercise index.

3. The intelligent interactive method for film and television teaching based on a large model according to claim 2, characterized in that, Also includes: The label generation window displays the subject labels and the knowledge point labels, allowing users to manually adjust them, and the adjusted subject labels and knowledge point labels are associated with and stored with the parsed content.

4. The intelligent interactive method for film and television teaching based on a large model according to claim 2, characterized in that, The teaching association window includes a label display area, a knowledge base matching control, a question bank matching control, and an index generation control; The display teaching association window, based on the subject tags and knowledge point tags, associates them with the big data knowledge base and the big data exercise base to generate the extended knowledge index and the extended exercise index, including the following steps: In the label display area, select the desired video teaching file, and its corresponding subject label and knowledge point label will be displayed; In response to the trigger command of the knowledge base matching control, based on the subject tag and the knowledge point tag, related extended knowledge is retrieved from the big data knowledge base, and a mapping relationship between the extended knowledge and the timeline or content location of the film and television teaching file is established as the first mapping relationship; In response to the trigger command of the exercise library matching control, based on the knowledge point tags, extended exercises containing the same or derived knowledge points are retrieved from the big data exercise library and associated with the exercise attributes. A mapping relationship between the extended exercises and the timeline or content location of the film and television teaching files is established as a second mapping relationship. The exercise attributes include the subject tags, the knowledge point tags, the difficulty tags, and the usage scenario tags. In response to a trigger command to the index generation control, the subject tags, knowledge point tags, extended knowledge, extended exercises and their corresponding exercise attributes, the first mapping relationship and the second mapping relationship are integrated to generate the extended knowledge index and the extended exercise index.

5. The intelligent interactive method for film and television teaching based on a large model according to claim 2, characterized in that, The process of calling the large model to automatically parse the video footage, audio signals, and subtitle text of the film and television teaching file, and generating and displaying the parsed content, includes the following steps: Keyframe sampling is performed on the video footage, and visual features are extracted through the visual processing module of the large model. Endpoint detection and noise reduction are performed on the speech signal, and the speech-to-text is extracted by the speech processing module of the large model. The subtitle text is segmented and entity recognized. Combined with the speech-to-text, the semantic features of the speech text are extracted using the text processing module of the large model. The cross-modal fusion module of the large model integrates the visual features and the semantic features of the speech and text to generate and display structured parsed content, including timestamp-associated keywords and their knowledge point descriptions, key image annotations, and speech summaries.

6. The intelligent interactive method for film and television teaching based on a large model according to claim 2, characterized in that, The process of generating the subject tags and knowledge point tags based on the parsed content, by calling the large model, includes the following steps: Based on the parsed content, the subject labels are output through the subject classification module of the large model; The educational knowledge graph module of the large model is used to semantically match the keywords in the parsed content with the preset educational knowledge graph to generate preliminary tags. The clustering and filtering module of the large model performs semantic clustering and redundancy filtering on the initial labels, retains the initial labels with a confidence level higher than 0.85 as the knowledge point labels, and supplements the hierarchical relationship between the knowledge point labels.

7. The intelligent interactive method for film and television teaching based on a large model according to claim 1, characterized in that, The main interface also includes a multi-screen split-screen control; In response to the trigger command of the multi-screen split-screen control, a multi-screen split-screen sub-interface is displayed, which synchronously displays the screen of the specified video teaching file and the whiteboard. Users can drag the dividing line to adjust the split-screen ratio and support cross-screen interaction.

8. The intelligent interactive method for film and television teaching based on a large model according to claim 1, characterized in that, The teaching materials export sub-interface includes an export format selection control, an export content filtering control, an export layout setting control, an export path selection control, and a batch export control; The sub-interface for exporting teaching materials automatically generates and exports teaching materials based on the parsed content of the video teaching file, the extended knowledge index, the extended exercise index, and the information of the teaching interaction trigger points, including the following steps: In response to a trigger command on the export format selection control, a format selection window is displayed to receive the handout export format selected by the user; In response to the trigger command of the exported content filtering control, a content selection list is displayed, and the content options selected by the user are received; the content selection list includes parsed content options, subject tag and knowledge point tag options, extended knowledge options, and extended exercise options; In response to a trigger command on the exported layout settings control, a layout settings window is displayed to receive layout parameters set by the user; In response to a trigger command on the export path selection control, a path selection window is displayed to receive the storage path selected by the user; In response to the trigger command of the batch export control, a batch export window is displayed. Based on the video teaching files specified by the user, the document generation engine is invoked to generate the teaching handouts according to the handout export format, the content options, and the layout parameters, and the handouts are exported in batches according to the storage path.

9. A film and television teaching intelligent interactive system based on a large model, characterized in that: Used to perform the intelligent interactive method for film and television teaching based on a large model as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Online video training method based on large model

    CN119600858A

  • Video processing method and device based on large model, equipment and storage medium

    CN119653206A