Teaching method and device based on recorded video, electronic equipment and storage medium

By extracting subtitles from recorded videos to generate teaching syllabus and teaching scripts, the problem of lack of interactivity in traditional recorded videos is solved, and the richness and interactiveness of recorded videos is improved.

CN120151583APending Publication Date: 2025-06-13广州今之港教育咨询有限公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510159287.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional recorded classes lack interactivity, and it is difficult for students to answer questions in a timely manner when encountering them during the learning process, and it is difficult for teachers to fully understand students' learning situation.

Method used

By obtaining the subtitles of the recorded video, generating a syllabus, and generating a teaching script based on the syllabus, using the script to play the recorded video, adding interactive and Q&A sessions.

Benefits of technology

It improves the richness and interactivity of recorded videos, enables students to learn more effectively, and saves teacher resources in after-school care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151583A_ABST
    Figure CN120151583A_ABST
Patent Text Reader

Abstract

The invention discloses a teaching method and device based on recorded and broadcast videos, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining the recorded and broadcast videos of a teaching course; subtitles in the recorded and broadcast video are extracted; generating a teaching outline of the recorded video according to the subtitles; generating a teaching script according to the teaching outline; and playing the recorded video by using the teaching script. The teaching program and the teaching script are generated according to the recorded and broadcast video of the teaching course, and then the recorded and broadcast video is played, so that the content richness and interactivity of the recorded and broadcast video are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a teaching method, device, electronic device, and storage medium based on recorded videos. Background Art

[0002] With the rapid development of online education, recorded courses have become an important teaching resource. However, traditional recorded courses lack interactivity. It is difficult for students to get timely answers to their questions during the learning process, and it is also difficult for teachers to comprehensively understand the learning situation of students. In addition to the video itself, traditional recorded courses often lack supporting teaching plans, as well as relevant content such as classroom interaction, after-class exercises, and question answering. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a teaching method, device, electronic device, and storage medium based on recorded videos to improve the interactivity of the recorded videos of teaching courses.

[0004] To achieve the above object, on the one hand, the embodiments of this application propose a teaching method based on recorded videos, and the method includes the following steps:

[0005] Obtain the recorded video of the teaching course;

[0006] Extract the subtitles in the recorded video;

[0007] Generate a teaching syllabus for the recorded video according to the subtitles;

[0008] Generate a teaching script according to the teaching syllabus;

[0009] Play the recorded video using the teaching script.

[0010] In some embodiments, the extracting the subtitles in the recorded video includes the following steps:

[0011] Extract the audio in the recorded video and save it as an audio file;

[0012] Extract the speech text from the audio file as the subtitles.

[0013] In some embodiments, the generating a teaching syllabus for the recorded video according to the subtitles includes the following steps:

[0014] Input the subtitles, prompt words, and description information of the output format into a large language model;

[0015] Using the large language model, output the teaching theme, teaching objectives, teaching knowledge points, teaching content, and teaching summary of the recorded video in sequence according to the prompt words and the description information of the output format as the teaching syllabus; wherein, using the large language model, determine the start time of each teaching content according to the time of the subtitles.

[0016] In some embodiments, generating the teaching script according to the teaching syllabus includes the following steps:

[0017] Generate a roll call session;

[0018] Divide the recorded video into corresponding video playback sessions according to each teaching content in the teaching syllabus;

[0019] Add an interaction session after each teaching knowledge point in each video playback session according to the total teaching duration;

[0020] Generate a post-class Q&A session;

[0021] Sequentially determine the roll call session, each video playback session, and the post-class Q&A session as the teaching script; wherein, each interaction session is set after the corresponding teaching knowledge point in each video playback session.

[0022] In some embodiments, the steps of generating the interaction session and the Q&A session include the following steps:

[0023] Generate the text content of the exercises corresponding to each teaching knowledge point through a large language model; wherein, the exercises include at least one of multiple-choice questions, true or false questions, or connection questions;

[0024] Convert the text content of the exercises into H5 content, and then add the H5 content to the interaction session and the Q&A session.

[0025] In some embodiments, playing the recorded video using the teaching script includes at least one of the following steps:

[0026] Obtain the list of students attending class, obtain the face images of each student in the classroom through a camera device, and identify whether each face image matches the list; if a student on the list is not recognized, then conduct a voice roll call for the unrecognized student's name; dynamically receive the classroom voice, detect whether the classroom voice contains a set keyword, if the keyword is detected, then sign in the unrecognized student, if the keyword is not detected, then output the corresponding roll call information;

[0027] Alternatively, capture images in the classroom dynamically; input the captured images into a pre-trained disciplinary behavior recognition model; if a disciplinary behavior is recognized through the disciplinary behavior recognition model, then identify the faces in the captured images through face recognition, and save the information of the students corresponding to the recognized faces and the captured images; if more than a set number of disciplinary behaviors are recognized within a set time using the disciplinary behavior recognition model, then pause the playback of the recorded video and output a reminder voice to remind the disciplinary students to maintain classroom discipline;

[0028] Alternatively, play, pause, fast forward, or rewind the recorded video;

[0029] Alternatively, in an interactive session, read aloud the exercises of teaching knowledge points in a voice broadcast manner and display the content of the exercises; capture classroom images dynamically, perform pose analysis and face recognition on the classroom images to detect whether a student raises their hand; if it is detected that a student raises their hand, randomly call on the student who raised their hand to answer the question; after receiving the answer submitted by the student, judge whether the answer is correct through a large language model, and then output whether the answer is correct in voice.

[0030] In some embodiments, the method further includes the following steps:

[0031] Send the exercises of the teaching knowledge points in the recorded video to multiple mobile terminals.

[0032] To achieve the above object, on the other hand, an embodiment of the present application proposes a teaching device based on a recorded video, and the device includes:

[0033] A video acquisition unit for acquiring a recorded video of a teaching course;

[0034] A subtitle extraction unit for extracting subtitles in the recorded video;

[0035] A teaching syllabus generation unit for generating a teaching syllabus of the recorded video according to the subtitles;

[0036] A teaching script generation unit for generating a teaching script according to the teaching syllabus;

[0037] A video teaching unit for playing the recorded video using the teaching script.

[0038] To achieve the above object, on the other hand, an embodiment of the present application proposes an electronic device, and the electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above method is implemented.

[0039] To achieve the above object, another aspect of the embodiments of the present application proposes a computer-readable storage medium storing a computer program, which when executed by a processor implements the above method.

[0040] The embodiments of the present application at least include the following beneficial effects:

[0041] The present application can obtain the recorded video of the teaching course; extract the subtitles in the recorded video; generate the teaching syllabus of the recorded video according to the subtitles; generate the teaching script according to the teaching syllabus; and play the recorded video using the teaching script. The present application generates the teaching syllabus and the teaching script based on the recorded video of the teaching course, and then plays the recorded video, greatly improving the content richness and interactivity of the recorded video. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 It is a schematic flowchart of the teaching method based on the recorded video provided by the embodiments of the present application;

[0044] Figure 2 It is an example diagram of the functional modules of the teaching method based on the recorded video provided by the embodiments of the present application;

[0045] Figure 3 It is an example diagram of a multiple-choice question in the H5 content form provided by the embodiments of the present application;

[0046] Figure 4 It is an example diagram of a true or false question in the H5 content form provided by the embodiments of the present application;

[0047] Figure 5 It is an example diagram of a matching question in the H5 content form provided by the embodiments of the present application;

[0048] Figure 6 It is a schematic structural diagram of the teaching device based on the recorded video provided by the embodiments of the present application;

[0049] Figure 7 It is a schematic hardware structure diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0051] It can be understood that the terms "first", "second", etc. used in the present application can be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "while...", or "in response to determining".

[0052] The terms "at least one", "multiple", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0054] Embodiments of the present application provide a teaching method, device, electronic device, and storage medium based on recorded videos, which relate to the field of artificial intelligence technology. The teaching method, device, electronic device, and storage medium based on recorded videos provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server, or can be software running on a terminal or a server. In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side may be configured as an independent physical server, may also be configured as a server cluster or a distributed system composed of multiple physical servers, or may also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server may also be a node server in a blockchain network; the software may be an application implementing the teaching method based on recorded videos, etc., but is not limited to the above forms.

[0055] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0056] Refer to Figure 1 , embodiments of the present application provide a teaching method based on recorded videos, and the method may include but is not limited to S100 to S140, specifically as follows:

[0057] S100: Obtain a recorded video of a teaching course.

[0058] It can be understood that in this embodiment, any pre-recorded teaching video can be obtained as the recorded video of the teaching course.

[0059] S110: Extract subtitles from the recorded video.

[0060] Further, S110 may include the following steps S111 to S112:

[0061] S111: Extract the audio from the recorded video and save it as an audio file;

[0062] S112: Extract the speech text from the audio file as the subtitle.

[0063] S120: Generate a teaching syllabus for the recorded video according to the subtitle.

[0064] Further, S120 may include the following steps S121 to S122:

[0065] S121: Input the subtitle, prompt words, and description information of the output format into a large language model;

[0066] S122: Use the large language model to sequentially output the teaching theme, teaching objectives, teaching knowledge points, teaching content, and teaching summary of the recorded video as the teaching syllabus according to the prompt words and the description information of the output format; wherein, the start time of each teaching content is determined according to the time of the subtitle by using the large language model.

[0067] S130: Generate a teaching script according to the teaching syllabus.

[0068] Further, S130 may include the following steps S131 to S135:

[0069] S131: Generate an attendance-taking session;

[0070] S132: Divide the recorded video into corresponding video playback sessions according to each teaching content in the teaching syllabus;

[0071] S133: Add an interaction session after each teaching knowledge point in each video playback session according to the total teaching duration;

[0072] S134: Generate a Q&A session after class;

[0073] S135: Sequentially determine the attendance-taking session, each video playback session, and the Q&A session after class as the teaching script; wherein, each interaction session is set after the corresponding teaching knowledge point in each video playback session.

[0074] Even further, the steps of generating the interaction session and the Q&A session include the following steps:

[0075] Generate the text content of the exercises corresponding to each teaching knowledge point through a large language model; wherein, the exercises include at least one of multiple-choice questions, true or false questions, or connection questions;

[0076] Convert the text content of the exercise into H5 content, and then add the H5 content to the interactive session and the Q&A session.

[0077] S140: Play the recorded video using the teaching script.

[0078] Furthermore, S140 may include at least one of the following steps S141 to S144:

[0079] S141: Obtain the list of students in class, obtain the face images of each student in the classroom through a camera device, and identify whether each face image matches the list; if a student on the list is not recognized, then conduct a voice roll call for the name of the unrecognized student; dynamically receive classroom voice, detect whether the classroom voice contains a set keyword, if the keyword is detected, then sign in the unrecognized student, if the keyword is not detected, then output the corresponding roll call information;

[0080] S142: Dynamically obtain the captured images in the classroom; input the captured images into a pre-trained disciplinary behavior recognition model; if a disciplinary behavior is recognized through the disciplinary behavior recognition model, then identify the face in the captured image through face recognition method, and save the information of the student corresponding to the recognized face and the captured image; if more than a set number of disciplinary behaviors are recognized within a set time using the disciplinary behavior recognition model, then pause the playback of the recorded video and output a reminder voice to remind the disciplinary students to maintain classroom discipline;

[0081] S143: Play, pause, fast forward or rewind the recorded video;

[0082] S144: In the interactive session, read the exercises of the teaching knowledge points by voice broadcast and display the content of the exercises; dynamically obtain classroom images, conduct pose analysis and face recognition on the classroom images to detect whether there are students raising their hands; if it is detected that there are students raising their hands, then randomly call on the students who raise their hands to answer questions; after receiving the answers submitted by the students, judge whether the answers are correct through a large language model, and then output by voice whether the answers are correct.

[0083] As a further implementation manner, the embodiment of the present application may further include the following step S150:

[0084] S150: Send the exercises of the teaching knowledge points in the recorded video to multiple mobile terminals.

[0085] It can be understood that by sending the exercises of the teaching knowledge points to multiple mobile terminals, students can practice during the teaching process of the recorded video or after the recorded video ends, further increasing the interactivity.

[0086] Next, the solutions of the embodiments of the present application will be introduced and described in detail in combination with specific application examples.

[0087] In after-school care, improving the teaching effect and learning experience of recorded lessons can greatly improve the teaching quality of after-school care. Therefore, it is of great significance to develop an artificial intelligence teaching assistant solution with intelligent assistance functions.

[0088] Based on this, in this embodiment, the audio of the recorded lesson video is extracted, corresponding subtitles are generated, the teaching syllabus is extracted through the big fish-eye model, and the teaching script, interactive content, and after-class exercises are generated to enrich the teaching content and improve the interactivity of the recorded video.

[0089] During teaching, teaching will be carried out according to the links of the teaching script. Through the application of artificial intelligence such as visual analysis, gesture analysis, speech synthesis, speech recognition, and face recognition, combined with the control of the player, diversified teaching is achieved to improve the teaching effect and experience of recorded lessons.

[0090] Exemplarily, the solution of this embodiment may include functional modules as Figure 2 shown.

[0091] Specifically, the solution of this embodiment may include the following content:

[0092] Recorded lesson analysis and processing module: This module analyzes the recorded video, sorts out the video content through the large language model, and grasps what knowledge points there are, the content of each link of the recorded lesson video, and the time corresponding to each link. Specifically, it includes the following solutions:

[0093] 1. Audio extraction:

[0094] Extract the audio of the video through ffmpeg for subsequent processing.

[0095] 2. Subtitle generation:

[0096] Extract the subtitles of the extracted audio file through the AI subtitle extraction module (medium model) based on the open-source item Whispe and save it as a txt file.

[0097] 3. Teaching syllabus extraction:

[0098] Input the extracted subtitles, prompt words, and description information of the output format into the large language model to extract the teaching syllabus. For the teaching content part, the start time of each link of the teaching content in the video will be supplemented according to the time of the subtitles. An example of the output format of the large model teaching syllabus is as follows:

[0099] I. Teaching theme;

[0100] II. Teaching objectives;

[0101] III. Key and Difficult Points in Teaching (Teaching Knowledge Points);

[0102] Key Points;

[0103] Difficult Points;

[0104] IV. Teaching Content;

[0105] Introduction of Scenario (hh:mm:ss:sss);

[0106] Teaching Explanation (hh:mm:ss:sss);

[0107] Knowledge Point 1: (hh:mm:ss:sss);

[0108] Knowledge Point 2: (hh:mm:ss:sss);

[0109] Knowledge Point 3: (hh:mm:ss:sss);

[0110] Summary (hh:mm:ss:sss);

[0111] V. Teaching Summary (hh:mm:ss:sss).

[0112] 4. Generation of Teaching Script:

[0113] The teaching script is used for video control and content playback during class according to the links and time in the script content.

[0114] The content of the teaching script includes the link name, duration, and type.

[0115] The types are: roll call, video playback, interaction, Q&A, and after-class exercises.

[0116] Steps to generate the teaching script:

[0117] 1. By default, the first link is a 2-minute roll call link.

[0118] 2. Generate links of the type of video playback according to the teaching content part in the teaching syllabus.

[0119] 3. According to the total teaching duration, add an interaction link after teaching each knowledge point. The duration of the interaction link = (total teaching duration - video duration) / number of knowledge points + 2). If the calculated result > 5 minutes, then take 5 minutes.

[0120] 4. Generate an after-class Q&A link at the end of the course.

[0121] Explanation of the data generation module:

[0122] First, generate corresponding interactive questions and after-class exercise content through a large language model. The generated content will be divided into: multiple-choice questions, true or false questions, and matching questions.

[0123] After the large language model generates the text content, convert the text content into H5 content for display in class. The format examples are as Figure 3 , Figure 4 and Figure 5 shown.

[0124] Teaching module: Conduct teaching according to the teaching script, specifically including:

[0125] 1. Take attendance before class.

[0126] Before taking attendance, obtain the list of students who have asked for leave in the system and take attendance for non-leaving students.

[0127] Capture the faces of the students in the classroom through the camera for face recognition. If all recognitions are successful, proceed to the next step.

[0128] If not all are recognized, enter the voice attendance step. After synthesizing the voices of the names through ChatTTS, broadcast the attendance through the voice device of the classroom supporting TV. Real-time collect the voices of the students in the classroom, convert the voice to text, and after extracting the keyword "present", sign in the students.

[0129] If no feedback from the students is received after more than 3 attendance attempts, notify the patrol and supervision teacher through forms such as text messages and APP push.

[0130] 2. Discipline management.

[0131] Use the open-source deep learning framework PyTorch. After a large number of photo data collected during after-school custody are manually marked, conduct model training.

[0132] After capturing the classroom pictures through the camera, conduct behavior analysis through the trained model. If a behavior violating classroom discipline is encountered, save the information and behavior of the student violating the discipline through face recognition. If the number of violations exceeds 10 times within 3 minutes during capture, pause the playback of the recorded lesson and remind the corresponding students to maintain classroom discipline through voice.

[0133] If the analysis result of the captured student behavior is unavailable, send the video of the 10 seconds before and after the capture to the server. After manual marking, continue to train the model.

[0134] 3. Player control.

[0135] Control the playback, pause, fast forward, and rewind of the recorded video according to the teaching script.

[0136] Roll call completed: Start playing the video.

[0137] Explanation of knowledge points completed: Pause the video and enter the interactive session.

[0138] Interactive session ended: Play the video.

[0139] 4. Classroom interaction model.

[0140] First, it will read the title of the interactive content through speech synthesis and display the H5 page generated by the data generation module.

[0141] At the same time, it will analyze which student has raised their hand to answer the question through a capture camera + pose analysis + face recognition.

[0142] Randomly call on students who have raised their hands to answer the question.

[0143] After the student submits the answer, a voice broadcast of correct or incorrect will be given through the large language model.

[0144] 5. After-class exercises.

[0145] After the course ends, send after-class exercise questions to the application on the mobile terminal.

[0146] The technical effects of this embodiment include:

[0147] Based on the recorded video of the teaching course, the corresponding teaching content can be improved and enriched. At the same time, through artificial intelligence solutions such as large language models, pose analysis, and speech recognition, it assists in the teaching process, improves the richness and interactivity of the recorded video, and can also save teacher resources in after-school care.

[0148] Refer to Figure 6 , this application embodiment also provides a teaching device based on the recorded video, which can implement the above-mentioned teaching method based on the recorded video. The device includes:

[0149] Video acquisition unit, used to acquire the recorded video of the teaching course;

[0150] Subtitle extraction unit, used to extract the subtitles in the recorded video;

[0151] Teaching syllabus generation unit, used to generate the teaching syllabus of the recorded video according to the subtitles;

[0152] Teaching script generation unit, used to generate a teaching script according to the teaching syllabus;

[0153] Video teaching unit, used to play the recorded video using the teaching script.

[0154] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0155] An embodiment of the present application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above teaching method based on the recorded video. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0156] It can be understood that the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0157] Please refer to Figure 7 , Figure 7 , which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0158] A processor 701, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0159] A memory 702, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 702 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 702 and are called by the processor 701 to execute the teaching method based on the recorded video in the embodiments of the present application;

[0160] An input / output interface 703, which is used to implement information input and output;

[0161] A communication interface 704, which is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0162] A bus 705 transmits information between various components of the device (such as a processor 701, a memory 702, an input / output interface 703, and a communication interface 704);

[0163] Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 achieve communication connections with each other inside the device through the bus 705.

[0164] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above teaching method based on recorded video is implemented.

[0165] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0166] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0167] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0168] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.

[0169] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0170] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.

[0171] As used in the description of the present application and the above accompanying drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0172] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or a similar expression thereof refers to any combination of these items, including any combination of single items or plural items. For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0173] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0174] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0175] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0176] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs and other various media that can store programs.

[0177] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. The teaching method based on recorded video is characterized by: The method comprises the following steps: Get the recorded videos of teaching courses; Extracting subtitles from the recorded video; Generating a teaching outline of the recorded video according to the subtitles; Generate a teaching script according to the teaching syllabus; The recorded video is played using the teaching script.

2. The teaching method based on recorded video according to claim 1, characterized in that: The step of extracting subtitles from the recorded video comprises the following steps: Extracting the audio from the recorded video and saving it as an audio file; Extracting speech text from the audio file as the subtitle.

3. The teaching method based on recorded video according to claim 1, characterized in that: The step of generating a teaching outline for the recorded video according to the subtitles comprises the following steps: Inputting the subtitles, prompt words and output format description information into a large language model; The large language model is used to output the teaching theme, teaching objectives, teaching knowledge points, teaching content and teaching summary of the recorded video in sequence as the teaching outline according to the prompt words and the description information of the output format; wherein the large language model is used to determine the start time of each teaching content according to the time of the subtitles.

4. The teaching method based on recorded video according to claim 1, characterized in that: Generating a teaching script according to the teaching syllabus includes the following steps: Generate a roll call session; Dividing the recorded video into corresponding video playback links according to various teaching contents in the teaching syllabus; According to the total teaching time, an interactive session is added after each teaching knowledge point in each video playback session; Generate post-class Q&A session; The roll call session, each of the video playing sessions and the after-class question-and-answer session are sequentially determined as the teaching script; wherein each of the interactive sessions is arranged after the corresponding teaching knowledge point in each of the video playing sessions.

5. The teaching method based on recorded video according to claim 4, characterized in that: The steps of generating the interactive session and the question-answering session include the following steps: Generate text content of exercises corresponding to each of the teaching knowledge points through a large language model; wherein the exercises include at least one of multiple-choice questions, true-or-false questions, or connecting-line questions; The text content of the exercise is converted into H5 content, and then the H5 content is added to the interactive session and the question-and-answer session.

6. The teaching method based on recorded video according to claim 1, characterized in that: The step of playing the recorded video using the teaching script comprises at least one of the following steps: Obtain a list of students in class, obtain facial images of each student in the classroom through a camera device, and identify whether each facial image matches the list; if a student on the list is not identified, then voice roll call the name of the unidentified student; dynamically receive classroom voice, detect whether the classroom voice contains set keywords, if the keyword is detected, then sign in the unidentified student, if the keyword is not detected, then output the corresponding roll call information; Alternatively, dynamically obtain a captured image of the classroom; input the captured image into a pre-trained violation recognition model; if a violation is identified by the violation recognition model, identify the face in the captured image by a face recognition method, and save the information of the student corresponding to the identified face and the captured image; If more than a set number of violations are identified by the violation identification model within a set time, the playback of the recorded video is paused, and a reminder voice is output to remind the violating students to maintain classroom discipline; Alternatively, play, pause, fast forward or rewind the recorded video; Alternatively, in the interactive session, exercises of the teaching knowledge points are read aloud through voice broadcast, and the content of the exercises is displayed; classroom images are dynamically acquired, and posture analysis and face recognition are performed on the classroom images to detect whether there are students raising their hands; if a student is detected raising his hand, the student who raised his hand is randomly called to answer the question; after receiving the answer submitted by the student, the large language model is used to determine whether the answer is correct, and then the voice output is used to determine whether the answer is correct.

7. The teaching method based on recorded video according to any one of claims 1 to 6, characterized in that: The method further comprises the following steps: Sending exercises of the teaching knowledge points in the recorded video to multiple mobile terminals.

8. A teaching device based on recorded video, characterized in that: The device comprises: A video acquisition unit, used to acquire recorded videos of teaching courses; A subtitle extraction unit, used to extract subtitles from the recorded video; A teaching syllabus generating unit, configured to generate a teaching syllabus for the recorded video according to the subtitles; A teaching script generating unit, used to generate a teaching script according to the teaching syllabus; The video teaching unit is used to play the recorded video using the teaching script.

9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • System for searching teaching video contents and method thereof

    CN102184259A

  • Method and system for editing and playing interactive video, and electronic learning device

    CN102833490A

  • Cloud-end framework based on teaching plan recording library and pseudo students in online teaching system

    CN109816571A

  • Data processing method and device in paperless scene, medium and electronic equipment

    CN110378665A

  • Terminal-based teacher-student remote interactive education platform

    CN113781853A