Digital human courseware automatic generation system

Through the automatic generation system of digital courseware, the teacher's teaching video and audio are automatically collected and processed, and the digital courseware is generated, which solves the problems of low efficiency and poor generalization of electronic courseware production, and realizes students' invisible review.

CN120378709APending Publication Date: 2025-07-25HANGZHOU BAONUO INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510454730.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the production of electronic courseware is inefficient and poorly applicable, making it difficult to meet students' after-class review needs.

Method used

The digital human courseware automatic generation system is adopted, and the teacher's teaching video and audio are automatically collected through the image and audio acquisition module, and the data processing module performs format conversion and filtering. The courseware generation module calls the preset digital human image generation courseware.

Benefits of technology

It realizes the inability to obtain teacher teaching content and automatically generate digital courseware. Students can fully review classroom content, improving the production efficiency and universality of electronic courseware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378709A_ABST
    Figure CN120378709A_ABST
Patent Text Reader

Abstract

The invention relates to a digital human courseware automatic generation system, and the system comprises an image collection module which is used for collecting a teacher close-up video, a blackboard writing video and a computer display video in the teaching process of a teacher; the audio acquisition module is used for acquiring teaching audio in the teaching process of the teacher; the data processing module is used for pushing the teacher close-up video, the blackboard writing video, the computer display video and the teaching audio to the courseware generation module; and the courseware generation module is used for calling a preset digital human image to generate digital human courseware according to the teacher close-up video, the blackboard writing video, the computer display video and the teaching audio. According to the method and the device, the classroom teaching content of the teacher can be acquired without feeling, the digital person courseware is automatically generated, when the student opens the digital person courseware to watch, the teaching content of the teacher in the classroom can be completely reviewed, and the problem of how to efficiently manufacture the electronic courseware with high universality is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a digital human courseware automatic generation system. Background Art

[0002] With the advent of the information age, all fields are gradually moving towards digitalization, and the education industry is no exception. Generally, digital teaching requires teachers to prepare electronic courseware in advance before teaching. The traditional method is for teachers to personally select teaching content and manually make courseware through specific software. This method has the problems of high cost and low efficiency. Moreover, this manually made electronic courseware is generally only suitable for teachers to use in the classroom. For students who want to use it for review after class, due to the lack of teachers' classroom explanations, the effect will be significantly reduced.

[0003] Currently, for the problem of how to efficiently produce electronic courseware with high versatility in related technologies, no effective solution has been proposed. Summary of the Invention

[0004] An embodiment of this application provides a digital human courseware automatic generation system to at least solve the problem of how to efficiently produce electronic courseware with high versatility in related technologies.

[0005] In a first aspect, an embodiment of this application provides a digital human courseware automatic generation system, which includes an image acquisition module, an audio acquisition module, a data processing module, and a courseware generation module;

[0006] The image acquisition module is used to acquire the teacher's close-up video, the blackboard writing video, and the computer display video during the teacher's teaching process;

[0007] The audio acquisition module is used to acquire the teaching audio during the teacher's teaching process;

[0008] The data processing module is used to push the teacher's close-up video, the blackboard writing video, the computer display video, and the teaching audio to the courseware generation module;

[0009] The courseware generation module is used to generate digital human courseware by calling a preset digital human image according to the teacher's close-up video, the blackboard writing video, the computer display video, and the teaching audio.

[0010] In some of these embodiments, the image acquisition module includes a first image acquisition module and a second image acquisition module;

[0011] The first image acquisition module is used to obtain the target area where the teacher is located through a target detection algorithm during the teacher's teaching process, acquire the teacher's close-up video within the target area, and acquire the blackboard writing video within a preset area;

[0012] The second image acquisition module is configured to acquire the computer display video during the teacher's teaching process.

[0013] In some embodiments, the system includes:

[0014] The data processing module is configured to convert the acquired blackboard writing video and the computer display video from the physical signal format into the picture media stream format;

[0015] The data processing module is configured to intercept video pictures of the video in the picture media stream format at intervals of three seconds respectively, and after screening the intercepted video pictures, push them to the courseware generation module.

[0016] In some embodiments, the data processing module is configured to push the first video picture intercepted from the picture to the courseware generation module, and calculate the similarity between adjacent video pictures starting from the second video picture. If the similarity is less than a preset threshold, it indicates that the video picture is in a changing state. If the similarity is greater than the preset threshold, it indicates that the video picture is in a stable state;

[0017] In the case where the video picture is in a changing state, the video picture is not pushed to the courseware generation module. Among them, if the video picture is in a changing state for ten consecutive times, a timeout push is triggered;

[0018] In the case where the video picture is in a stable state for three consecutive times, the video picture is pushed to the courseware generation module.

[0019] In some embodiments, the courseware generation module is configured to call a preset digital human image to generate a digital human courseware according to the video pictures of the blackboard writing video and the computer display video respectively intercepted and screened, as well as the teacher's close-up video and the teaching audio.

[0020] In some embodiments, the system includes:

[0021] The audio acquisition module includes a first audio acquisition module and a second audio acquisition module;

[0022] The first audio acquisition module is configured to acquire the teacher's lecture audio during the teacher's teaching process;

[0023] The second audio acquisition module is configured to acquire the computer playback audio during the teacher's teaching process.

[0024] In some of these embodiments, the data processing module is further configured to perform frontal face detection on the teacher close-up video through a face detection algorithm to obtain the frontal face image of the teacher appearing in the teacher close-up video, and push the frontal face image of the teacher to the courseware generation module.

[0025] In some of these embodiments, the courseware generation module is further configured to perform face recognition on the frontal face image of the teacher appearing in the teacher close-up video. If the teacher passes the face recognition, the teacher is the target teacher.

[0026] In some of these embodiments, the courseware generation module is configured to determine the courseware timeline of the digital human courseware according to the duration of the lecture audio, and convert the lecture audio into lecture text;

[0027] The courseware generation module is configured to call the preset digital human image of the target teacher to generate a digital human teaching video with a duration consistent with the courseware timeline, behaviors consistent with the teacher close-up video, and lip movements consistent with the lecture text;

[0028] The courseware generation module is configured to generate the digital human courseware of the target teacher according to the blackboard writing video, the computer display video, the lecture audio, the computer playback audio, and the digital human teaching video.

[0029] In some of these embodiments, the system includes:

[0030] The data processing module is configured to convert the collected teacher close-up video from the physical signal format into a video media stream format, and then push it to the courseware generation module;

[0031] The data processing module is configured to convert the collected lecture audio from the physical signal format into a sound media stream format, and then push it to the courseware generation module.

[0032] Compared with the related art, a digital human courseware automatic generation system provided by an embodiment of the present application includes: an image acquisition module for acquiring a teacher close-up video, a blackboard writing video, and a computer display video during the teacher's teaching process; an audio acquisition module for acquiring the lecture audio during the teacher's teaching process; a data processing module for pushing the teacher close-up video, the blackboard writing video, the computer display video, and the lecture audio to the courseware generation module; and a courseware generation module for calling a preset digital human image according to the teacher close-up video, the blackboard writing video, the computer display video, and the lecture audio to generate a digital human courseware, realizing the non-sensing acquisition of the teacher's classroom teaching content and automatically generating the digital human courseware. When a student opens the digital human courseware for viewing, they can completely review the teacher's teaching content in the classroom, solving the problem of how to efficiently produce electronic courseware with high versatility. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings described herein are provided to further understand the present application and form a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0034] Figure 1 is a structural block diagram of a digital human courseware automatic generation system according to an embodiment of the present application;

[0035] Figure 2 is a schematic diagram of the interface of a digital human courseware according to an embodiment of the present application;

[0036] Figure 3 is a schematic flow diagram of intercepting the frontal face of a teacher according to an embodiment of the present application;

[0037] Figure 4 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be described and explained below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.

[0039] Obviously, the drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, the present application can be applied to other similar scenarios based on these drawings without making creative efforts. In addition, it can also be understood that although the efforts made in such a development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.

[0040] Referring to "embodiment" in the present application means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.

[0041] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one", "the" and the like involved in this application do not indicate a quantity limitation and may represent a singular or plural number. The terms "comprise", "include", "have" and any variations thereof involved in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include steps or units not listed, or may further include other steps or units inherent to these processes, methods, products or devices. The words such as "connect", "be connected", "couple" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0042] Embodiment 1

[0043] An embodiment of this application provides a digital human courseware automatic generation system. Figure 1 It is a structural block diagram of the digital human courseware automatic generation system according to the embodiment of this application, as Figure 1 shown. The system includes an image acquisition module, an audio acquisition module, a data processing module and a courseware generation module.

[0044] The image acquisition module is used to acquire the teacher's close-up video, blackboard writing video and computer display video during the teacher's teaching process.

[0045] Preferably, the image acquisition module includes a first image acquisition module and a second image acquisition module. The first image acquisition module is used to obtain the target area where the teacher is located through a target detection algorithm during the teacher's teaching process, acquire the teacher's close-up video within the target area, and acquire the blackboard writing video within a preset area. The second image acquisition module is used to acquire the computer display video during the teacher's teaching process. It should be emphasized that the digital human courseware of the embodiment of this application is automatically recorded and generated according to the class schedule information of each classroom at the start time of the class. The whole process does not require manual intervention and is non-intrusive and automated. In other words, when the class time in the class schedule arrives, the image acquisition module will automatically start acquiring data without manual intervention.

[0046] For example, the above first image acquisition module can be a network camera, which is fixed in the middle of the classroom and directly faces the panoramic view of the teacher's class; the area of the blackboard writing is framed within the panoramic view. After the framing is completed, the camera outputs the blackboard writing video of the framed area, and then uses a target detection algorithm to intelligently detect the target area where the teacher is located, generates a dynamic close-up that can follow the teacher's movement, and outputs the teacher's close-up video; the camera (the first image acquisition module) is connected to the data processing module through a local area network to push the teacher's close-up video and the blackboard writing video to the data processing module through the network. The above second image acquisition module can be a desktop computer, which is connected to the data processing module through an HDMI signal cable and transmits the collected computer display video to the data processing module.

[0047] An audio acquisition module for acquiring the teaching audio during the teacher's class;

[0048] Preferably, the audio acquisition module includes a first audio acquisition module and a second audio acquisition module; the first audio acquisition module is used for acquiring the teacher's lecture audio during the teacher's class; the second audio acquisition module is used for acquiring the computer-played audio during the teacher's class. It should be emphasized that, similarly, when the class time in the class schedule arrives, the audio acquisition module will automatically start collecting data without manual intervention, and the generation of the courseware is completely non-intrusive and automated.

[0049] For example, the above first audio acquisition module can be a pickup, which is installed in the podium area of the classroom and uses an omnidirectional microphone to collect the teacher's lecture audio in real time, and transmits the sound to the data processing module through a 3.5mm audio cable. The above second audio acquisition module can be a desktop computer, which transmits the collected computer-played audio to the data processing module through an HDMI signal cable.

[0050] A data processing module for pushing the teacher's close-up video, the blackboard writing video, the computer display video, and the teaching audio to the courseware generation module;

[0051] Preferably, the data processing module is used to convert the collected teacher close-up video, blackboard writing video, and computer display video from the physical signal format into the picture media stream format, and then push them to the courseware generation module; convert the collected teaching audio from the physical signal format into the sound media stream format, and then push it to the courseware generation module. Further, for the collected blackboard writing video, the data processing module needs to optimize the picture content through an image quality enhancement algorithm. Specifically: ① Picture preprocessing: Automatically identify the four corners of the whiteboard picture through the algorithm, and use the perspective correction algorithm to correct the picture by fusing the straight line detection of the whiteboard edge to eliminate the shooting angle distortion; then use the Retinex theory to decompose the reflection / illumination component for illumination equalization processing to ensure that the illumination level of the blackboard writing picture is consistent. ② Picture enhancement: Establish a teaching-specific color card library (chalk / whiteboard pen color gamut database), adaptively enhance the saturation of the color of the blackboard writing trajectory, and at the same time achieve differential sharpening of the edges of strokes of different thicknesses through dynamic convolution. ③ Post-processing optimization: Perform blackboard / whiteboard scene detection and dynamically switch the color enhancement strategy (emphasize contrast for blackboard / suppress reflection for whiteboard). The optimized blackboard writing video is obtained through the above three steps.

[0052] The courseware generation module is used to generate a digital human courseware by calling a preset digital human image according to the above-mentioned teacher close-up video, blackboard writing video, computer display video, and teaching audio.

[0053] It should be noted by way of example that the above courseware generation module can be a digital human courseware platform, which is connected to the data processing module through the Internet and receives the teacher close-up video, blackboard writing video, computer display video, and teaching audio pushed by the data processing module.

[0054] Through the embodiments of the present application, when the class time in the class schedule arrives, the image acquisition module, audio acquisition module, data processing module, and courseware generation module will automatically start collecting data and generating courseware without manual intervention. The generation of the courseware is completely non-sensing and automated, realizing the non-sensing acquisition of the teacher's classroom teaching content and automatically generating a digital human courseware. When students open the digital human courseware for viewing, they can completely review the teacher's teaching content in class, solving the problem of how to efficiently produce electronic courseware with high versatility.

[0055] Embodiment 2

[0056] In some specific embodiments, on the basis of the above Embodiment 1, this specific embodiment provides a digital human courseware automatic generation system, which includes:

[0057] The data processing module is used to convert the collected blackboard writing video and computer display video from the physical signal format into the picture media stream format;

[0058] The data processing module is used to intercept video frames of the video in the format of a picture media stream at intervals of three seconds, and after screening the intercepted video frames, push them to the courseware generation module.

[0059] Further, the data processing module screens the intercepted video frames and then pushes them to the courseware generation module specifically as follows:

[0060] The data processing module is used to push the first video frame intercepted from the picture to the courseware generation module, and start calculating the similarity between adjacent video frames from the second video frame. If the similarity is less than the preset threshold, it means the video frame is in a changing state. If the similarity is greater than the preset threshold, it means the video frame is in a stable state;

[0061] In the case where the video frame is in a changing state, the video frame is not pushed to the courseware generation module. Among them, if the video frame is in a changing state for ten consecutive times, a timeout push is triggered;

[0062] In the case where the video frame is in a stable state for three consecutive times, the video frame is pushed to the courseware generation module.

[0063] It should be noted that the data processing module intercepts video frames of the video in the format of a picture media stream (blackboard writing video, computer display video) at intervals of 3 seconds. The first intercepted video frame is fixedly pushed to the platform as the starting page. Starting from the second frame, the picture is converted into YUV components for similarity calculation. When the similarity is less than the preset threshold, the video frame is in a changing state, and when it is greater than the threshold, the video frame is in a stable state; in the case where the video frame is in a changing state, the video frame is not pushed to the courseware generation module until the video frame is in a stable state for three consecutive times. At this time, the video frame is marked with course time information and pushed to the courseware generation module; in addition, if the video frame is in a changing state for ten consecutive times (in a changing state for 30 seconds), a timeout push will be triggered to prevent the video frame from not being pushed to the courseware generation module for too long.

[0064] Even further, the courseware generation module is used to call a preset digital human image to generate a digital human courseware according to the video frames of the blackboard writing video and the computer display video that have been intercepted and screened respectively, as well as the teacher's close-up video and teaching audio. In other words, Figure 2 is the schematic diagram of the interface of the digital human courseware according to the embodiment of the present application, as Figure 2As shown, the playback interface of the generated digital human courseware includes the video image after cropping and screening of the blackboard writing video (played in slide mode following the course timeline), the video image after cropping and screening of the computer display video (played in slide mode following the course timeline), the text converted from the lecture audio, and the digital human teaching video. It can be seen that compared with the current mainstream video playback form, the mode of playing key frames in slide show of this courseware resource will greatly reduce the consumption of storage space and the dependence on network bandwidth, and has extremely high economic utilization value. Optionally, any one of the three windows of the computer screen, the blackboard screen, and the lecture text also supports being switched to the main window (or enlarged full screen).

[0065] Through the data processing module and the courseware generation module in the embodiments of the present application, when a student opens the digital human courseware for viewing, the computer screen and the blackboard screen (the computer display video image and the blackboard writing video image after image cropping and screening) will be automatically played in slide show in sequence according to the course timeline, and at the same time, the lecture audio and the digital human teaching video (generated based on the teacher's close-up video and the lecture audio) will be played. By changing the computer display video and the blackboard writing video from video format to a slide mode of displaying pictures frame by frame, compared with traditional video resources, it greatly reduces the storage space required for courseware resources and can retain the teaching content during classroom teaching to a great extent.

[0066] Embodiment 3

[0067] In some specific embodiments, based on the above Embodiment 1, this specific embodiment provides a digital human courseware automatic generation system, which includes:

[0068] A data processing module, which is further used to perform frontal face detection on the teacher's close-up video through a face detection algorithm to obtain the frontal face image of the teacher appearing in the teacher's close-up video, and push the frontal face image of the teacher to the courseware generation module.

[0069] A courseware generation module, which is further used to perform face recognition on the frontal face image of the teacher appearing in the teacher's close-up video. If the teacher passes the face recognition, the teacher is the target teacher.

[0070] It should be noted by way of example that Figure 3 is a schematic flowchart of intercepting the frontal face image of the teacher according to the embodiments of the present application, as Figure 3As shown in the figure, the close-up video of the teacher is intercepted at 3-second intervals, and the face recognition algorithm is used to determine whether the face of the current teacher is in the frontal face state. The recognized frontal face image of the teacher is pushed to the courseware generation module; the courseware generation module performs face recognition analysis and returns whether it matches the teacher identity data on the platform successfully. If the match is successful, the data processing module will no longer intercept and push the frontal face image of the teacher. If the match is unsuccessful, continuous sampling will be performed until the match is successful or the course ends.

[0071] Furthermore, the courseware generation module is used to determine the courseware timeline of the digital human courseware according to the duration of the lecture audio, and convert the lecture audio into lecture text; then call the preset digital human image of the target teacher to generate a digital human teaching video with the same duration as the courseware timeline, the same behavior as the teacher close-up video, and the same lip shape as the lecture text; finally, generate the digital human courseware of the target teacher according to the blackboard writing video, computer display video, lecture audio, computer playback audio, and digital human teaching video. It should be emphasized that the digital human courseware is automatically recorded and generated according to the class schedule information of each classroom at the start time of the class. The whole process does not require manual intervention and is non-intrusive and automated.

[0072] For example, the courseware generation module first determines the playback duration of this courseware according to the duration of the lecture audio to obtain the courseware timeline, and at the same time converts the lecture audio into lecture text through the speech recognition algorithm; then automatically arranges the processed computer display video and blackboard writing video on the courseware timeline according to the course time information; then according to the result of face recognition, automatically retrieve the digital human image of the target teacher from the digital human resource library, and generate a digital human teaching video with the same duration as the courseware timeline, the same behavior as the teacher close-up video, and the same lip shape as the lecture text according to the teacher close-up video and lecture text, and overlay this digital human teaching video in the lower right corner of the courseware screen. When the student opens the digital human courseware to watch, the platform automatically performs a slide show of the computer display video and blackboard writing video in sequence according to the courseware timeline, and at the same time plays the teaching audio and digital human teaching video.

[0073] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combination form.

[0074] Embodiment 4

[0075] In one embodiment, Figure 4 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application, as Figure 4As shown, an electronic device is provided. The electronic device may be a server, and its internal structure diagram may be as shown in Figure 4 . The electronic device includes a processor, a network interface, an internal memory, and a non-volatile memory connected by an internal bus. Among them, the non-volatile memory stores an operating system, a computer program, and a database. The processor is used to provide computing and control capabilities. The network interface is used to communicate with external terminals through a network connection. The internal memory is used to provide an environment for the operation of the operating system and the computer program. When the computer program is executed by the processor, the data processing steps in the above embodiments are implemented. The database is used to store data.

[0076] Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0077] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0078] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, all possible combinations of the technical features in the above embodiments are not described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0079] The embodiments described above merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A digital human courseware automatic generation system, characterized in that The system includes an image acquisition module, an audio acquisition module, a data processing module, and a courseware generation module; The image acquisition module is used to acquire the teacher's close-up video, the blackboard writing video, and the computer display video during the teacher's teaching process; The audio acquisition module is used to acquire the teaching audio during the teacher's teaching process; The data processing module is used to push the teacher's close-up video, the blackboard writing video, the computer display video, and the teaching audio to the courseware generation module; The courseware generation module is used to generate a digital human courseware by calling a preset digital human image according to the teacher's close-up video, the blackboard writing video, the computer display video, and the teaching audio.

2. The system according to claim 1, wherein The image acquisition module includes a first image acquisition module and a second image acquisition module; The first image acquisition module is used to obtain the target area where the teacher is located through a target detection algorithm during the teacher's teaching process, acquire the teacher's close-up video within the target area, and acquire the blackboard writing video within a preset area; The second image acquisition module is used to acquire the computer display video during the teacher's teaching process.

3. The system according to claim 2, wherein The system includes: The data processing module is used to convert the acquired blackboard writing video and computer display video from the physical signal format into a video stream format of the picture media; The data processing module is used to intercept the video pictures of the video stream format of the picture media at intervals of three seconds respectively, and push the intercepted video pictures to the courseware generation module after screening.

4. The system according to claim 3, wherein The data processing module is used to push the first video picture intercepted from the picture to the courseware generation module, and calculate the similarity between adjacent video pictures starting from the second video picture. If the similarity is less than a preset threshold, it means that the video picture is in a changing state. If the similarity is greater than the preset threshold, it means that the video picture is in a stable state; In the case where the video picture is in a changing state, the video picture is not pushed to the courseware generation module. Among them, if the video picture is in a changing state for ten consecutive times, a timeout push is triggered; In the case where the video picture is in a stable state for three consecutive times, the video picture is pushed to the courseware generation module.

5. The system according to claim 4, wherein The courseware generation module is used to generate a digital human courseware by calling a preset digital human image according to the video pictures of the blackboard writing video and the computer display video after being intercepted and screened respectively, as well as the teacher's close-up video and the teaching audio.

6. The system according to claim 1, wherein The system includes: The audio acquisition module includes a first audio acquisition module and a second audio acquisition module; The first audio acquisition module is used to acquire the teacher's lecture audio during the teacher's teaching process; The second audio acquisition module is used to acquire the computer playback audio during the teacher's teaching process.

7. The system according to claim 6, characterized in that, The data processing module is further used to perform a frontal face detection on the teacher's close-up video through a face detection algorithm, obtain the frontal face picture of the teacher appearing in the teacher's close-up video, and push the frontal face picture of the teacher to the courseware generation module.

8. The system according to claim 7, wherein The courseware generation module is further configured to perform face recognition on the frontal face image of the teacher that appears in the teacher close-up video. If the teacher passes the face recognition, the teacher is the target teacher.

9. The system according to claim 8, wherein The courseware generation module is configured to determine the courseware timeline of the digital human courseware according to the duration of the lecture audio, and convert the lecture audio into lecture text. The courseware generation module is configured to call the preset digital human image of the target teacher to generate a digital human teaching video with a duration consistent with the courseware timeline, behavior consistent with the teacher close-up video, and lip shape consistent with the lecture text. The courseware generation module is configured to generate the digital human courseware of the target teacher according to the blackboard writing video, the computer display video, the lecture audio, the computer playback audio, and the digital human teaching video.

10. The system according to claim 1, wherein The system includes: The data processing module is configured to convert the collected teacher close-up video from the physical signal format into the picture media stream format, and then push it to the courseware generation module. The data processing module is configured to convert the collected lecture audio from the physical signal format into the sound media stream format, and then push it to the courseware generation module.