Display apparatus, video generation method, storage medium, and program product

By acquiring the speaker's voice and image signals through display devices, and using generative technology to generate a speaker video that is then integrated with the video of the teaching materials, the problem of high cost and low efficiency in generating high-quality courses has been solved, achieving cost reduction and efficiency improvement.

CN120825554APending Publication Date: 2025-10-21HISENSE COMML DISPLAY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410447717.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-15
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

The current technology for generating high-quality courses is costly and inefficient, mainly because it requires dedicated cameras and classrooms to record live teachers and integrate them with slide videos.

Method used

The speaker's voice and image signals are obtained through the display device, and the speaker video is generated using generative technology. It is then fused with the explanation material video to generate the target video.

Benefits of technology

It reduces the cost of generating high-quality courses, improves generation efficiency, and reduces reliance on dedicated hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120825554A_ABST
    Figure CN120825554A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the display technology, and provides a display device, a video generation method, a storage medium and a program product, the display device comprises a display screen and a processor, and the processor is configured to obtain a voice signal for explaining an explaining material by an explainer and a frame of image signal of the explainer; obtaining an explainer video according to the voice signal and the frame of image signal, and obtaining an explanation material video according to the explanation material; the length of the explainer video, the length of the voice signal and the length of the explaining material video are consistent; an explainer in the explainer video is consistent with an explainer in one frame of image signal; the mouth shape of the explainer in the explainer video is consistent with the voice in the voice signal; and fusing the explainer video and the explaining material video to generate a target video. By processing the sound signal and the frame of image signal, the quality and efficiency of target video generation can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of display technology, and more specifically, to a display device, a video generation method, a storage medium, and a program product. Background Art

[0002] Quality courses are an important way to demonstrate educational quality and disseminate educational resources. Quality courses are typically presented in a dual-screen format: a courseware screen displays the corresponding content (e.g., slides), while the lecturer screen explains the content displayed on the courseware screen.

[0003] In related technologies, the production of high-quality courses involves using dedicated cameras and dedicated classrooms to record videos of real teachers giving lectures in the classrooms. The pre-produced slideshow video is then combined with the real teacher's lecture video to create a high-quality course. This method of producing high-quality courses is costly and inefficient. Summary of the Invention

[0004] The exemplary embodiments of the present application provide a display device, a video generation method, a storage medium, and a program product, which can effectively reduce the cost of generating high-quality courses and improve the efficiency of generating high-quality courses.

[0005] In a first aspect, the present application provides a display device, comprising: a display screen and a processor, wherein the processor is configured to:

[0006] Acquire a voice signal of a lecturer explaining the lecture material, and a frame of image signal of the lecturer;

[0007] Acquire a video of a lecturer based on the voice signal and the one-frame image signal, and acquire a video of a lecture material based on the lecture material; the length of the lecturer video, the length of the voice signal, and the length of the lecture material video are consistent; the lecturer in the lecturer video is consistent with the lecturer in the one-frame image signal; and the lip shape of the lecturer in the lecturer video is consistent with the voice in the voice signal;

[0008] The lecturer video and the lecture material video are fused to generate a target video.

[0009] In some embodiments, the processor is configured to:

[0010] Obtaining text information of the explanation material and the voiceprint information of the explainer,

[0011] Performing speech synthesis on the text information and the voiceprint information to obtain the speech signal;

[0012] Alternatively, the processor is configured to:

[0013] Acquire a voice signal of the lecturer giving a real-life explanation of the lecture material.

[0014] In some embodiments, the processor is configured to:

[0015] Using generative technology to process the voice signal and the one frame of image signal to generate the narrator video;

[0016] or,

[0017] The voice signal and the one-frame image signal are sent to a server, and the lecturer video sent by the server is received; the lecturer video is generated by the server using a generative technology to process the voice signal and the one-frame image signal.

[0018] In some embodiments, the processor is configured to:

[0019] Acquire indication information; the indication information is used to indicate the posture of the speaker in the speaker video and / or the tone of the speaker;

[0020] The lecturer video is generated according to the instruction information, the voice signal and the one frame image signal.

[0021] In some embodiments, the processor is configured to:

[0022] determining target content in the explanation material video according to one or more of the explanation material, the voice signal, and the explainer video;

[0023] The target content is marked in the explanation material video.

[0024] In some embodiments, the processor is configured to:

[0025] The lecturer video and the lecture material video are merged according to a preset format to generate a target video; the preset format includes one or more of a typesetting format, an interface design format, and a video format.

[0026] In some embodiments, the processor is configured to:

[0027] displaying the target video;

[0028] Obtaining evaluation information of the target video; the evaluation information is used to indicate problems existing in the target video;

[0029] The target video is optimized according to problems existing in the target video.

[0030] In a second aspect, an embodiment of the present application provides a video generation method, comprising:

[0031] Acquire a voice signal of a lecturer explaining the lecture material, and a frame of image signal of the lecturer;

[0032] Generate a video of a lecturer based on the voice signal and the one-frame image signal, and generate a video of a lecture material based on the lecture material; the length of the lecturer video, the length of the voice signal, and the length of the lecture material video are consistent; the lecturer in the lecturer video is consistent with the lecturer in the one-frame image signal; the lip shape of the lecturer in the lecturer video is consistent with the voice in the voice signal;

[0033] The lecturer video and the lecture material video are fused to generate a target video.

[0034] In a third aspect, an embodiment of the present application provides a video generation device, comprising:

[0035] An acquisition module, configured to acquire a voice signal of a lecturer explaining the lecture material, and a frame of image signal of the lecturer;

[0036] a processing module configured to generate a video of a lecturer based on the voice signal and the one-frame image signal, and to generate a video of a lecture material based on the lecture material; wherein the length of the lecturer video, the length of the voice signal, and the length of the lecture material video are consistent; the lecturer in the lecturer video is consistent with the lecturer in the one-frame image signal; and the lip shape of the lecturer in the lecturer video is consistent with the voice in the voice signal;

[0037] The fusion module is used to fuse the lecturer video and the lecture material video to generate a target video.

[0038] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the second aspect.

[0039] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the method described in the second aspect when executed by a processor.

[0040] In a sixth aspect, an embodiment of the present application provides a chip, the chip including a processor, the processor being used to call a computer program in a memory to execute the method described in the second aspect.

[0041] The display device, video generation method, storage medium, and program product provided by the embodiments of the present application include: a display screen and a processor, wherein the processor is configured to: obtain a voice signal of a lecturer explaining a lecture material, and a one-frame image signal of the lecturer; obtain a lecturer video based on the voice signal and the one-frame image signal, and obtain a lecture material video based on the lecture material; the length of the lecturer video, the length of the voice signal, and the length of the lecture material video are consistent; the lecturer in the lecturer video is consistent with the lecturer in the one-frame image signal; the lecturer's lip shape in the lecturer video is consistent with the voice in the voice signal; and the lecturer video and the lecture material video are fused to generate a target video. By processing the sound signal and the one-frame image signal, the quality and efficiency of target video generation can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the implementation methods in the embodiments of the present application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0043] Figure 1 A schematic diagram of an existing display scene;

[0044] Figure 2 A schematic structural diagram of a display device provided in an embodiment of the present application;

[0045] Figure 3 A schematic diagram of a video generation method provided in an embodiment of the present application Figure 1 ;

[0046] Figure 4 A schematic diagram of a video generation process provided in an embodiment of the present application;

[0047] Figure 5 A schematic diagram of a video generation method provided in an embodiment of the present application Figure 2 ;

[0048] Figure 6 A schematic diagram of the stages of a video generation process provided in an embodiment of the present application;

[0049] Figure 7 A schematic diagram of the structure of a video generation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] In order to make the purpose, implementation mode and advantages of the present application clearer, the exemplary implementation mode of the present application will be clearly and completely described below in conjunction with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only part of the embodiments of the present application, not all of the embodiments.

[0051] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.

[0052] In addition, the terms "comprises" and "comprising" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to those components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0053] Quality courses are an important form of educational quality assessment, educational level display, and educational resource dissemination. The country vigorously develops the construction of quality courses, and schools attach importance to the production of quality courses. Quality courses have also become an important indicator for evaluating teachers' teaching quality.

[0054] The presentation format of high-quality courses is generally in the form of a dual screen of courseware + lecturer on the display device. The courseware screen displays the corresponding content (for example, slides), and the lecturer screen explains the content displayed on the courseware screen.

[0055] Figure 1 This is a possible schematic diagram of displaying high-quality courses provided in the embodiment of the present application. Figure 1 As shown, the user can control the display device 200 through the smart device 300 or the control device 100, so that the display device 200 executes the display of high-quality courses according to the user's instructions.

[0056] In some embodiments, the control device 100 may be a remote controller. Communication between the remote controller and the display device may include infrared protocol communication, Bluetooth protocol communication, or other short-range communication methods, and the display device 200 may be controlled wirelessly or wired. The user may control the display device 200 by inputting user commands through buttons on the remote controller, voice input, control panel input, and the like.

[0057] In some embodiments, the display device 200 can also be controlled using a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.). For example, the display device 200 can be controlled using an application running on the smart device. In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, the display device can receive user control through touch or gestures. For another example, the user's voice command control can be directly received through a module for obtaining voice commands configured within the display device 200, or the user's voice command control can be received through a voice control device provided external to the display device 200.

[0058] In some embodiments, the display device 200 also communicates data with the server 400. The display device 200 may be connected to a local area network (LAN), a wireless local area network (WLAN), or other networks. The server 400 may provide various content and interactions to the display device 200. The server 400 may be a single cluster or multiple clusters, and may include one or more types of servers.

[0059] Among them, the display device provided in the embodiment of the present application can have various implementation forms, for example, it can be a television, a smart TV, a laser projection device, a monitor, an electronic bulletin board, an electronic table, etc., and the embodiment of the present application does not limit this.

[0060] The following is an introduction to the hardware configuration of the display devices mentioned above.

[0061] Figure 2 This is a possible hardware configuration diagram of a display device provided in this application. Figure 2 As shown, in some embodiments, the display device 200 may include at least one of a tuner and demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a power supply 280, a memory 290, and a user interface 2100.

[0062] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, RAM, ROM, and first to nth interfaces for input / output.

[0063] The display 260 includes a display screen component for presenting images, and a driving component for driving image display, a component for receiving image signals output from a controller, and a component for displaying video content, image content, and a menu control interface and a user control UI interface.

[0064] The display 260 may be a liquid crystal display, an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLed, a Micro-OLED, a quantum dot light-emitting diode (QLED), or a projection display, etc. It may also be a projection device and a projection screen.

[0065] Communicator 220 is a component used to communicate with external devices or servers using various communication protocols. For example, the communicator may include at least one of a WiFi module, a Bluetooth module, a wired Ethernet module, or other network communication protocol chip, a near-field communication protocol chip, and an infrared receiver. Display device 200 can use communicator 220 to send and receive control signals and data signals with external control device 100 or server 400.

[0066] The user interface 2100 may be configured to receive a control signal from the control device 100 (eg, an infrared remote controller, etc.).

[0067] Detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 includes a light receiver, a sensor for collecting ambient light intensity; or detector 230 includes an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or detector 230 includes a sound collector, such as a microphone, for receiving external sounds.

[0068] The external device interface 240 may include, but is not limited to, any one or more of the following: a high-definition multimedia interface (HDMI), an analog or digital high-definition component input interface (component), a composite video input interface (CVBS), a USB input interface (USB), an RGB port, etc. It may also be a composite input / output interface formed by multiple of the above interfaces.

[0069] The tuner-demodulator 210 receives broadcast television signals via a wired or wireless reception method, and demodulates audio and video signals, such as EPG data signals, from a plurality of wireless or wired broadcast television signals.

[0070] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0071] Controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory 290. Controller 250 controls the overall operation of display device 200. For example, in response to receiving a user command to select a UI object for display on display 260, controller 250 may perform operations related to the object selected by the user command.

[0072] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM Random Access Memory (RAM), ROM (Read-Only Memory, ROM), a first interface to an nth interface for input / output, a communication bus (Bus), etc.

[0073] The user may input a user command through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input command through the graphical user interface (GUI). Alternatively, the user may input a user command through a specific voice or gesture, and the user input interface may recognize the voice or gesture through a sensor to receive the user input command.

[0074] A user interface is the medium for interaction and information exchange between an application or operating system and the user. It converts information between its internal form and a user-friendly format. A common user interface is the graphical user interface (GUI), which refers to a graphical user interface related to computer operations. It can be an icon, window, control, or other interface element displayed on an electronic device's display. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0075] In related technologies, the production of high-quality courses involves using dedicated cameras and dedicated classrooms to record videos of real teachers giving lectures in the classrooms. The pre-produced slideshow video is then combined with the real teacher's lecture video to create a high-quality course. This method of producing high-quality courses is costly and inefficient.

[0076] In view of this, the embodiments of the present application provide a display device, a video generation method, a storage medium and a program product, which can effectively reduce the cost of producing high-quality courses and improve the efficiency of high-quality courses.

[0077] The following detailed description of the technical solution of the present application is provided in conjunction with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0078] Figure 3 Schematic diagram of the process of the video generation method provided in the embodiment of the present application Figure 1 ,like Figure 3 As shown, the steps include:

[0079] S301: Acquire a voice signal of a lecturer explaining a lecture material and a frame of image signal of the lecturer.

[0080] The execution subject of the embodiment of the present application may be a display device, a processor of the display device, a control chip of the display device, a video generation system deployed in the display device, etc., and the embodiment of the present application is not limited thereto. The following description will take the processor of the display device as an example.

[0081] In some embodiments, the display device may include an image acquisition device (eg, a camera) and an audio acquisition device (eg, a microphone).

[0082] In some embodiments, the display device may further include a text capture device (eg, a keyboard).

[0083] In some embodiments, the lecturer can be the protagonist of the generated video, for example, a teacher who lectures on a high-quality course. The lecture materials can be the materials that need to be lectured, for example, a PPT corresponding to the high-quality course.

[0084] For example, Figure 4 As shown, the processor can obtain a frame of image signal of the lecturer through the image acquisition device, and obtain the voice signal of the lecturer explaining the explanation material through the audio acquisition device or the text acquisition device.

[0085] Exemplarily, the processor may take a picture of the lecturer through an image acquisition device to obtain a frame of image signal of the lecturer, or may record a portrait video of the lecturer through an image acquisition device, extract a frame of picture from the portrait video, and obtain a frame of image signal of the lecturer.

[0086] The lecturer can provide a live explanation of the explanatory material, and the processor can capture the lecturer's voice through an audio acquisition device to obtain the voice signal. Alternatively, the processor can capture text information of the lecture material through a text acquisition device and convert the text information into a voice signal of the lecturer's explanation of the lecture material. For example, when the processor acquires the text information, it can also acquire the lecturer's voiceprint information and use speech synthesis technology to process the text and voiceprint information to obtain the voice signal. The voiceprint information can be pre-recorded, or the lecturer's voiceprint information can be captured through the audio acquisition device when the text information is acquired. The speech synthesis technology can use Text to Speech (TTS) technology.

[0087] In some embodiments, the display device may further include a peripheral interface (e.g., a USB interface), and the processor may obtain an externally inputted picture of the lecturer or a video of the lecturer's portrait through the peripheral interface to obtain a frame of image signal of the lecturer. The display device may also obtain text through a touch device, or transcribe speech into text through a voice input device, although this embodiment of the application is not limited thereto.

[0088] S302: Acquire a video of the lecturer according to the voice signal and the one frame of image signal, and acquire a video of the lecture material according to the lecture material.

[0089] In some embodiments, when the processor obtains the voice signal and the one-frame image signal, it can generate a narrator video based on the voice signal and the one-frame image signal.

[0090] In a possible implementation, the processor may use generative technology (eg, artificial intelligence generated content (AIGC) technology) to process the voice signal and the one-frame image signal to generate the narrator video.

[0091] For example, the processor can call a pre-trained video generation model built based on AIGC, input the voice signal and the one-frame image signal into the video generation model, and obtain the narrator video output by the model. The video generation model can be trained based on the ChatGPT model, Sadtalker model, EMO model, styletalk model, etc. The process of training the video generation model can refer to the implementation method in the prior art, and the embodiments of the present application will not be described in detail.

[0092] The length of the lecturer video is consistent with the length of the voice signal; the lecturer in the lecturer video is consistent with the lecturer in the one-frame image signal; and the lecturer's lip shape in the lecturer video is consistent with the voice in the voice signal. For example, the voice signal can be the teacher's voice explaining the material, and the image signal can be a selfie of the teacher. The processor inputs the voice and photo and uses the EMO model method to generate a portrait lecture video (lecturer video) based on the teacher's selfie. The lip shape of the teacher's portrait in the video is aligned with the input voice, realizing a "real-life portrait lecture video."

[0093] For another possible implementation, please refer to Figure 4 The processor can send the voice signal and the one-frame image signal to a server (e.g., a cloud server). Upon receiving the voice signal and the one-frame image signal, the server can use generative technology to process the voice signal and the one-frame image signal to generate the narrator video. The server then sends the generated narrator video to the processor of the display device. In this way, the narrator video can be quickly generated even when the display device has poor performance.

[0094] In some embodiments, the processor can generate a video of the explanation material based on pre-input explanation materials. For example, the processor can generate the video of the explanation material by copying an image, recording a desktop, or the like. The length of the explanation material video is consistent with the length of the video of the speaker, and the content of the explanation material video matches the content of the video of the speaker. For example, if the explanation material is a PowerPoint presentation, and the video of the speaker is explaining the content of page M of the PowerPoint presentation, the video of the explanation material displays the content of page M.

[0095] In some embodiments, after acquiring the explainer video, the processor may generate an explanation material video based on the length and content of the explainer video.

[0096] In some embodiments, the processor may send the explanation material to the server, so that the server generates an explanation material video that matches the explainer video.

[0097] S303: Fusing the lecturer video and the lecture material video to generate a target video.

[0098] In an embodiment of the present application, when the processor obtains the lecturer video and the lecture material video, it can use video editing to merge the lecturer video and the lecture material video to generate a target video (for example, a high-quality course).

[0099] In some embodiments, the fusion of the lecturer video and the lecture material video may also be performed by the server, and the server sends the generated target video to the processor of the display device.

[0100] In some embodiments, after the processor generates the target video, the target video may be displayed on a display screen of a display device, or the target video may be sent or stored to a preset address.

[0101] The video generation method provided in the embodiment of the present application obtains a voice signal of a lecturer explaining a lecture material, and a frame of image signal of the lecturer; obtains a lecturer video based on the voice signal and the frame of image signal, and obtains a lecture material video based on the lecture material; the length of the lecturer video, the length of the voice signal, and the length of the lecture material video are consistent; the lecturer in the lecturer video is consistent with the lecturer in the frame of image signal; the lip shape of the lecturer in the lecturer video is consistent with the voice in the voice signal; and the lecturer video and the lecture material video are fused to generate a target video. The target video can be generated by processing the voice signal and the frame of image signal using generative technology, which reduces the dependence on dedicated hardware in the process of generating the target video, thereby reducing the cost of generating the target video and improving the efficiency of generating the target video.

[0102] Based on the above embodiments, Figure 5 The video generation method provided in the embodiment of the present application is further explained.

[0103] Figure 5 Schematic diagram of the process of the video generation method provided in the embodiment of the present application Figure 2 ,like Figure 5 As shown, the steps include:

[0104] S501: Acquire a voice signal of a lecturer explaining a lecture material and a frame of image signal of the lecturer.

[0105] The implementation of S501 in the embodiment of the present application is similar to Figure 3 The implementation of S301 in the illustrated embodiment is similar and will not be described again here.

[0106] In some embodiments, when using speech synthesis technology to convert text information into voice signals, the processor can also select voiceprint parameters such as speech speed and intonation based on user instructions, thereby providing voice signals of various styles. It should be understood that if the processor does not receive the speech speed, intonation, or other parameters selected for the voiceprint, it can select the corresponding default parameters.

[0107] S502: Acquire indication information; the indication information is used to indicate the posture of the speaker in the speaker video and / or the tone of the speaker.

[0108] In the embodiments of this application, the speaker's posture may refer to the speaker's movements and speed in the speaker's video. For example, the speaker's head position, head shaking direction, and speed; the speaker's body position, body tilt, and speed; the speaker's hand gesture position, gesture type, and speed. The speaker's voice tone may refer to the configuration and changes in the speaker's voice's rhythm, pitch, heaviness, slowness, and speed during the speaker's presentation.

[0109] In some embodiments, the processor can obtain the instruction information through an external interface of the display device (e.g., a user input interface). By limiting the posture and / or tone of the speaker in the speaker video through the instruction information, the generated speaker video can be made more realistic and natural.

[0110] In some embodiments, when the voice signal of the lecturer explaining the explanation material is based on the lecturer's real voice, the voice signal already includes the lecturer's tone, and the indication information may only include the lecturer's posture.

[0111] It should be understood that S502 is an optional step.

[0112] S503: Generate the lecturer video according to the instruction information, the voice signal and the one frame of image signal.

[0113] In an embodiment of the present application, the processor may input the instruction information, the voice signal, and the one-frame image signal into a pre-trained video generation model based on AIGC to obtain a narrator video output by the video generation model. The instruction information may serve as a restriction condition in the process of generating the narrator video.

[0114] In some embodiments, if the instruction information is not obtained, the processor may generate the narrator video based on the voice signal and the one frame image signal.

[0115] In some embodiments, the processor may send the instruction information, the voice signal, and the one-frame image signal to the server, and the server generates the narrator video according to the received instruction information, the voice signal, and the one-frame image signal.

[0116] In some embodiments, when the processor generates the explainer video based on the voice signal and the one-frame image signal, if the processor receives the instruction information, the processor may regenerate the explainer video based on the instruction information, the voice signal, and the one-frame image signal. Alternatively, the processor may adjust the generated explainer video based on the instruction information to obtain the desired explainer video.

[0117] In some embodiments, if a server is used to generate a commentator video, and if the instruction information, the voice signal, and the one-frame image signal are not transmitted to the server simultaneously, upon receiving the instruction information, if the server has already generated the commentator video, the server may regenerate the commentator video based on the instruction information or adjust the already generated commentator video. If the server has not yet generated the commentator video, the commentator video may be generated based on the instruction information, the voice signal, and the one-frame image signal.

[0118] S504: Obtain a video of the explanation material according to the explanation material.

[0119] The implementation of S504 in the embodiment of this application is similar to Figure 3 The implementation of S302 in the illustrated embodiment is similar and will not be described again here.

[0120] In some embodiments, when generating the explanation video, the processor may also obtain style information of the explanation video based on user instructions, thereby generating the explanation video in the style desired by the user. The user instructions may be obtained via a text capture device, a voice capture device, or a user input interface of a display device, and this embodiment of the application is not limited thereto.

[0121] It should be understood that the length of the generated explainer video, the length of the voice signal, and the length of the explanation material video are consistent; the explainer in the explainer video is consistent with the explainer in the one frame image signal; the explainer's lip shape in the explainer video is consistent with the voice in the voice signal.

[0122] S505: Determine target content in the explanation material video according to one or more of the explanation material, the voice signal, and the explainer video.

[0123] In the embodiment of the present application, the target content may refer to some key contents in the explanation material video, or content that needs to be emphasized or paid special attention to.

[0124] In some embodiments, the target content may be marked in the explanation material, for example, the target content may be marked in bold or red, etc. The processor may perform content recognition on the explanation material to determine the target content.

[0125] In some embodiments, the voice signal may carry a description of the target content. For example, the content corresponding to the voice signal is "X is the key content of this part." The processor may recognize the voice signal and determine the target content.

[0126] In some embodiments, the intonation in the speech signal can also represent the target content. For example, for a part of the content, if there are more emphasis or more pauses when explaining the part of the content in the speech signal, the part of the content can be identified as the target content. Alternatively, if a part of the content is mentioned more times in the speech signal, the part of the content can be used as the target content.

[0127] In some embodiments, the speaker's voice and movements in the explainer video can also represent the target content. For example, if the explainer emphasizes a certain part of the content more often or pauses more often when explaining the part, the part can be identified as the target content. Alternatively, if the explainer mentions a certain part of the content more often, the part can be identified as the target content. Alternatively, if the explainer makes large movements or makes many movements when explaining the part, the part can be identified as the target content. The processor can identify the explainer video and determine the target content.

[0128] S506: Mark the target content in the explanation material video.

[0129] In some embodiments, when the processor acquires the target content, it can mark the target content in the explanation material video. For example, the target content can be marked in red, crossed out, annotated, circled, etc. The embodiments of this application do not limit the method of marking the target content. By marking the target content, the content of the explainer video and the material video can be synchronized and interacted, thereby improving the quality of the generated video.

[0130] It should be understood that S505 and S506 are optional steps.

[0131] S507: Fusing the lecturer video and the lecture material video according to a preset format to generate a target video.

[0132] In an embodiment of the present application, the preset format is a format that the target video needs to meet, and the preset format includes one or more of a typesetting format, an interface design format, and a video format.

[0133] In some embodiments, the preset format may be pre-defined in the processor, or acquired based on a text acquisition device, or a voice acquisition device, or a user input interface of a display device, and the embodiments of the present application do not limit this.

[0134] The processor may use a fusion algorithm to fuse the lecturer video and the lecture material video to generate a target video in a preset format. The fusion algorithm may be a gradient fusion algorithm, a video splicing algorithm, etc., which is not limited in the present embodiment.

[0135] In some embodiments, since the explanation material may consist of multiple sections and the logic of different sections may be inconsistent, a lecturer video and a material video may be generated for each section separately, and then all the videos are fused to generate a target video. For example, a voice signal is generated for each section separately, and the processor processes each voice signal and a frame of image signal to generate a lecturer video for that section. The processor also processes the material video for that section to generate a lecture material video for that section. The lecturer video and the lecture material video for that section are then fused to generate the target video corresponding to that section.

[0136] S508: Obtain evaluation information of the target video; the evaluation information is used to indicate problems existing in the target video.

[0137] In an embodiment of the present application, when a processor generates a target video, the target video may be displayed on a display screen. A user may view the generated target video and obtain user evaluation information of the target video through a user input interface of the display device. Upon obtaining the evaluation information, the processor may identify the evaluation information and determine any issues with the target video.

[0138] S509: Optimize the target video according to the problems existing in the target video.

[0139] In some embodiments, the evaluation information may be an evaluation of the target video as a whole, or an evaluation of a specific portion thereof, such as the style of the target video, the style of the teaching material video, or the posture of the speaker in the teaching video.

[0140] When the processor detects a problem in the target video, it can optimize the entire target video based on the problem, for example, by using a relevant optimization algorithm to optimize the target video. Alternatively, it can optimize the problematic portion separately. For example, if the speaker in the lecture video has a posture problem, the processor can optimize the lecture video, merge the optimized lecture video with the lecture material video, and input the optimized target video.

[0141] In summary, the video generation method provided in the embodiment of the present application includes three stages: the first stage is the generation of the lecturer video, the second stage is the generation of the lecture material video, and the third stage is the fusion process of the lecturer video and the lecture material video.

[0142] like Figure 6As shown in FIG, the first stage mainly includes a signal acquisition module, a speech synthesis module, an AIGC model, and a model control module.

[0143] Signal acquisition module: mainly collects text, voice and images, relying on Figure 6 The peripherals shown in the figure can be used, such as keyboard, microphone and camera, but text collection can also be input through touch or voice input can be transcribed into text. Voice input can use near-field microphones such as smart pens to achieve near-field sound pickup effects. Images can also be obtained by importing from other devices.

[0144] Text-to-speech (TTS) module: This module primarily implements text-to-speech conversion, generally using registered voiceprint information to convert text into a specific person's voice signal. This module is editable and can flexibly provide variable-style speech generation services by controlling the voiceprint type, speed, and intonation.

[0145] The AIGC model primarily utilizes generative technology for video generation. Based on a diffusion model, it inputs image and audio feature information, fuses them, and employs a continuous noise reduction process to produce a guide video that meets the requirements. The generated guide video includes image information, with the voice and lip movements of the characters aligned with the input audio. This is an editable module that can generate guide videos of different styles through pre-training.

[0146] Model Control Module: This module provides control over the model-generated videos. By inputting guidance information, this module can guide the generated narrator videos to achieve fixed movements and speeds, such as head position, head shake direction, and speed; body position, body tilt line, and speed; and gesture position, type, and speed. This is an editable module that controls the generated narrator videos by inputting these behavioral features. The control information features provided by this module can be jointly trained with the AIGC module to achieve model controllability.

[0147] The second stage mainly includes: explanatory material video generation module and alignment module.

[0148] Explanation Video Generation Module: This module generates corresponding explanation videos for the accompanying explanation materials. First, the length of the audio explanation signal for this part of the material is referenced, and the original material video of the corresponding length is obtained through methods such as image copying and desktop recording. This module is editable and can be used to adjust the style, duration, and other video information of the material video.

[0149] Alignment Module: This module primarily relies on and references the lecturer video to edit and optimize the explanation material video, enabling content interaction between the explanation material video and the lecturer video. Based on the content explained in the lecturer video of the same length, it annotates the content in the material video that needs to be emphasized, achieving synchronous interaction between the lecturer video and the explanation material video, and finally outputs an aligned explanation material video. This module is editable and can be used to emphasize and annotate the material content, such as highlighting, underlining, and circling.

[0150] The third stage mainly includes: video comprehensive module.

[0151] The video integration module primarily integrates the aligned explainer video and explanation material video to produce the target video. This module uses video editing to integrate the explainer video and explanation material video according to the required format, creating a composite video that meets the required formatting requirements. This is the final target video. This is an editable module that provides integrated video layouts in different typographic styles to enrich the target video's format.

[0152] Based on the above embodiments, an embodiment of the present application further provides a video generating device.

[0153] Figure 7 A schematic diagram of the structure of a video generating device 70 provided in an embodiment of the present application is shown in FIG. Figure 7 Shown, including:

[0154] The acquisition module 701 is used to acquire a voice signal of a lecturer explaining the explanation material, and a frame of image signal of the lecturer.

[0155] The processing module 702 is used to obtain a video of the lecturer based on the voice signal and the one-frame image signal, and to obtain a video of the explanation material based on the explanation material; the length of the lecturer video, the length of the voice signal, and the length of the explanation material video are consistent; the lecturer in the lecturer video is consistent with the lecturer in the one-frame image signal; and the lip shape of the lecturer in the lecturer video is consistent with the voice in the voice signal.

[0156] The fusion module 703 is used to fuse the lecturer video and the lecture material video to generate a target video.

[0157] In some embodiments, the acquisition module 701 is further configured to acquire text information explaining the explanation material and the voiceprint information of the explainer, and perform speech synthesis on the text information and the voiceprint information to obtain the speech signal;

[0158] In some embodiments, the acquisition module 701 is further configured to acquire a voice signal of the lecturer giving a real-life explanation of the explanation material.

[0159] In some embodiments, the processing module 702 is further configured to process the voice signal and the one frame image signal using a generative technology to generate the narrator video.

[0160] In some embodiments, the processing module 702 is further used to send the voice signal and the one-frame image signal to the server, and receive the lecturer video sent by the server; the lecturer video is generated by the server using generative technology to process the voice signal and the one-frame image signal.

[0161] In some embodiments, the processing module 702 is further used to obtain indication information; the indication information is used to indicate the posture of the speaker in the speaker video and / or the tone of the speaker; and the speaker video is generated based on the indication information, the voice signal and the one frame image signal.

[0162] In some embodiments, the processing module 702 is further configured to determine target content in the explanation material video based on one or more of the explanation material, the voice signal, and the explainer video; and mark the target content in the explanation material video.

[0163] In some embodiments, the fusion module 703 is further used to fuse the lecturer video and the lecture material video according to a preset format to generate a target video; the preset format includes one or more of a typesetting format, an interface design format, and a video format.

[0164] In some embodiments, the processing module 702 is further used to display the target video; obtain evaluation information of the target video; the evaluation information is used to indicate problems existing in the target video; and optimize the target video according to the problems existing in the target video.

[0165] The video generation device provided in this application is used to execute the technical solution of the video generation method provided in any of the aforementioned embodiments. Its implementation principle and technical effect are similar and will not be described in detail.

[0166] It should be noted that it should be understood that the division of the various modules of the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. Each module can be a separately established processing element, or it can be integrated into a chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called by a processing element of the above device to execute the functions of the above modules. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by the hardware integrated logic circuit in the processor element or software instructions.

[0167] The present application also provides a computer-readable storage medium, which may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes. Specifically, the computer-readable storage medium stores program instructions, and the program instructions are used for the methods in the above embodiments.

[0168] The present application also provides a program product, the program product including execution instructions stored in a readable storage medium. At least one control module of a display device can read the execution instructions from the readable storage medium, and at least one control module executes the execution instructions to cause the display device to implement the video generation method provided in the various embodiments described above.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

[0170] For ease of explanation, the above description has been presented in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations are possible. The above embodiments have been selected and described to better explain the principles and practical applications, thereby enabling those skilled in the art to better utilize the embodiments and various variations of the embodiments suitable for specific use considerations.

Claims

1. A display device, characterized in that: The display device includes: a display screen and a processor, wherein the processor is configured to: Acquire a voice signal of a lecturer explaining the lecture material, and a frame of image signal of the lecturer; Acquire a video of a lecturer based on the voice signal and the one-frame image signal, and acquire a video of a lecture material based on the lecture material; the length of the lecturer video, the length of the voice signal, and the length of the lecture material video are consistent; the lecturer in the lecturer video is consistent with the lecturer in the one-frame image signal; and the lip shape of the lecturer in the lecturer video is consistent with the voice in the voice signal; The lecturer video and the lecture material video are fused to generate a target video.

2. The display device according to claim 1, wherein The processor is configured to: Obtaining text information of the explanation material and the voiceprint information of the explainer, Performing speech synthesis on the text information and the voiceprint information to obtain the speech signal; Alternatively, the processor is configured to: Acquire a voice signal of the lecturer giving a real-life explanation of the lecture material.

3. The display device according to claim 2, wherein The processor is configured to: Using generative technology to process the voice signal and the one frame of image signal to generate the narrator video; or, The voice signal and the one-frame image signal are sent to a server, and the lecturer video sent by the server is received; the lecturer video is generated by the server using a generative technology to process the voice signal and the one-frame image signal.

4. The display device according to claim 3, wherein The processor is configured to: Acquire indication information; the indication information is used to indicate the posture of the speaker in the speaker video and / or the tone of the speaker; The lecturer video is generated according to the instruction information, the voice signal and the one frame image signal.

5. The display device according to any one of claims 1 to 4, characterized in that: The processor is configured to: determining target content in the explanation material video according to one or more of the explanation material, the voice signal, and the explainer video; The target content is marked in the explanation material video.

6. The display device according to claim 5, wherein: The processor is configured to: The lecturer video and the lecture material video are merged according to a preset format to generate a target video; the preset format includes one or more of a typesetting format, an interface design format, and a video format.

7. The display device according to claim 1 or 6, characterized in that: The processor is configured to: displaying the target video; Obtaining evaluation information of the target video; wherein the evaluation information is used to indicate problems existing in the target video; The target video is optimized according to problems existing in the target video.

8. A video generation method, characterized in that: include: Acquire a voice signal of a lecturer explaining the lecture material, and a frame of image signal of the lecturer; Generate a video of a lecturer based on the voice signal and the one frame of image signal, and generate a video of a lecture material based on the lecture material; the length of the lecturer video, the length of the voice signal, and the length of the lecture material video are consistent; The speaker in the speaker video is consistent with the speaker in the one frame of image signal; The lip shape of the speaker in the speaker video is consistent with the voice in the voice signal; The lecturer video and the lecture material video are fused to generate a target video.

9. A storage medium, characterized in that: The storage medium stores computer-executable instructions, which are used to implement the method according to claim 8 when executed by a processor.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to claim 8 when the computer program is executed by a processor.