Digital human video generation method and device, electronic equipment and storage medium
The method of generating digital human videos through large models solves the problems of high creation threshold and high material matching cost in the creation of medical popular science content, and realizes the efficient generation of high-quality popular science videos.
Patent Information
- Application Number
- CN202510653550.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies make it difficult to efficiently generate high-quality videos in the creation of medical popular science content, and there are problems with high creation thresholds and high material matching costs.
A method for generating digital human videos through a large model includes determining popular science demand information, generating prompt words, generating oral scripts and video storyboard scripts, obtaining target materials and digital human oral audio, and generating popular science videos in combination with video storyboard scripts.
It lowers the creation threshold, improves the efficiency and quality of video generation, ensures the accuracy and display effect of content, and reduces the cost of material matching.
Smart Images

Figure CN120640095A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, specifically to the fields of large models, artificial intelligence and content creation, and in particular to a method, device, electronic device and storage medium for generating digital human videos. Background Art
[0002] With the development of large-scale model, audio, image and video processing technologies, it has become possible to use artificial intelligence capabilities to assist content production and improve production efficiency. Faced with the urgent need to improve the efficiency of health science popularization content production, how to create high-quality medical content through the simplest input while solving the problems of content quality and production efficiency is an urgent problem to be solved. Summary of the Invention
[0003] The present disclosure provides a method, device, electronic device and storage medium for generating a digital human video.
[0004] According to one aspect of the present disclosure, a method for generating a digital human video is provided, the method comprising:
[0005] Determining science popularization demand information, and generating prompt words for a large model based on the science popularization demand information;
[0006] Generate a spoken script and a video storyboard script based on the prompt words through the large model;
[0007] According to the oral script, the target materials and the oral audio of the digital human required for popular science are obtained;
[0008] A popular science video of the digital human is generated based on the video storyboard script, the target material and the spoken audio of the digital human.
[0009] According to another aspect of the present disclosure, there is provided a digital human video generating apparatus, comprising:
[0010] A first acquisition module is used to determine science popularization demand information and generate prompt words of a large model according to the science popularization demand information;
[0011] A second acquisition module is used to generate a spoken script and a video storyboard script according to the prompt words using the large model;
[0012] The third acquisition module is used to obtain the target materials required for popular science and the digital human's spoken audio according to the spoken script;
[0013] A generation module is used to generate a popular science video of the digital human based on the video storyboard script, the target material and the spoken audio of the digital human.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the embodiment of the first aspect.
[0018] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the embodiment of the first aspect.
[0019] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the steps of the method described in the embodiment of the first aspect.
[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0022] Figure 1 is a schematic diagram of a digital human video generation method provided by an embodiment of the present disclosure;
[0023] Figure 2 is a schematic diagram of another method for generating a digital human video provided by an embodiment of the present disclosure;
[0024] Figure 3 is a schematic diagram of another method for generating a digital human video provided by an embodiment of the present disclosure;
[0025] Figure 4 is a schematic diagram of another method for generating a digital human video provided by an embodiment of the present disclosure;
[0026] Figure 5 This is a logic diagram of a digital human video generation system provided by an embodiment of the present disclosure;
[0027] Figure 6 This is a structural block diagram of a digital human video generation device provided by an embodiment of the present disclosure;
[0028] Figure 7 A schematic block diagram of an electronic device for implementing an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0029] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] Data processing is the collection, storage, retrieval, processing, transformation and transmission of data. Its basic purpose is to extract and derive valuable and meaningful data for certain specific people from a large amount of data that may be disorganized and difficult to understand.
[0031] A large model refers to a machine learning model with large-scale parameters and complex computing structure. It is usually built by a deep neural network and has billions or even hundreds of billions of parameters. The purpose is to improve the model's expressive power and predictive performance and to be able to handle more complex tasks and data.
[0032] Artificial Intelligence (AI) is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. Its purpose is to produce a new intelligent machine that can respond in a way similar to human intelligence.
[0033] Content creation is an activity that takes creativity as its core and systematically produces and disseminates information that meets the needs of the audience through media forms such as text, pictures, videos, and audio. It is a process of integrating, creating, and editing materials in various forms such as text, pictures, audio, and video to form information with dissemination value.
[0034] Figure 1 FIG. 1 is a schematic diagram of a method for generating a digital human video provided by an embodiment of the present disclosure. Figure 1 As shown, the method includes:
[0035] S101, determining science popularization demand information, and generating prompt words of a large model according to the science popularization demand information.
[0036] Optionally, this embodiment may be applicable to the generation of digital human videos for popular science, such as medical science, astronomy science, or technology science.
[0037] It is understandable that the science popularization demand information is the demand information required for generating science popularization videos; optionally, the demand information may include but is not limited to information such as video length, video content, and video theme.
[0038] The prompt words of the large model are generated according to the popular science demand information. The prompt words of the large model are used to guide the large model to understand user needs and output more accurate results. In this embodiment, the prompt words of the large model are generated according to the popular science demand information, so that the large model can generate content corresponding to the popular science demand information according to the prompt words.
[0039] In some embodiments, the popular science demand information can be spliced and filled into a preset prompt word template to obtain the prompt words of the large model; for example, the popular science demand information is: 120 seconds of video length, medical theme, and content introduction of brain tissue structure, then prompt words can be generated according to the popular science demand information: Please generate a medical video with a video length of 120 seconds according to the following requirements, and the video content is: brain tissue structure.
[0040] S102, generate a spoken script and a video storyboard script based on the prompt words through the large model.
[0041] The oral script is the text content used to guide oral output, and at least includes the text content that needs to be spoken; the video storyboard script is the content that converts the text script into a visual shooting plan. For example, in the process of medical surgery, the footage can be divided into three parts for display: before surgery, during surgery, and after surgery.
[0042] In some embodiments, in order to improve the accuracy of the output results of the large model, more detailed generation instructions can be added to the prompt words. For example, the prompt words can be: Please generate a spoken script and a video storyboard script for a medical video with a video length of 120 seconds according to the following requirements. The video content theme is: brain organizational structure.
[0043] It can be understood that the keywords in the prompt words are understood and extracted through the big model, such as "brain", "organizational structure", "120 seconds", "spoken script" and "video storyboard script" and other keywords are understood and processed to generate spoken scripts and video storyboard scripts that meet the needs of popular science information. Compared with manually written scripts, this embodiment generates spoken scripts and video storyboard scripts through the big model, which greatly improves the processing efficiency. When the spoken script or video storyboard script does not meet the expected needs, the prompt words can be adjusted to achieve rapid iteration of the spoken script and video storyboard script, reducing trial and error losses and lowering the creation threshold.
[0044] S103, according to the oral script, obtain the target materials required for popular science and the oral audio of the digital human.
[0045] In some embodiments, keyword extraction can be performed on the spoken script to obtain keywords in the spoken script, and target materials can be matched based on the keywords, where the target materials can be materials such as pictures or videos, so as to more intuitively display and understand the corresponding organizational structure matching target materials.
[0046] For example, assuming the spoken script is "The human brain is composed of the cerebral cortex, cerebellum and brainstem. The cerebral cortex dominates thinking and memory, the cerebellum coordinates movement, and the brainstem maintains vital signs", keywords such as "cerebral cortex", "cerebellum" and "brainstem" are extracted from the spoken script. Based on the keywords, corresponding target materials are matched from the material library, such as pictures or videos corresponding to each organizational structure such as "cerebral cortex", "cerebellum" and "brainstem", so as to display the popular science content more intuitively.
[0047] In some embodiments, the spoken audio of the digital person can be generated by a large model. For example, the spoken script and audio generation requirements are combined into prompt words, and the prompt words are input into the large model, and the large model generates a spoken video of the digital person that meets the requirements.
[0048] For example, assuming that the audio generation requirements include: a female voice of about 25 years old, a clear and sweet timbre, and a slower speaking speed; then the audio generation requirements and the oral script are spliced together to generate prompt words, and the prompt words are input into the large model to obtain the oral audio of the digital human. The oral audio is a clear and sweet female voice of about 25 years old, which broadcasts the oral script at a slower speaking speed.
[0049] In some embodiments, if real-person voice broadcast is required, real-person voice features and spoken habits can be collected, and prompt words can be obtained based on the real-person voice features, spoken habits and spoken script, and the prompt words can be input into the large model to obtain the digital person's spoken audio.
[0050] S104, generating a popular science video of the digital human based on the video storyboard script, target material and the digital human's spoken audio.
[0051] In some embodiments, a popular science video of a digital human can be generated based on a large model. The image of the digital human can be a real person image or a virtual character image. The lip shape of the digital human image is driven to provide explanations and matched with the digital human's spoken audio for playback.
[0052] It is understandable that the science popularization video of the digital human also includes the display of target materials. For example, during the oral broadcast of the science popularization video, the digital human will display the target materials corresponding to the content in the oral audio. For example, when the digital human broadcasts to the cerebellum, the target materials such as pictures or videos corresponding to the cerebellum will be displayed in the science popularization video. In addition, during the playback of the science popularization video, the storyboard processing is carried out according to the video storyboard script to obtain the science popularization video of the digital human with better display effect, which lowers the creation threshold while ensuring the generation effect of the science popularization video of the digital human.
[0053] In this embodiment, science popularization demand information is determined, and prompt words of the large model are generated based on the science popularization demand information. The large model generates a matching oral script and video storyboard script based on the prompt words, thereby lowering the professional threshold requirements for video creation. The target materials and digital human oral audio required for science popularization are obtained according to the generated oral script, and the digital human science popularization video is generated based on the video storyboard script, the target materials and the digital human oral audio, thereby ensuring the correctness of the oral audio in the science popularization video. The display effect of the science popularization video is improved through the video storyboard script, and the target materials are combined with the target materials during the display process to improve the science popularization effect, thereby improving the richness of the science popularization video and the video generation efficiency.
[0054] Figure 2 FIG. 1 is a schematic diagram of another method for generating a digital human video provided by an embodiment of the present disclosure. Figure 2 As shown, the method includes:
[0055] S201, determining science popularization demand information, and generating prompt words of a large model according to the science popularization demand information.
[0056] In the embodiment of the present disclosure, the implementation method of step S201 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0057] S202, generate a spoken script and a video storyboard script based on the prompt words through the large model.
[0058] In the embodiment of the present disclosure, the implementation method of step S202 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0059] S203, according to the oral script, obtain the target materials required for popular science and the oral audio of the digital human.
[0060] In the embodiment of the present disclosure, the implementation method of step S203 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0061] S204: Determine the shot information of the popular science video according to the video storyboard script.
[0062] In some embodiments, a video shooting plan can be extracted from a video storyboard script, and the lens information of the popular science video can be determined based on the video shooting plan; wherein the video storyboard script may include but is not limited to shooting angles, camera movement methods, storyboard cutting methods, storyboard shot sequences, shot durations, materials and material layouts, and other shooting plans; accordingly, according to the video storyboard script, the lens information can be determined, including but not limited to shot sequences, the content of each shot in the shot sequence, materials and material layouts, shot durations, shooting angles and camera movement methods, and other information.
[0063] As an example, the shooting plan including shooting angle, camera movement, frame division and lighting requirements can be determined from the video storyboard script. For example, if the shooting angle is to simulate the main perspective of the observer, the camera movement is to promote and emphasize key information, the frame division is to switch the frame display every 20 seconds, and the lighting requires natural light, then the lens information of the popular science video can be determined as: natural light, simulating the main perspective of the observer, promoting camera movement and switching the frame display every 20 seconds.
[0064] S205, obtaining the image of the popular science figure, and generating a digital human video based on the spoken audio and the image according to the lens information, thereby generating a target spoken video of the digital human.
[0065] In some embodiments, the character information of the science popularization figure can be determined based on the science popularization demand information; that is, the science popularization demand information can also include the demand for science popularization figures, and the character information of the science popularization figure can include character profile information such as the character's name, position or skills, and the character image of the science popularization figure is determined from the character database based on the character information.
[0066] Optionally, the character image of the science popularization character can be uploaded by the user character himself and stored in the character database.
[0067] For example, assuming that a digital human video in the medical field is generated, the digital human image can be constructed based on the image of a real doctor. The character information can be determined from the popular science demand information, and the doctor's name, position, and areas of expertise and other related information can be determined based on the character information. The corresponding doctor can be matched from the doctor database based on the character information to obtain the doctor's character image, and the doctor's character image can be used as the digital human image to improve the realism of the video.
[0068] In some embodiments, the spoken audio may be divided according to the shot information; for example, the spoken audio may be divided into one or more segments according to the shot division method in the shot information.
[0069] In some embodiments, the human image can be driven based on the voice-over audio part corresponding to the shot to generate the initial voice-over video of the digital human; optionally, the voice basic units in the voice-over audio part corresponding to the shot and the pronunciation order of the voice basic units can be obtained; according to the voice basic units and the corresponding pronunciation order, a lip movement sequence corresponding to the human image is generated.
[0070] In this embodiment, the voice basic unit can be a character, and the pronunciation order of the voice basic unit can be the lip shape order. For example, for the character "好" (good), its pronunciation order mainly depends on the two parts of its pinyin "h" and "ao". The lip shape for "h" is slightly open, and the lip shape for "ao" is round. Therefore, the pronunciation order is slightly open - round. Further, based on the voice basic unit and the corresponding pronunciation order, a lip movement sequence corresponding to the human image is generated, that is, based on the character order and the pronunciation order corresponding to each character, a lip movement sequence is generated.
[0071] Furthermore, the human image is driven based on the lip movement sequence to generate the initial voice-over video of the digital human. In this embodiment, a lip movement video generation algorithm of visual technology is used to drive the lips of the human image to move, so that the lip movement for pronunciation corresponds to the voice-over audio part corresponding to the current shot. For example, if the voice-over audio part corresponding to the shot is "小脑" (cerebellum), the lip movement for pronunciation of the human image should also be the same as the lip movement when a person really pronounces "小脑" to obtain the initial voice-over video of the digital human.
[0072] Furthermore, in this embodiment, a foreground matting task is performed on the initial voice-over video to obtain the target voice-over video of the digital human; for example, a foreground matting task is performed based on an optimized Relevance Vector Machine (RVM) model. Through data-driven feature learning, temporal dynamic modeling, and targeted edge optimization, problems such as motion blur, semi-transparent regions, and complex background interference are solved, and clear edge separation and matting are achieved. The key pixel points are identified as the foreground boundary, separating the digital human from the background to obtain the target voice-over video of the digital human.
[0073] In some embodiments, the digital human's head drive sequence can also be determined according to the spoken script; the target spoken video of the digital human can be adjusted according to the head drive sequence; optionally, the head drive action can include nodding, shaking the head or tilting the head. For example, in the question sentence part of the spoken script, the digital human's head action can be slightly tilted to reflect that the character is asking questions or thinking, then the head drive can be determined to be original-side tilt. Similarly, the digital human's head action is determined according to the scene changes in the spoken script, thereby obtaining the digital human's head drive sequence, and the target spoken video of the digital human is adjusted according to the head drive sequence, and the digital human's head action is realized to obtain the target spoken video of the digital human that is more in line with the real person's reaction, thereby improving the realism and expressiveness of the digital human in the target spoken video.
[0074] S206: Layout the target material according to the lens information to determine the material layout information of the lens, and generate a background video for the popular science video based on the material layout information of the lens.
[0075] The material layout information of the shot is determined based on the shot sequence in the shot information and the materials and material layout in the shot. In this embodiment, the material layout information at least includes the materials included in the shot and the layout positions of the materials.
[0076] In some embodiments, key point detection can be performed on the material included in the lens to obtain a key point set of the material; for example, if the material is a picture, the key points in the picture are detected to obtain a key point set; for example, in the material picture of a medical beauty scene, the key points can be key parts such as the edge of the face, lips and eyes, and key point detection is performed on the picture material to obtain a key point set in the picture material.
[0077] According to the layout position, the key point set is point-matched to determine the display position of the material in the background; that is, according to the layout position of the material, the key point set of the material is point-matched in the video to obtain the precise display position of the material in the video background, and the material is rendered based on the display position and the key points, that is, the material is rendered based on the key points at the display position to generate the background video of the popular science video. The background video may include one or more display positions, each of which displays the corresponding material.
[0078] S207, generating a popular science video based on the target spoken video and the background video.
[0079] Optionally, auxiliary information of the popular science video can also be obtained; in this embodiment, the auxiliary information of the popular science video can be obtained by determining the video header information and the video ending information based on the lens information, and performing dynamic graphic rendering on the video header information and the video ending information respectively to generate a header video sequence and an ending video sequence.
[0080] In this embodiment, the video opening information can be a summary of the main content in the popular science video, and the video ending information can be a brief introduction to the relevant information of the digital human corresponding to the character; dynamic graphics rendering is performed according to the lens requirements in the lens information and the video opening information and video ending information to obtain the opening video sequence and the ending video sequence. The opening video sequence displays the main content in the popular science video, and the ending video sequence displays the relevant brief information of the task.
[0081] In some embodiments, obtaining the auxiliary information of the science popularization video may also be obtaining the subtitles of the science popularization video; that is, the subtitles of the oral audio of the science popularization video are used as the auxiliary information.
[0082] In some embodiments, obtaining auxiliary information of the science popularization video can also be obtaining a cover image of the science popularization video as auxiliary information. The cover image can be rendered by the theme name of the science popularization video, thereby improving the diversity of the auxiliary information.
[0083] Furthermore, the target spoken video, background video and auxiliary information are combined to generate a popular science video; in some embodiments, the priority and layout information of the target spoken video, background video and auxiliary information can be determined; according to the priority and layout information, the target spoken video, background video and auxiliary information are combined according to the timeline to generate a popular science video, thereby improving the generation effect and display level of the popular science video.
[0084] As an example, assuming that the auxiliary information includes a title video sequence, an end video sequence, and a cover image, the priority and layout information of the target voice-over video, background video, title video sequence, end video sequence, and cover image are determined. For example, the cover image has the highest priority, followed by the title video sequence, followed by the target voice-over video and background video with the same priority, and finally the end video sequence. The layout positions of the cover image, title video sequence, and end video sequence are the entire video screen. The background video is laid out as the background of the target voice-over video in the entire video screen, and the target voice-over video is laid out as the foreground video in the lower right or lower left corner of the video screen. Based on the above priority and layout information, the target voice-over video, background video, and auxiliary information are combined according to the timeline to generate a popular science video. The playback order of the popular science video is cover image, title video sequence, target voice-over video and background video, and end video sequence. In this embodiment, the popular science video supports multi-layer display of foreground and background, and multi-type combination rendering and synthesis of subtitles, animation, and video, so that the popular science video generation effect is better.
[0085] In this embodiment, after obtaining the target material and the spoken audio of the digital human, the lens information can be determined according to the video storyboard script, and the character image of the popular science character can be obtained. The digital human video is generated based on the character image, lens information and spoken audio to obtain the target spoken video of the digital human. During the generation of the target spoken video, the character image is driven by the spoken audio part to determine the lip movement sequence. The lip movement sequence is combined with the lens information of the popular science needs to obtain a more accurate target spoken video. The head drive sequence of the digital human can also be obtained to improve the realism of the digital human in the digital human video. The target spoken video is used as the foreground part, and the target material is laid out in combination with the lens information to obtain the background video. Finally, the popular science video is generated based on the target spoken video of the foreground part, the background video and a variety of auxiliary information. The audio and video generation capability of the spoken script based on the real voice sample is realized, the sound is natural, and there is no need for the character to dub the script for different scripts. This solves the problem of high cost of video script creation and material matching for other groups, and improves the generation effect of the popular science video.
[0086] Figure 3 FIG. 1 is a schematic diagram of another method for generating a digital human video provided by an embodiment of the present disclosure. Figure 3 As shown, the method includes:
[0087] S301, determining science popularization demand information, and generating prompt words of a large model according to the science popularization demand information.
[0088] In the embodiment of the present disclosure, the implementation method of step S301 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0089] S302, generate a spoken script and a video storyboard script based on the prompt words through the large model.
[0090] In the embodiment of the present disclosure, the implementation method of step S302 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0091] S303, obtaining entity words of the oral script, and obtaining target materials required for popular science based on the entity words.
[0092] In some embodiments, entity words can be extracted from the spoken script to obtain one or more key entity words; entity words refer to words with clear reference objects, such as "mobile phone", "Beijing" or "2025" and other entity words with specific references to objects, place names or time.
[0093] For example, assuming the spoken script is "The human brain is composed of the cerebral cortex, cerebellum and brainstem. The cerebral cortex dominates thinking and memory, the cerebellum coordinates movement, and the brainstem maintains vital signs", then entity words are extracted from the spoken script, and the key entity words obtained may be words such as "human brain", "cerebral cortex", "cerebellum" and "brainstem".
[0094] In some embodiments, keywords displayed according to the target mode can be extracted from the spoken script, and key entity words can be extracted from the keywords; the target mode display can be a display mode such as highlighting or bolding. For example, when the spoken script is generated by a large model, the keywords in the spoken script are highlighted or bolded; when extracting key entity words, the key entity words are extracted from the keywords displayed in the target mode that are highlighted or bolded, thereby improving the efficiency of key entity word extraction.
[0095] Furthermore, video material information can be determined based on key entity words; the video material information can include relevant materials corresponding to the key entity words. For example, if the key entity word is "cerebral cortex", the video material information is material information related to "cerebral cortex", such as relevant videos or relevant pictures of "cerebral cortex".
[0096] According to the video material information, the target materials required for popular science are obtained from the material library; the material library may include picture materials and video materials for display, and matching materials are determined from the material library according to the video material information, such as multi-angle picture materials and multi-view video materials corresponding to the "cerebral cortex", so as to obtain accurate target materials required for popular science.
[0097] S304: Acquire an audio sample, and acquire the digital human's spoken audio according to the audio sample and the spoken script.
[0098] In some embodiments, the audio sample can be an audio sample of a real person corresponding to the digital human image. For example, the audio sample is obtained by recording the audio of the real person corresponding to the digital human image. The duration of the audio sample can be 5-10 seconds, which can reflect the speaking characteristics of the real person corresponding to the digital human image.
[0099] Optionally, after obtaining the audio sample, the timbre features of the audio sample can be obtained; the timbre features can include real-person speaking habits, such as sentence pauses and other habits, and can also include timbre and accent to fully obtain real-person speaking characteristics.
[0100] Furthermore, the spoken script can be converted into audio based on the timbre characteristics to obtain the spoken audio of the digital person; the spoken audio of the digital person can fully simulate the speaking habits, timbre and accent characteristics of a real person to obtain real and accurate digital human spoken audio.
[0101] Optionally, a fine-tuned and trained non-autoregressive text-to-speech model based on flow matching (AFairytaler that Fakes Fluent and Faithful Speech with Flow Matching, F5-TTS) can be used to convert the spoken script into audio to obtain the spoken audio of the digital human.
[0102] In some embodiments, the spoken script and the spoken audio can also be aligned to obtain the time relationship between the same text in the spoken script and the spoken audio; the time relationship is used to align the audio and subtitles in the spoken video of the digital person; for example, the audio text in the spoken audio, that is, the subtitles of the spoken audio, is obtained through voice recognition, and the audio text and the text between the spoken script are matched one by one to ensure that the audio text of the voice recognition is consistent with the spoken script; further obtain the time relationship between the same text in the spoken script and the spoken audio, so as to align the text timestamp of the spoken script with the timestamp of the same text in the spoken audio, and realize the time alignment of the digital person's spoken audio and the spoken script, thereby realizing the alignment of the audio and subtitles in the spoken video through text alignment and time alignment, and improving the video generation effect.
[0103] Optionally, this embodiment can use a stable timestamp synchronization tool (stable-ts) and a large-scale network supervised pre-trained speech recognition model (whisper) to perform time alignment and speech recognition, and achieve character-level speech recognition result alignment through acoustic feature analysis of polyphone matching and forced alignment correction.
[0104] S305, generating a popular science video of the digital human based on the video storyboard script, target material and the digital human's spoken audio.
[0105] In the embodiment of the present disclosure, the implementation method of step S305 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0106] In this embodiment, after generating the spoken script and the video storyboard script, keywords can be obtained based on the highlighted annotations in the spoken script, and the key entity words therein can be extracted to match the target material to obtain richer target material, and audio samples recorded by real people can be obtained. The spoken audio of the digital person is generated based on the timbre characteristics in the audio sample, fully simulating the speaking habits, timbre and accent characteristics of the real person to obtain real and accurate digital human spoken audio. Based on the video storyboard script, the target material and the spoken audio of the digital person, a popular science video of the digital person is generated to improve the generation effect of the popular science video.
[0107] Figure 4 FIG. 1 is a schematic diagram of another method for generating a digital human video provided by an embodiment of the present disclosure. Figure 4 As shown, the method includes:
[0108] S401, determining science popularization demand information, and generating prompt words of a large model according to the science popularization demand information.
[0109] In the embodiment of the present disclosure, the implementation method of step S401 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0110] S402, generate a spoken script and a video storyboard script based on the prompt words through the large model.
[0111] In the embodiment of the present disclosure, the implementation method of step S402 can be implemented by adopting any method in each embodiment of the present disclosure, which is not limited here and will not be described in detail.
[0112] S403, obtaining entity words of the oral script, and obtaining target materials required for popular science based on the entity words.
[0113] In the embodiment of the present disclosure, the implementation method of step S403 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0114] S404: Acquire an audio sample, and acquire the spoken audio of the digital human according to the audio sample and the spoken script.
[0115] In the embodiment of the present disclosure, the implementation method of step S404 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0116] S405: Determine the shot information of the popular science video according to the video storyboard script.
[0117] In the embodiment of the present disclosure, the implementation method of step S405 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0118] S406, obtaining the image of the popular science figure, and generating a digital human video based on the spoken audio and the image according to the lens information, thereby generating a target spoken video of the digital human.
[0119] In the embodiment of the present disclosure, the implementation method of step S406 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0120] S407: Layout the target material according to the lens information to determine the material layout information of the lens, and generate a background video for the popular science video based on the material layout information of the lens.
[0121] In the embodiment of the present disclosure, the implementation method of step S407 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0122] S408: Generate a popular science video based on the target spoken video and the background video.
[0123] In the embodiment of the present disclosure, the implementation method of step S408 can be implemented by using any of the methods in the embodiments of the present disclosure, which is not limited here and will not be described in detail.
[0124] In this embodiment, the popular science demand information is determined, and the prompt words of the large model are generated according to the popular science demand information. The large model generates a matching oral script and video storyboard script according to the prompt words, which reduces the professional threshold requirements for video creation. Keywords are obtained according to the highlighted annotations in the oral script, and the key entity words therein are extracted to match the target material to obtain richer target materials. In addition, audio samples recorded by real people are obtained, and the oral audio of the digital person is generated according to the timbre characteristics in the audio samples, which fully simulates the speaking habits, timbre and accent characteristics of the real person to obtain real and accurate digital human oral audio, which can be further determined according to the video storyboard script. The system uses lens information and obtains the image of the popular science figure. Based on the image, lens information and oral audio, digital human video is generated to obtain the target oral video of the digital human. The target oral video is used as the foreground part, and the target material is laid out in combination with the lens information to obtain the background video. Finally, a popular science video is generated based on the target oral video of the foreground part, the background video and a variety of auxiliary information. The ability to generate audio and video of oral scripts based on real-person voice samples is realized. The sound is natural, and there is no need for characters to dub scripts for different scripts. This solves the problem of high cost of creating video scripts and matching materials for other groups, and improves the generation effect of popular science videos.
[0125] Figure 5 This is a logic diagram of a digital human video generation system provided by an embodiment of the present disclosure. Figure 5 As shown, taking the generation of medical popular science videos as an example, the system may include components such as a production platform / manual script, a content processor (content-processor), a content agent (content-agent), an artificial intelligence platform (AI-platform), a health policy service, a rendering service, a corpus, a visual identity system (VIS), and a doctor service (doctorservice). Content submission is obtained through the production platform / manual script. In this embodiment, the content may be prompt words determined according to popular science needs. The prompt words are submitted to the content-processor through artificial intelligence-generated content aigc-submit to generate a spoken script and a video storyboard script. Specifically, the content-agent can assemble and infer the prompt words in real time based on the prompt words. Furthermore, based on the spoken script and the video storyboard script, the content-processor sequentially performs the following processes: querying basic information such as the doctor's photo, name, and department; single image cutout; generating digital human video material; generating head-turning video; generating opening and ending video; generating cover image; generating pre- and post-operative pictures; generating subtitles; proofreading word-by-word automatic speech recognition (ASR); and integrating and post-processing video synthesis materials.
[0126] Among them, basic information such as doctor's photo, name, department, etc. can be obtained through data query on the doctor service end doctorservice; single-image matting is based on the ai-platform to synchronously call the vis model to obtain the character image; digital human video material generation includes digital human oral audio synthesis (TTS audio), single-image digital human lip movement video generation and video matting and stroking. TTS audio synthesis and single-image digital human lip movement video generation are asynchronous tasks submitted to the ai-platform for task storage and queue execution, and the TTS audio in the health strategy service and the video generation results in vis are obtained; video matting and stroking can be achieved through the application programming interface (Application Programming Interface) The asynchronous task is submitted to the ai-platform for task storage and queue execution, and the video matting and stroking results in the health strategy service are obtained; head turning video generation refers to the generation of digital human head movement video, and the asynchronous task is submitted to the ai-platform for task storage and queue execution, and the head turning video generation results in vis are obtained; the opening and ending video generation includes the assembly of opening and ending materials, the generation of opening videos and the generation of ending videos, and both the opening and ending videos are rendered by the rendering service; the cover image generation includes the assembly of video cover materials and the synthesis of cover images, and the cover image synthesis is rendered by the rendering service to obtain the cover image; the pre-operative and post-operative image generation includes the detection of key facial points in the pre-operative and post-operative images, the matching of point coordinates of script part analysis points, and the rendering of pre-operative and post-operative part maps. The point detection is implemented by the ai-platform, and the rendering of pre-operative and post-operative part maps is rendered by the rendering service; the subtitle generation is rendered by the rendering service; the word-by-word ASR proofreading can be implemented by calling the stable-ts in the health strategy service through the application program interface; the integration of video synthesis materials can be implemented through the application program interface using the Fast Forward Moving Picture Experts (FMPE) in the corpus Group, FFmpeg) synthesizes the output with the main video to achieve material integration; post-processing can include result feedback callback, parsing paser, audit, risk assessment risker, storage store and distribution distribute process, which is stored in the distributed publish-subscribe messaging system pulsar, and pulsar is used to store and distribute the results.
[0127] Figure 6 This is a structural block diagram of a digital human video generation device provided by an embodiment of the present disclosure. Figure 6 As shown, the digital human video generating device 600 includes:
[0128] The first acquisition module 601 is used to determine science popularization demand information and generate prompt words of the large model according to the science popularization demand information;
[0129] The second acquisition module 602 is used to generate a spoken script and a video storyboard script according to the prompt words using the large model;
[0130] The third acquisition module 603 is used to acquire target materials and digital human audio required for popular science according to the oral script;
[0131] The generation module 604 is used to generate a popular science video of the digital human based on the video storyboard script, the target material and the spoken audio of the digital human.
[0132] In some embodiments, the third acquisition module 603 includes:
[0133] Entity words are extracted from the spoken script to obtain one or more key entity words;
[0134] Determine video material information based on key entity words;
[0135] According to the video material information, the target material required for popular science is obtained from the material library.
[0136] In some embodiments, the third acquisition module 603 includes:
[0137] Extract keywords displayed according to the target mode from the spoken script, and extract key entity words from the keywords.
[0138] In some embodiments, the third acquisition module 603 includes:
[0139] Get audio samples;
[0140] Get the timbre characteristics of the audio sample;
[0141] According to the timbre characteristics, the spoken script is converted into audio to obtain the spoken audio of the digital person.
[0142] In some embodiments, the third acquisition module 603 further includes:
[0143] Align the spoken script and the spoken audio to obtain the time relationship between the same text in the spoken script and the spoken audio. The time relationship is used to align the audio and subtitles in the digital human's spoken video.
[0144] In some embodiments, the generating module 604 includes:
[0145] Determine the lens information of the science popularization video based on the video storyboard script;
[0146] Obtain the image of the popular science figure, and generate a digital human video based on the spoken audio and the figure's image according to the lens information, thus generating the target spoken video of the digital human;
[0147] Layout the target material according to the lens information to determine the material layout information of the lens, and generate the background video of the popular science video based on the material layout information of the lens;
[0148] Generate a popular science video based on the target spoken video and background video.
[0149] In some embodiments, the generating module 604 includes:
[0150] Determine the information of science popularization figures based on science popularization demand information;
[0151] According to the character information, the character image of the science popularization character is determined from the character database.
[0152] In some embodiments, the generating module 604 includes:
[0153] Divide the spoken audio according to the lens information;
[0154] Based on the spoken audio portion corresponding to the shot, the character image is driven to generate the initial spoken video of the digital human;
[0155] Perform the foreground cutout task on the initial spoken video to obtain the target spoken video of the digital human.
[0156] In some embodiments, the generating module 604 further includes:
[0157] Determine the digital human's head driving sequence based on the spoken script;
[0158] According to the head drive sequence, the target oral video of the digital human is adjusted.
[0159] In some embodiments, the generating module 604 includes:
[0160] Obtaining the basic speech units and the pronunciation order of the basic speech units in the spoken audio portion corresponding to the shot;
[0161] Generate lip movement sequences corresponding to the character image based on the basic speech units and the corresponding pronunciation order;
[0162] The character image is driven based on the lip movement sequence to generate the initial spoken video of the digital human.
[0163] In some embodiments, the material layout information includes the materials included in the shot and the layout positions of the materials. The generation module 604 includes:
[0164] Perform key point detection on the material included in the shot to obtain a set of key points of the material;
[0165] According to the layout position, the key point set is matched to determine the display position of the material in the background;
[0166] Render the material based on the display position and key points to generate the background video of the popular science video.
[0167] In some embodiments, the generating module 604 includes:
[0168] Get auxiliary information for popular science videos;
[0169] Combine the target spoken video, background video and auxiliary information to generate a popular science video.
[0170] In some embodiments, the generating module 604 includes at least one of the following operations:
[0171] Determine the video title information and the video ending information according to the shot information, and perform dynamic graphic rendering on the video title information and the video ending information respectively to generate a title video sequence and an ending video sequence;
[0172] Get subtitles for popular science videos;
[0173] Get the cover image of the popular science video.
[0174] In some embodiments, the generating module 604 includes:
[0175] Determine the priority and layout of the target spoken video, background video, and auxiliary information;
[0176] According to the priority and layout information, the target spoken video, background video and auxiliary information are combined along the timeline to generate a popular science video.
[0177] In this embodiment, the popular science demand information is determined, and the prompt words of the large model are generated according to the popular science demand information. The large model generates a matching oral script and video storyboard script according to the prompt words, which reduces the professional threshold requirements for video creation. Keywords are obtained according to the highlighted annotations in the oral script, and the key entity words therein are extracted to match the target material to obtain richer target materials. In addition, audio samples recorded by real people are obtained, and the oral audio of the digital person is generated according to the timbre characteristics in the audio samples, which fully simulates the speaking habits, timbre and accent characteristics of the real person to obtain real and accurate digital human oral audio, which can be further determined according to the video storyboard script. The system uses lens information and obtains the image of the popular science figure. Based on the image, lens information and oral audio, digital human video is generated to obtain the target oral video of the digital human. The target oral video is used as the foreground part, and the target material is laid out in combination with the lens information to obtain the background video. Finally, a popular science video is generated based on the target oral video of the foreground part, the background video and a variety of auxiliary information. The ability to generate audio and video of oral scripts based on real-person voice samples is realized. The sound is natural, and there is no need for characters to dub scripts for different scripts. This solves the problem of high cost of creating video scripts and matching materials for other groups, and improves the generation effect of popular science videos.
[0178] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0179] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0180] Figure 7 A schematic block diagram of an electronic device for implementing an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0181] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0182] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0183] The computing unit 701 can be a variety of general and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the digital human video generation method. For example, in some embodiments, the digital human video generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the digital human video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the digital human video generation method by any other appropriate means (e.g., by means of firmware).
[0184] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0185] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0186] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0187] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0188] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0189] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0190] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0191] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for generating a digital human video, wherein: The method comprises: Determining science popularization demand information, and generating prompt words for a large model based on the science popularization demand information; Generate a spoken script and a video storyboard script based on the prompt words through the large model; According to the oral script, the target materials and the oral audio of the digital human required for popular science are obtained; A popular science video of the digital human is generated based on the video storyboard script, the target material and the spoken audio of the digital human.
2. The method according to claim 1, wherein According to the oral script, obtain the target materials required for science popularization, including: Extracting entity words from the spoken script to obtain one or more key entity words; Determining video material information according to the key entity words; According to the video material information, the target material required for the science popularization is obtained from the material library.
3. The method according to claim 2, wherein: The entity word extraction is performed on the oral script to obtain one or more key entity words, including: Keywords displayed according to the target mode are extracted from the oral script, and the key entity words are extracted from the keywords.
4. The method according to claim 1, wherein According to the oral script, obtaining the oral audio of the digital human includes: Get audio samples; Acquiring the timbre characteristics of the audio sample; According to the timbre characteristics, the spoken script is converted into audio to obtain the spoken audio of the digital human.
5. The method according to claim 1 or 4, wherein After obtaining the spoken audio of the digital human, the method further includes: The spoken script and the spoken audio are aligned to obtain the time relationship between the same characters in the spoken script and the spoken audio, wherein the time relationship is used to align the audio and subtitles in the spoken video of the digital human.
6. The method according to claim 1, wherein Generating a popular science video of the digital human based on the video storyboard, the target material and the spoken audio of the digital human includes: Determining shot information of the popular science video according to the video storyboard script; Acquire the image of the popular science figure, and generate a digital human video based on the spoken audio and the figure image according to the lens information, to generate a target spoken video of the digital human; Laying out the target material according to the lens information to determine the material layout information of the lens, and generating a background video for the popular science video based on the material layout information of the lens; The popular science video is generated according to the target spoken video and the background video.
7. The method according to claim 6, wherein: The step of obtaining the image of the popular science figure includes: Determining the information of the science popularization figure according to the science popularization demand information; According to the character information, the character image of the science popularization character is determined from a character database.
8. The method according to claim 6, wherein: The step of generating a digital human video from the spoken audio and the character image according to the lens information to generate a target spoken video of the digital human includes: Dividing the spoken audio according to the shot information; Based on the spoken audio portion corresponding to the shot, the character image is driven to generate an initial spoken video of the digital human; A foreground cutout task is performed on the initial spoken video to obtain a target spoken video of the digital human.
9. The method according to claim 8, wherein After performing the foreground cutout task on the initial spoken video to obtain the target spoken video of the digital human, the method further includes: Determining the head driving sequence of the digital human according to the oral script; According to the head driving sequence, the target spoken video of the digital human is adjusted.
10. The method according to claim 8 or 9, wherein: The method of driving the character image based on the spoken audio portion corresponding to the lens to generate the initial spoken video of the digital human includes: Obtaining basic speech units in the spoken audio portion corresponding to the shot and a pronunciation order of the basic speech units; generating a lip movement sequence corresponding to the character image according to the basic speech units and the corresponding pronunciation order; The character image is driven based on the lip movement sequence to generate an initial spoken video of the digital human.
11. The method according to any one of claims 6 to 9, wherein: The material layout information includes the materials included in the shot and the layout positions of the materials. The generating of the background video of the popular science video based on the material layout information of the shot includes: Performing key point detection on the material included in the shot to obtain a key point set of the material; According to the layout position, performing point matching on the key point set to determine the display position of the material in the background; The material is rendered based on the display position and the key points to generate a background video for the popular science video.
12. The method according to any one of claims 6 to 9, wherein: Generating the popular science video according to the target spoken video and the background video includes: Obtaining auxiliary information of the popular science video; The target spoken video, background video and auxiliary information are combined to generate the popular science video.
13. The method according to claim 12, wherein: The obtaining of auxiliary information of the popular science video includes at least one of the following operations: Determine video header information and video end information according to the shot information, and perform dynamic graphics rendering on the video header information and the video end information respectively to generate a header video sequence and an end video sequence; Obtain subtitles of the popular science video; Get the cover image of the popular science video.
14. The method according to claim 13, wherein The step of combining the target spoken video, the background video, and the auxiliary information to generate the popular science video includes: Determining the priority and layout information of the target spoken video, background video, and auxiliary information; According to the priority and layout information, the target spoken video, background video and auxiliary information are combined according to the time axis to generate the popular science video.
15. A digital human video generation device, comprising: A first acquisition module is used to determine science popularization demand information and generate prompt words of a large model according to the science popularization demand information; A second acquisition module is used to generate a spoken script and a video storyboard script according to the prompt words using the large model; The third acquisition module is used to obtain the target materials required for popular science and the digital human's spoken audio according to the spoken script; A generation module is used to generate a popular science video of the digital human based on the video storyboard script, the target material and the spoken audio of the digital human.
16. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.
18. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 14.
Citation Information
Cited By
Digital human audio and video processing method and device, storage medium and program product
CN122205198A