Video generation method and device, electronic equipment and storage medium
By extracting features from the reference video to generate digital human benchmark images and voice, combined with voice-driven models, the digital human live video is automatically generated, which solves the problem of complex and costly construction of digital human live broadcast rooms, and achieves efficient and realistic live broadcast effects.
Patent Information
- Application Number
- CN202510535203.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-11
AI Technical Summary
The construction process of digital people's live broadcast room is cumbersome, dependent on professional and technical personnel, time-consuming and labor-intensive, cost-effective, and difficult to optimize the live broadcast quality and interactive experience.
By extracting local area features and speech features in the reference video, a reference local image and target speech of digital people are generated, and a lip shape change is generated in combination with the voice-driven model, and a digital people live video is automatically generated.
显著简化直播间搭建流程,提高效率,降低成本,增强直播的真实感和交互体验,提供智能问答交互功能。
Smart Images

Figure CN120302120A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to the fields of deep learning, natural language processing, computer vision, and large model technology, and can be used in application scenarios such as intelligent assistants, virtual assistants, and intelligent e-commerce. More specifically, the present disclosure provides a video generation method, apparatus, electronic device, storage medium, and computer program product. Background Art
[0002] The construction of a digital human live broadcast room usually includes multiple processes such as script writing, scene design, digital human image and voice customization, and interaction function configuration. The production process is cumbersome and the construction efficiency is very low. It relies heavily on professional technicians, and it is difficult to optimize the final presented live broadcast quality and interaction experience. Summary of the Invention
[0003] The present disclosure provides a video generation method, apparatus, electronic device, storage medium, and computer program product.
[0004] According to a first aspect, a video generation method is provided. The method includes: extracting local region features and speech features of a target object in a reference video; generating a reference local image of a digital human for the target object according to the local region features; generating a target speech corresponding to the target text according to the speech features and the target text; generating a local image sequence according to the target speech and the reference local image, where the local image sequence represents the lip movement changes of the digital human when uttering the target speech; and generating a target video of the digital human according to the reference video and the local image sequence.
[0005] According to a second aspect, a video generation apparatus is provided. The apparatus includes: an extraction module for extracting local region features and speech features of a target object in a reference video; an image generation module for generating a reference local image of a digital human for the target object according to the local region features; a speech generation module for generating a target speech corresponding to the target text according to the speech features and the target text; an image sequence generation module for generating a local image sequence according to the target speech and the reference local image, where the local image sequence represents the lip movement changes of the digital human when uttering the target speech; and a video generation module for generating a target video of the digital human according to the reference video and the local image sequence.
[0006] According to a third aspect, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the present disclosure.
[0007] According to a fourth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided according to the present disclosure.
[0008] According to a fifth aspect, there is provided a computer program product including a computer program stored on at least one of a readable storage medium and an electronic device, the computer program, when executed by a processor, implementing the method provided according to the present disclosure.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0011] Figure 1 is a schematic diagram of an exemplary system architecture to which a video generation method and apparatus can be applied according to an embodiment of the present disclosure;
[0012] Figure 2 is a flowchart of a video generation method according to an embodiment of the present disclosure;
[0013] Figure 3 is a schematic diagram of a target video according to an embodiment of the present disclosure;
[0014] Figure 4 is a flowchart of a video generation method according to another embodiment of the present disclosure;
[0015] Figure 5 is a block diagram of a video generation apparatus according to an embodiment of the present disclosure; and
[0016] Figure 6 is a block diagram of an electronic device for a video generation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following describes exemplary embodiments of the present disclosure with reference to the drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.
[0018] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0019] In the technical solutions of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting the user's personal information.
[0020] In some examples, the construction of a digital human live broadcast room usually relies on complex manual processes, including script writing, scene design, digital human image and voice customization, and interactive function configuration. These processes are not only time-consuming and laborious, but also require the participation of professional technical personnel, resulting in high construction costs and low efficiency. Based on this, the present disclosure proposes a video generation method for quickly constructing a digital human live broadcast room according to a live broadcast reference video, aiming to solve the problems of cumbersome construction processes, high labor and time costs in the current digital human live broadcast room, and providing efficient and professional technical support for digital human live broadcasts.
[0021] Figure 1 It is a schematic diagram of an exemplary system architecture to which the video generation method and device according to an embodiment of the present disclosure can be applied. It should be noted that Figure 1 The figure shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0022] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0023] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 may be various electronic devices, including but not limited to smart phones, tablet computers, laptop portable computers, etc.
[0024] Server 105 may be a server that provides various services. For example, it can be a background management server (only for example) that supports the websites browsed by users using terminal devices 101, 102, and 103. The background management server can analyze data such as video information uploaded by users, process it through the models deployed on server 105, and feedback the generated processing results to the terminal devices. Server 105 may deploy a pre-trained predetermined model, and the predetermined model can be a Large Language Model (LLM), a speech synthesis model, a speech-driven model, etc.
[0025] The video generation method provided by the embodiments of the present disclosure can generally be executed by server 105. Correspondingly, the video generation device provided by the embodiments of the present disclosure can generally be set in server 105.
[0026] Figure 2 It is a flowchart of a video generation method according to an embodiment of the present disclosure.
[0027] As Figure 2 shown, the video generation method 200 includes operations S210 to S250.
[0028] In operation S210, extract the local region features and speech features of the target object in the reference video.
[0029] In the embodiments of the present disclosure, the reference video may refer to a live replay video including the live broadcast content of the anchor, or any recorded video. The target object may refer to the anchor in the reference video. The local region features may refer to the partial region features of the target object, such as facial expression features, mouth features, limb movement features, etc. The speech features may refer to the voiceprint features of the target object.
[0030] For example, the reference video is a live replay video of an anchor selling electronic products. The target object may refer to the anchor in the reference video. The facial region of the anchor in the reference video can be intercepted to obtain a facial image, and then the facial region features can be extracted from the facial image; the speech data of the anchor during the live broadcast in the reference video can be extracted to obtain the speech features.
[0031] In operation S220, generate a reference local image of the digital human for the target object according to the local region features.
[0032] In the embodiments of the present disclosure, the digital human may refer to a digital human image similar to the anchor's image. The reference local image may refer to an image of a local region of the digital human. The local region of the digital human corresponds to the local region of the target object in the reference video.
[0033] For example, the local area feature refers to the mouth feature of the target object. The reference local image can be the mouth image of the digital human generated according to the mouth feature of the target object.
[0034] For another example, the local area feature refers to the facial feature of the target object. The reference local image can be the facial image of the digital human generated according to the facial feature of the target object.
[0035] In an embodiment of the present disclosure, the reference local image of the digital human can be obtained according to the local area feature of the target object in the natural static state.
[0036] For example, when in the natural static state, the local area feature of the target object is manifested as the upper and lower lips being closed, and the corresponding reference local image can be the mouth image with the upper and lower lips closed.
[0037] For example, when in the natural static state, the local area feature of the target object is manifested as the two arms hanging naturally and the palms being open, and the corresponding reference local image can be the upper body image with the two arms hanging naturally and the palms being open.
[0038] In operation S230, according to the voice feature and the target text, a target voice corresponding to the target text is generated.
[0039] In an embodiment of the present disclosure, the target text can refer to the copywriting information for live broadcast. The target voice can be the voice data obtained by performing voice conversion on the target text and having the voice feature of the target object.
[0040] In operation S240, according to the target voice and the reference local image, a local image sequence is generated, where the local image sequence can represent the corresponding change of the local area when the digital human emits the target voice.
[0041] In an embodiment of the present disclosure, the local area feature can refer to the mouth feature of the target object, the reference local image can be the mouth image of the digital human with the upper and lower lips naturally closed, and the local image sequence can represent the lip shape change when the digital human emits the target voice.
[0042] In another embodiment, the local area feature can refer to the upper body feature of the target object, the reference local image can be the upper body image of the digital human in a natural and relaxed state, and the local image sequence can represent the mouth and upper limb movement changes when the digital human emits the target voice.
[0043] In operation S250, according to the reference video and the local image sequence, a target video of the digital human is generated.
[0044] In an embodiment of the present disclosure, the target video can be a new live broadcast video generated based on the reference video. In the target video, the local area feature of the digital human can be determined based on the local image sequence.
[0045] For example, a partial region of a target object in a reference video can be replaced with a partial image sequence to obtain a target video.
[0046] According to an embodiment of the present disclosure, by extracting the local region features and speech features of the host from a reference video, generating a reference local image of a digital human according to the local region features, generating a target speech with specific voice features according to the speech features and the target text, and then driving the reference local image of the digital human to undergo corresponding lip shape changes based on the target speech to obtain a partial image sequence, and generating a target video based on the partial image sequence on the basis of the reference video. In the digital human live broadcast scenario, it is possible to automatically generate a digital human live broadcast video based on the reference video, improving the efficiency of live broadcast video generation.
[0047] According to an embodiment of the present disclosure, generating a digital human live broadcast video with an automated video generation process can significantly simplify the live broadcast room setup process, improve the live broadcast room setup efficiency, and reduce the setup threshold and production cost of the digital human live broadcast room.
[0048] In an embodiment of the present disclosure, generating a target video of a digital human according to a reference video and a partial image sequence includes: replacing a partial region of a target object in the reference video with a partial image in the partial image sequence.
[0049] For example, the target speech can be added to the target video, and the mouth region of the target object in the reference video can be replaced with a partial image sequence that undergoes lip shape changes with the target speech, so that the speech in the digital human live broadcast room matches the lip shape of the digital human, improving the realism of the digital human live broadcast.
[0050] According to an embodiment of the present disclosure, extracting the local region features and speech features of a target object in a reference video includes: identifying multiple objects and the state of each object in each image from multiple images of the reference video; for each object, determining the number of image frames in which the object is in the first state; and determining the target object from multiple objects according to the number of image frames in which each object is in the first state.
[0051] In an embodiment of the present disclosure, the reference video may include multiple objects including the host. The images of the reference video can be recognized to determine the object with the identity of the host as the target object from multiple objects. The first state may refer to a state in which the object is speaking with an open mouth or a state in which the face is facing the camera.
[0052] For example, the reference video includes a host and two assistant hosts. Three objects and the state of the three objects in each image can be obtained through multiple images of the reference video. By counting the number of image frames in which the three objects are speaking with an open mouth, the object with the most image frames of speaking with an open mouth is determined as the target object.
[0053] According to an embodiment of the present disclosure, generating a reference local image of a digital human for a target object based on local region features includes: intercepting a reference image of the target object in a second state from multiple images of a reference video; and generating a reference local image of the digital human for the target object according to the local region features of the target object in the reference image.
[0054] In an embodiment of the present disclosure, the second state may refer to a state where the object is in a natural static state or facing the camera frontally.
[0055] For example, if a reference image of the target object facing the camera frontally with the mouth naturally closed is intercepted from the images of the reference video, a corresponding reference local image of the digital human facing the camera frontally with the mouth closed can be generated according to the mouth features of the target object in the reference image.
[0056] According to an embodiment of the present disclosure, the voice feature includes a voiceprint feature; generating a target voice corresponding to the target text according to the voice feature and the target text includes: using a voice synthesis model, based on the voiceprint feature and the target text, converting the target text into a target voice with the voiceprint feature.
[0057] In an embodiment of the present disclosure, the voice synthesis model can convert the target text into a target voice. Based on the voiceprint feature, the voice synthesis model can convert the target text into a target voice with a specific timbre.
[0058] For example, the voice synthesis model converts the target text into a target voice with the timbre of an anchor, so that the timbre of the target voice emitted by the digital human is consistent with the timbre of the anchor in the reference video, increasing the realism of the digital human live broadcast.
[0059] According to an embodiment of the present disclosure, the local region includes a lip region; generating a local image sequence according to the target voice and the reference local image includes: using a voice-driven model to drive the lip region in the reference local image to generate a lip shape change corresponding to the prosody feature of the target voice, obtaining the local image sequence.
[0060] In an embodiment of the present disclosure, when the target object is speaking, the lip region of the target object will generate corresponding lip shape changes based on the spoken voice. Therefore, in order to make the lip region of the digital human more conform to the real human image, the shape of the lip region in the reference local image can be changed through a voice-driven model based on the prosody feature of the target voice, so that the lip region of the digital human can generate corresponding lip shape changes according to the target voice.
[0061] For example, if the target voice is "Good evening", the voice-driven model can generate multiple lip image corresponding to "wan", "shang", and "hao" respectively according to the target voice, obtain a local image sequence, and after replacing the local area of the target object in the reference video according to the local image sequence, a lip movement corresponding to the target voice "Good evening" can be obtained.
[0062] According to an embodiment of the present disclosure, the target video of the digital human not only includes the digital human image, but may also include the scene image during the live broadcast and the display text describing the live broadcast content. Next, in combination with Figure 3 The target video of the digital human of the present disclosure will be further described.
[0063] Figure 3 It is a schematic diagram of a target video according to an embodiment of the present disclosure.
[0064] As Figure 3 shown, the target video 300 includes multiple digital humans 310, an interactive information display area 320, and a background image 330. Among them, the digital human 311 can be the host.
[0065] In an embodiment of the present disclosure, the reference video is a live video; the target text can be generated according to the live broadcast copywriting and the copywriting style in the reference video; or the target text can be generated according to the live broadcast copywriting, live broadcast item information, and comment information in the historical live video of the target object.
[0066] For example, based on user needs, the copywriting style can be determined to be professional explanation, humorous interaction, promotional sales, etc., and different types of target texts can be generated.
[0067] In an embodiment of the present disclosure, the target text can also be processed by a large model for sensitive word filtering, false propaganda detection, etc. to obtain legal and compliant text information.
[0068] In an embodiment of the present disclosure, the content in the target text can be marked, and attribute tags such as "interaction attribute", "item characteristics", "selling price", etc. can be added based on the characteristics of each sentence in the target text. When the digital human 311 broadcasts the text with the attribute tag, the corresponding information can be displayed in the live broadcast room. For example, when the digital human 311 broadcasts to "item characteristics", a sticker containing the characteristic information can be inserted in the target video.
[0069] According to an embodiment of the present disclosure, the scene image can also be generated according to the live broadcast item information and the background style of the reference video; and the background image in the target video can be replaced with the scene image.
[0070] For example, if the live item information is disinfected laundry detergent and the background style of the reference video is neat and fresh, a scene image related to the disinfected laundry detergent and with a fresh style can be generated according to the live item information and the background style of the reference video, and the background image 330 in the target video 300 can be replaced with the generated scene image.
[0071] In an embodiment of the present disclosure, the target video may further include live lighting and live props, and information such as the brightness, color, and position of the live lighting, and the position and size of the live props can be adjusted based on user requirements.
[0072] According to an embodiment of the present disclosure, by intelligently generating the live scene and providing a real-time adjustment function, the live effect is optimized during the digital human live broadcast, the operation process is significantly simplified, and the labor and time costs are reduced.
[0073] According to an embodiment of the present disclosure, when a live broadcast start instruction is received, the target video may also be played; live analysis data of the target video may be generated; and at least one of the target text and the scene image may be adjusted according to the live analysis data.
[0074] The live broadcast start instruction may be initiated by the user after confirming the target video. After receiving the live broadcast start instruction, the live broadcast room may be opened and the target video of the digital human may be played. The live analysis data may refer to the analysis data obtained by analyzing data such as the number of viewers of the live broadcast, the interaction rate, and the sales amount. In the case where the live analysis data indicates that the live broadcast experience is not good, the live broadcast effect may be optimized by adjusting the target text or the scene image.
[0075] According to an embodiment of the present disclosure, the reference video is a live broadcast video; reference interaction information may be generated according to the interaction information in the reference video; and in response to receiving the live broadcast start instruction, the target video may be played and the reference interaction information may be displayed in the target video.
[0076] The interaction information may include questions from the viewers watching the live broadcast in the reference video and the host's reply information, and may also include the host's live broadcast introduction information and the viewers' response information. The reference interaction information corresponding to the target video and the live item information may be generated according to the style and content of the interaction information in the reference video. The reference interaction information may be displayed in the interaction information display area 320 in the target video 300.
[0077] According to an embodiment of the present disclosure, it further includes: in response to receiving question information, generating a response information for the question information based on a question and answer library, where the question and answer library is generated based on the interaction information in the historical live broadcast video; and displaying the question information and the response information in the target video.
[0078] Question information can be initiated by the audience watching the live broadcast of the digital human. For the digital human live broadcast room, based on the Q&A library, response information for the question information can be obtained. The Q&A library can be generated according to the interaction information of historical live videos, or can be generated according to the information of live items and the configured intelligent Q&A model.
[0079] For example, the question information of the audience in the live broadcast room is "How much is this piece of clothing after the discount". Based on the Q&A library, the response information can be determined as "This piece of clothing is xx yuan". According to the live item information, the original price of this piece of clothing can be determined as 100 yuan and the discount is 90%. According to the Q&A library and the live item information, the response information can be determined as "This piece of clothing is 90 yuan".
[0080] The above text information and response information can also be displayed in the interaction information display area 320.
[0081] In the embodiments of the present disclosure, when the audience in the live broadcast room selects a live item, the target video can actively provide the audience with information such as pictures, names, prices, promotion content, etc. related to the live item.
[0082] According to the embodiments of the present disclosure, the intelligent Q&A interaction function can be provided while the digital human is live broadcasting, reducing the labor and time costs, providing efficient and convenient technical support for the digital human live broadcast, and improving the live broadcast quality.
[0083] Figure 4 It is a flowchart of a video generation method according to another embodiment of the present disclosure.
[0084] As Figure 4 shown, the video generation method includes operations S401~S406.
[0085] In operation S401, a reference video is obtained.
[0086] In operation S402, according to the live copywriting and copywriting style in the reference video, a target text is generated.
[0087] In operation S403, a target video of the digital human is generated according to the reference video.
[0088] In operation S404, according to the interaction information in the reference video, reference interaction information is generated.
[0089] In operation S405, according to the live item information and the background style of the reference video, a scene image is generated, and the background image in the target video is replaced with the scene image.
[0090] In operation S406, in response to receiving a live broadcast instruction, the target video is played and the reference interaction information is displayed in the target video.
[0091] According to an embodiment of the present disclosure, the present disclosure also provides a video generating device.
[0092] Figure 5 is a block diagram of a video generating apparatus according to an embodiment of the present disclosure.
[0093] like Figure 5 As shown, the video generating device 500 includes an extraction module 510 , an image generating module 520 , a speech generating module 530 , an image sequence generating module 540 and a video generating module 550 .
[0094] The extraction module 510 is used to extract local area features and speech features of the target object in the reference video.
[0095] The image generation module 520 is used to generate a reference local image of a digital human for a target object according to local area features.
[0096] The speech generation module 530 is used to generate a target speech corresponding to the target text according to the speech features and the target text.
[0097] The image sequence generation module 540 is used to generate a local image sequence according to the target speech and the reference local image, wherein the local image sequence represents the lip shape changes of the digital human when making the target speech.
[0098] The video generation module 550 is used to generate a target video of a digital human according to a reference video and a local image sequence.
[0099] According to an embodiment of the present disclosure, the extraction module 510 includes a recognition submodule, a frame number determination submodule, and an object determination submodule.
[0100] The recognition submodule is used to recognize multiple objects and the state of each object in each image from multiple images of the reference video. The frame number determination submodule is used to determine the number of image frames in which the object is in the first state for each object. The object determination submodule is used to determine the target object from multiple objects based on the number of image frames in which each object is in the first state.
[0101] According to an embodiment of the present disclosure, the image generation module 520 includes a capture submodule and an image generation submodule.
[0102] The capture submodule is used to capture a reference image of the target object in the second state from multiple images of the reference video. The image generation submodule is used to generate a reference local image of the digital human for the target object according to the local area features of the target object in the reference image.
[0103] According to an embodiment of the present disclosure, the voice feature includes a voiceprint feature, and the voice generation module 530 includes a voice generation sub-module for converting the target text into a target voice with the voiceprint feature based on the voiceprint feature and the target text using a voice synthesis model.
[0104] According to an embodiment of the present disclosure, the local area includes a lip area, and the image sequence generation module 540 includes a voice-driven sub-module for driving the lip area in the reference local image to generate a lip shape change corresponding to the prosody feature of the target voice using a voice-driven model, thereby obtaining a local image sequence.
[0105] According to an embodiment of the present disclosure, the video generation module 550 includes a replacement sub-module for replacing the local area of the target object in the reference video with the local image in the local image sequence.
[0106] According to an embodiment of the present disclosure, the reference video is a live video, and the video generation device 500 further includes a first text generation module and a second text generation module.
[0107] The first text generation module is used to generate a target text according to the live copywriting and the copywriting style in the reference video. The second text generation module is used to generate a target text according to the live copywriting, the live item information, and the comment information in the historical live video of the target object.
[0108] According to an embodiment of the present disclosure, the video generation device 500 further includes a scene generation module and a scene replacement module.
[0109] The scene generation module is used to generate a scene image according to the live item information and the background style of the reference video. The scene replacement module is used to replace the background image in the target video with the scene image.
[0110] According to an embodiment of the present disclosure, the video generation device 500 further includes a playback module, a data generation module, and an adjustment module.
[0111] The playback module is used to play the target video in response to receiving a live broadcast instruction. The data generation module is used to generate live analysis data of the target video. The adjustment module is used to adjust at least one of the target text and the scene image according to the live analysis data.
[0112] According to an embodiment of the present disclosure, the reference video is a live video, and the video generation device 500 further includes an information generation module and a first display module.
[0113] The information generation module is used to generate reference interaction information according to the interaction information in the reference video. The first display module is used to play the target video and display the reference interaction information in the target video in response to receiving a live broadcast instruction.
[0114] According to an embodiment of the present disclosure, the video generation device 500 further includes a response module and a second display module.
[0115] The response module is configured to generate a response message to the question message based on a question-answer library in response to receiving the question message, where the question-answer library is generated based on interaction information in historical live videos. The second display module is configured to display the question message and the response message in the target video.
[0116] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0117] Figure 6 A schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0118] As Figure 6 shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0119] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0120] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the video generation method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the video generation method by any other suitable means (e.g., by means of firmware).
[0121] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0122] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0124] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0125] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0126] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of the client and the server is generated by computer programs running on the respective computers and having a client-server relationship to each other.
[0127] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0128] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A video generation method, comprising: Extracting local region features and speech features of a target object in a reference video; Generating a reference local image of a digital human for the target object according to the local region features; Generating a target speech corresponding to the target text according to the speech features and the target text; Generating a local image sequence according to the target speech and the reference local image, wherein the local image sequence represents the lip movement changes of the digital human uttering the target speech; and Generating a target video of the digital human according to the reference video and the local image sequence.
2. The method according to claim 1, wherein The extracting local region features and speech features of a target object in a reference video comprises: Identifying a plurality of objects and the states of each object in each image from a plurality of images of the reference video; For each object, determining the number of image frames in which the object is in a first state; and Determining the target object from the plurality of objects according to the number of image frames in which each object is in the first state.
3. The method according to claim 1, wherein, The generating a reference local image of a digital human for the target object according to the local region features comprises: Cropping a reference image in which the target object is in a second state from a plurality of images of the reference video; and Generating a reference local image of a digital human for the target object according to the local region features of the target object in the reference image.
4. The method according to claim 1, wherein The speech features include voiceprint features; The generating a target speech corresponding to the target text according to the speech features and the target text comprises: Using a speech synthesis model, based on the voiceprint features and the target text, converting the target text into a target speech with the voiceprint features.
5. The method according to claim 1, wherein The local region includes a lip region; the generating a local image sequence according to the target speech and the reference local image comprises: Using a voice-driven model to drive the lip region in the reference local image to generate lip movement changes corresponding to the prosody features of the target speech, and obtaining the local image sequence.
6. The method according to claim 1, wherein The generating a target video of the digital human according to the reference video and the local image sequence comprises: Replacing the local region of the target object in the reference video with the local image in the local image sequence.
7. The method according to claim 1, wherein The reference video is a live video; the method further comprises: Generating the target text according to the live copywriting and the copywriting style in the reference video; or Generating the target text according to the live copywriting, live item information and comment information in the historical live video of the target object.
8. The method according to claim 7, further comprising: Generating a scene image according to the live item information and the background style of the reference video; And Replacing the background image in the target video with the scene image.
9. The method according to claim 8, further comprising: In response to receiving a live broadcast instruction, playing the target video; Generating live analysis data of the target video; And Adjusting at least one of the target text and the scene image according to the live analysis data.
10. The method according to claim 1, wherein, The reference video is a live video; the method further comprises: Generate reference interaction information based on the interaction information in the reference video; and In response to receiving a live broadcast instruction, play the target video and display the reference interaction information in the target video.
11. The method according to claim 10, further comprising: In response to receiving question information, generate a response message for the question information based on a question and answer library, where the question and answer library is generated based on the interaction information in historical live video; and Display the question information and the response message in the target video.
12. A video generation device, comprising: An extraction module, configured to extract local region features and voice features of a target object in a reference video; An image generation module, configured to generate a reference local image of a digital human for the target object according to the local region features; A voice generation module, configured to generate a target voice corresponding to the target text according to the voice features and the target text; An image sequence generation module, configured to generate a local image sequence according to the target voice and the reference local image, where the local image sequence represents the lip movement changes of the digital human uttering the target voice; and A video generation module, configured to generate a target video of the digital human according to the reference video and the local image sequence.
13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program, the computer program being stored on at least one of a readable storage medium and an electronic device, and the computer program, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Cited By
Video generation method, electronic equipment and storage medium
CN121151655A
Large model-based audio and video data generation method, training method and intelligent agent
CN121306089A
Audio and video data generation methods, training methods, and intelligent agents based on large models
CN121306089B