Video generation method and device, equipment and storage medium
Reference images, text or audio are obtained through the configuration interface, and when generating video content, not only the face area is driven, but also other areas are driven, solving the problem of poor video reality in the prior art and improving the quality of the video.
Patent Information
- Application Number
- CN202510246535.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-06
AI Technical Summary
When using voice content to drive characters' actions, the prior art can only control the facial area, resulting in a poor sense of reality in the generated video.
By providing a configuration interface, obtaining reference images and reference text or reference audio, when generating video content, not only drives the face area, but also other areas in the image, improving the video quality.
It realizes the use of text or audio to drive multiple areas in the reference image to improve the realism and quality of the generated video.
Smart Images

Figure CN120111319A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for providing video generation. Background Art
[0002] With the development of computer technology, some machine learning technologies support users to drive images through control signals to generate dynamic video content. For example, users can drive digital humans to perform actions corresponding to audio signals by inputting audio signals. Summary of the invention
[0003] In a first aspect of the present disclosure, a method for video generation is provided. The method includes: in response to receiving a video generation request, presenting a configuration interface, the configuration interface including a first control and a second control; obtaining configuration information via the configuration interface, the configuration information including: a reference image obtained via the first control, and a reference text or reference audio obtained via the second control, wherein the reference image includes a preset object; and providing video content generated based on the configuration information, the video content including a first motion associated with a first area of the reference image and a second motion associated with a second area of the reference image, the first area including a facial area, and the first motion corresponding to the reference text or reference audio.
[0004] In a second aspect of the present disclosure, a device for video generation is provided. The device includes: a presentation module configured to present a configuration interface in response to receiving a video generation request, the configuration interface including a first control and a second control; an acquisition module configured to acquire configuration information via the configuration interface, the configuration information including: a reference image acquired via the first control, and a reference text or reference audio acquired via the second control, wherein the reference image includes a preset object; and a providing module configured to provide video content generated based on the configuration information, the video content including a first motion associated with a first area of the reference image and a second motion associated with a second area of the reference image, the first area including a facial area, and the first motion corresponding to the reference text or reference audio.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device executes the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.
[0007] It should be understood that the contents described in this content section are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0009] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;
[0010] FIG. 2A to FIG. 2D A schematic diagram of a video generation interface according to some embodiments of the present disclosure is shown;
[0011] Figure 3 Schematic diagrams of video generation interfaces according to other embodiments of the present disclosure are shown.
[0012] Figure 4 A flowchart illustrating an example process of video generation according to some embodiments of the present disclosure;
[0013] Figure 5 A schematic structural block diagram of a device for video generation according to some embodiments of the present disclosure is shown;
[0014] Figure 6 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0016] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this article, and any type of embodiment may be included under any section / subsection. In addition, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0018] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects are subject to the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user knows and confirms. Accordingly, when implementing each embodiment of the present disclosure, the type, scope of use, usage scenario, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method can vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0019] In this specification and the embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or it is necessary to perform a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.
[0020] Traditionally, in a scene where voice content is used to drive character actions, usually only the facial area of the character can be controlled to present corresponding actions, such as mouth movements. However, the video generated in this way has poor realism.
[0021] The embodiment of the present disclosure proposes a scheme for video generation. According to the scheme, in response to receiving a video generation request, a configuration interface is presented, the configuration interface includes a first control and a second control; configuration information is obtained through the configuration interface, the configuration information includes: a reference image obtained through the first control, and a reference text or reference audio obtained through the second control, wherein the reference image includes a preset object; and video content generated based on the configuration information is provided, the video content includes a first motion associated with a first area of the reference image and a second motion associated with a second area of the reference image, the first area includes a facial area, and the first motion corresponds to the reference text or reference audio.
[0022] In this way, the embodiments of the present disclosure can support the use of text or audio to drive the reference image, and can not only drive the facial area to present corresponding motion, but also drive other areas in the reference image to present corresponding motion. Thus, the embodiments of the present disclosure can improve the quality of the generated video content.
[0023] Example Environment
[0024] Figure 1 1 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, example environment 100 may include electronic device 110 .
[0025] In this example environment 100, an electronic device 110 may run an application 120 that supports generating video content. The application 120 may be any suitable type of application for generating video content, examples of which may include, but are not limited to, video applications, editing applications, or other suitable applications. A user 140 may interact with the application 120 via the electronic device 110 and / or its attached devices.
[0026] exist Figure 1 In the environment 100 , if the application 120 is in an active state, the application 120 may provide a presentation interface 150 for the user 140 .
[0027] In some embodiments, the electronic device 110 communicates with the server 130 to provide services for the application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio receiver, an e-book device, a game device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the target user (such as a "wearable" circuit, etc.).
[0028] The server 130 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The server 130 may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc. The server 130 may provide background services for the application 120 that supports interface interaction in the electronic device 110.
[0029] A communication connection may be established between the server 130 and the electronic device 110. The communication connection may be established in a wired manner or a wireless manner. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the server 130 and the electronic device 110 may implement signaling interaction through the communication connection between the two.
[0030] It should be understood that the structure and function of the various elements in the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.
[0031] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0032] Example Interaction
[0033] Example 1
[0034] The following will refer to Figure 2A-2D To describe an example interaction process according to an embodiment of the present disclosure. FIG. 2A to FIG. 2D Example interfaces 200A to 200D according to some embodiments of the present disclosure are shown. Interfaces 200A to 200D may be composed of, for example, Figure 1 The electronic device 110 is provided as shown.
[0035] In some scenarios, the electronic device 110 may install an application 120 for generating a video, or the electronic device 110 may also support accessing a preset website through a browser to view a corresponding interface. As an example, such an application 120 may correspond to a media editing application or other appropriate type of application, which may support a user to initiate a video generation request.
[0036] As an example, the electronic device 110 may receive a user's trigger for a video generation portal (or video generation service) provided by the application 120, and may accordingly receive the user's video generation request. As an example, such a video generation portal may correspond to a "line performance" service. Accordingly, the electronic device 110 may present a video of the video being generated. Figure 2A Configuration interface 200A is shown.
[0037] like Figure 2A As shown, the configuration interface 200A may include multiple controls to determine the configuration information used to generate the video content. Such configuration information may also be referred to as video generation parameters. The following will specifically introduce the process of obtaining the configuration information.
[0038] like Figure 2A As shown, the configuration interface 200A may include a control 205 (also referred to as a first control). The control 205 may be used to obtain a reference image specified by a user.
[0039] In some examples, after receiving a click on control 205, electronic device 110 may present a media selection interface to support the user in selecting a picture or a video from a media library (eg, a local photo album of electronic device 110) as a reference image.
[0040] As another example, the electronic device 110 may also present an image capturing interface based on a click on the control 205 , and may determine the captured photo or video as a reference image.
[0041] In some scenarios, in order to better realize the driving of the reference image, the electronic device 110 or the server 130 may also determine whether the acquired reference image includes a specific object. As an example, such a specific object may include a facial object of a person or an animal.
[0042] If no facial object is detected from the reference image, the electronic device 110 may present a prompt message to instruct the user to upload or take a reference image again.
[0043] Additionally, if Figure 2A As shown, the configuration interface 200A may also provide a set of preset models, for example, model 210 and model 215. In some scenarios, different models may correspond to different video generation qualities.
[0044] As an example, model 210 may support global driving of the reference image. That is, model 210 may not only generate animations associated with the facial region, but also cooperatively generate animations associated with other regions. In contrast, model 215 may, for example, only support facial driving of the reference image, so that the facial region in the reference image presents the associated animation.
[0045] In some embodiments, model 210 and model 215 may correspond to different model architectures. For example, different models may support the generation of videos of different lengths or different resolutions. Alternatively, model 210 may also support more types of drive signals than model 215. As an example, model 210 may also support the user to input a prompt word to generate an animation corresponding to a non-facial area.
[0046] Continue to refer Figure 2A The configuration interface 200A may further include a control 220 (also referred to as a second control). The control 220 may be used to receive a reference text or a reference audio.
[0047] As an example, the control 220 may include a text input area, and the user may enter text content as reference text in the text input area of the control 220. In some examples, the text input area may have a preset upper limit on the number of characters to allow receiving text content that does not exceed the upper limit.
[0048] In some embodiments, such text content may also be referred to as "line performance content", which may correspond to the text content to be read aloud. For example, if the text content input by the user includes "The weather is good today", the character in the reference image is driven to show the movement (or animation) corresponding to reading "The weather is good today".
[0049] In some examples, the electronic device 110 may also support the user to upload reference audio through the control 220. As an example, the electronic device 110 may receive a user click on the button 225 in the control 220, and may provide options corresponding to one or more ways of uploading the reference audio.
[0050] In some examples, the electronic device 110 may support the user to record or upload a piece of media content, also referred to as the first media content, by clicking button 225. As an example, the first media content may include video content or audio content. Further, the electronic device 110 may determine the reference audio based on the acquired first media content.
[0051] In some embodiments, the electronic device 110 may also support the user to configure the reference audio by uploading a link. As an example, the electronic device 110 may receive link information input by the user, and the link information may correspond to the second media content. As an example, the second media content may be video content or audio content.
[0052] Further, similar to the first media content, the electronic device 110 may determine the reference audio based on the acquired second media content.
[0053] For example, when the length of the first media content or the second media content is less than a preset length, the audio portion corresponding to the first media content or the second media content can be determined as the reference audio. Conversely, when the length of the first media content or the second media content reaches the preset length, the electronic device 110 can automatically cut the first media content, or the electronic device 110 can support the user to cut the first media content or the second media content to determine the reference audio.
[0054] Specifically, after acquiring the first media content or the second media content, the electronic device 110 may present an audio editing interface, which may provide an audio trimming control to support trimming corresponding segments from the first media content or the second media content.
[0055] In some examples, the electronic device 110 may determine the start position and end position of the desired audio segment based on the user's operation of the audio trimming control, and may determine the trimmed audio segment as the reference audio.
[0056] In some embodiments, in order to facilitate users to perform more efficient cropping, the electronic device 110 can also be associated with the audio cropping control to present the text corresponding to the first media content or the second media content to help the user more intuitively perceive the text content corresponding to the retained audio segment.
[0057] by Figure 2B As an example, after acquiring the reference audio, the electronic device 110 may present a waveform diagram of the reference audio in the control 220, and may support the user to play the configured reference audio by clicking a play button. Additionally, the electronic device 110 may also support the user to delete the configured reference audio by clicking a delete button on the right, for example.
[0058] Continue to refer Figure 2A The configuration interface 200A may also include a tone selection control (also referred to as a third control). Figure 2A As shown, the timbre selection control can be used to determine the target audio parameters to be used. As an example, such target audio parameters can include timbre parameters. Additionally, such target audio parameters can also include parameters such as pitch, rhythm, etc.
[0059] like Figure 2A As shown, the timbre selection control may provide a plurality of candidate audio parameters, for example, timbre 235-1, timbre 235-2, and timbre 235-3. As an example, such candidate audio parameters may include audio parameters preset in the application 120. As another example, in the case of obtaining user authorization, such timbre may also include a timbre model constructed in advance based on a user configuration operation (e.g., an audio recording operation).
[0060] In some examples, the timbre selection control may also provide an upload entry 230 to support the user to determine the corresponding audio parameters by recording or uploading an audio segment. As an example, after clicking the upload entry 230, the electronic device 110 may present an audio recording interface and may provide a preset text and recording controls.
[0061] Furthermore, the user may record the audio content corresponding to the preset text by clicking the recording control, so as to determine the audio parameters to be used.
[0062] In some scenarios, when reference audio is acquired via the control 220 , the timbre selection control may also support, for example, multiplexing audio parameters of the reference audio.
[0063] Further, after acquiring the above configuration information, the electronic device 110 may receive a user click on the generate button 240 to trigger generation of video content based on the above configuration information.
[0064] In some embodiments, the video content may be generated using a generative model deployed at the electronic device 110 or the server 130 .
[0065] Specifically, the video content generated by the model may include a first motion associated with a first region of the reference image to be driven. Specifically, the first region may include a facial region in the reference image, and the first motion may include an animation corresponding to the facial region and a reading of a reference text or a reference audio, such as a mouth movement.
[0066] Additionally, the video content may also include a second motion associated with a second region of the reference image. In some scenarios, the second region may include other appropriate regions different from the facial region, such as a background region of a non-human object or an animal object. Such a background region may, for example, be driven to present a second motion that cooperates with the first motion.
[0067] In some embodiments, the facial region is associated with a preset object (e.g., a human object or an animal object) in the reference image, and the second region includes a limb region of the preset object. That is, the video content generated by the model includes not only facial animations of the human or animal, but also limb animations of the human or animal.
[0068] Additionally, the video content may also include audio content corresponding to the configured target audio parameter (e.g., timbre parameter). As an example, when the configuration information indicates that the reference text is "The weather is good today" and the target audio parameter is "timbre 1", the generated video content may include audio content read aloud in "timbre 1" (i.e., "The weather is good today").
[0069] Additionally, when the configuration information indicates that the reference audio is an audio segment corresponding to “How is the weather today?” and the target audio parameter is “Timbre 2,” the generated video content may include audio content read in “Timbre 2” (ie, “How is the weather today?”).
[0070] In this way, embodiments of the present disclosure can achieve global driving for reference images, thereby improving the quality of generated video content.
[0071] In some embodiments, the electronic device 110 may also receive a user's request to publish the provided video content, and may display the video content accordingly. Figure 2C Publishing interface 200C is shown.
[0072] like Figure 2C As shown, publishing interface 200C can present generated video content 245 and can provide one or more configuration items, for example, configuration item 250, configuration item 255, and configuration item 260. In some embodiments, configuration items 250 to 260 can be used to set which parts of the configuration information used to generate video content 245 can be made public.
[0073] For example, configuration item 250 can be used to set whether to disclose the reference image; configuration item 255 can be used to set whether to disclose the reference audio; and configuration item 260 can be used to set whether to disclose the used timbre.
[0074] Further, upon receiving a selection for the publish button 265, the electronic device 110 may trigger the publishing of the generated video content 245. As an example, the published video content 245 may be associated with Figure 2D The viewing interface 200D is shown.
[0075] like Figure 2D As shown, in the viewing interface 200D, the electronic device 110 can present description information associated with the video content 245. As an example, the electronic device 110 can display the description words corresponding to the reference image. As another example, the electronic device 110 can present the read text 275 corresponding to the video content 245.
[0076] Additionally, the viewing interface 200D may also include a video generation portal 280 (eg, to make the same model). Other users may initiate a video generation request by clicking on the video generation portal 280 and may access the video generation portal 280. Figure 2A The configuration interface shown.
[0077] When the configuration interface is accessed through the video generation portal 280, the configuration interface may present at least part of the configuration information for generating the video content 245 by default. As an example, such at least part may be determined based on the configuration items 250 to 260.
[0078] For example, when the user chooses to disclose the “reference image”, the configuration interface triggered by the video generation entry 280 may present the reference image corresponding to the video content 245 in the first control by default.
[0079] Similarly, when the user chooses to disclose the “reference audio”, the configuration interface triggered and displayed by the video generation entry 280 may provide the reference audio corresponding to the video content 245 in the second control by default.
[0080] Similarly, when the user chooses to disclose “my timbre”, the configuration interface triggered by the video generation entry 280 can provide a timbre selection entry corresponding to the video content 245 in the timbre selection control by default.
[0081] In this way, the embodiments of the present disclosure can support other users to quickly create similar video content, thereby improving the efficiency of video creation.
[0082] Example 2
[0083] The following will refer to Figure 3 To describe an example interaction process according to an embodiment of the present disclosure. As an example, FIG. 2A to FIG. 2D It can correspond to the interactive interface of the mobile terminal. Figure 3 The interface 300 shown may correspond to an interface provided by a personal computer, for example.
[0084] like Figure 3 As shown, interface 300 is also referred to as a configuration interface, which may include multiple controls similar to interface 200A. As an example, control 305 may be used to obtain a reference image. Interface 300 may also provide a plurality of generated models of different qualities for selection, for example, model 310, model 315, and model 320.
[0085] Additionally, similar to control 220, control 330 can be used to obtain reference text or reference audio. Control 335 can be used to specify target audio parameters (e.g., timbre). Additionally, control 340 can also be used to adjust the reading speed. Similarly, after obtaining configuration information via interface 300, electronic device 110 can receive a click on generate button 345 and can trigger the generation of video content accordingly.
[0086] Generate video content
[0087] In some embodiments, the electronic device 110 or the server 130 may determine an audio control signal based on a reference text or a reference audio, and may provide a reference image and the audio control signal to a generative model to generate video content.
[0088] As an example, when the configuration information includes reference text and timbre parameters, the audio generation model may be used to generate corresponding audio content as an audio control signal.
[0089] In some embodiments, in response to the length of the reference audio exceeding a threshold, the reference audio may be segmented to retain audio segments that meet the time length requirement. As an example, the server 130 may determine the corresponding highlight segment based on audio information (e.g., volume, rhythm, etc.) and / or text information of the reference audio, and may generate an audio control signal based on the highlight segment to drive the reference image.
[0090] In some embodiments, the model may include multiple attention layers. The attention layer may include a multimodal attention unit, which may have three input channels. Specifically, the first channel of the multimodal attention unit may correspond to the input video feature.
[0091] The model may be a generative model based on a diffusion model. Accordingly, a video token may be obtained by encoding a preset video content, and noise may be superimposed on the video token. As will be described below, the preset video content may include, for example, other video content associated with the video content to be generated, for example, a portion of video frames in a previous video clip that is continuous in time.
[0092] Furthermore, the skeleton information may also be determined based on the reference video content. As an example, the skeleton information 225 may represent a set of key points corresponding to a preset object of the reference video content.
[0093] Further, the skeleton information may be processed by the encoding unit to determine the skeleton features. As an example, the skeleton features may have the same feature size as the video token. Further, the video token and the skeleton features may be concatenated in the channel dimension to obtain the input video features. As an example, the encoding unit may include, for example, a convolutional neural network or other suitable network.
[0094] In addition, the multimodal attention unit may further include a second channel. The second channel may correspond to an input text feature. As an example, in the case where the control signal includes reference text content (e.g., a prompt word), the input text feature may include a text token determined by encoding the reference text content.
[0095] Additionally, the multimodal attention unit may further include a third channel. The third channel may correspond to an input image feature. As an example, the input image feature may include a reference token determined by encoding a reference image.
[0096] Further, depending on the modality type of the control signal, the multimodal attention unit may perform an attention-based update process on at least two types of input features received via multiple input channels. For example, in the case where the control signal includes reference video content, the multimodal attention unit may update the input video features and the input image features based on the multimodal attention mechanism.
[0097] Further, the attention layer may also include a cross attention unit. The cross attention unit may be associated with two input channels. For example, the cross attention unit may include a fifth channel corresponding to the input video features updated via the multimodal attention unit.
[0098] In addition, the cross-attention unit may further include a sixth channel. The sixth channel may correspond to the input audio features. The input audio features may include audio tokens generated based on the reference audio content.
[0099] Specifically, in the case where the control signal includes reference audio content, audio features of the reference audio content may be determined. As an example, the audio features may include Mel-spectrogram features of the audio content. Further, the audio features may be processed using a coding unit to obtain an audio token. As an example, the coding unit may include, for example, a multilayer perceptron (MLP) or other appropriate models.
[0100] Further, the cross attention unit can update the input video features based on the cross attention mechanism and based on the audio token. Further, the updated input video features can be provided to the multimodal attention unit in the next attention layer.
[0101] Similarly, the input text features and input image features updated by the multimodal attention unit in the attention layer can also be further provided to the multimodal attention unit in the next attention layer. Similarly, the multimodal attention unit can update the input video features, input text features, and input image features accordingly.
[0102] In this way, after processing through multiple attention layers, the model can output the final video features to determine the target video.
[0103] Example Process
[0104] Figure 4FIG. 4 is a flowchart showing an example process 400 for video generation according to some embodiments of the present disclosure. The process 400 may be implemented at the electronic device 110. Figure 1 4. The process 400 is described below.
[0105] As shown in the figure, in box 410, in response to receiving a video generation request, the electronic device 110 presents a configuration interface, where the configuration interface includes a first control and a second control.
[0106] In box 420, the electronic device 110 obtains configuration information via the configuration interface, where the configuration information includes: a reference image obtained via the first control, and a reference text or reference audio obtained via the second control, wherein the reference image includes a preset object.
[0107] In box 430, the electronic device 110 provides video content generated based on the configuration information, the video content includes a first motion associated with a first area of the reference image and a second motion associated with a second area of the reference image, the first area includes a facial area, and the first motion corresponds to a reference text or a reference audio.
[0108] In some embodiments, the facial region is associated with a preset object in the reference image, and the second region includes a limb region of the preset object.
[0109] In this way, the embodiments of the present disclosure can achieve not only facial drive but also body drive, thereby improving the quality of the generated video.
[0110] In some embodiments, obtaining configuration information via the configuration interface includes: obtaining text content input in the second control as reference text.
[0111] In this way, the embodiments of the present disclosure can support the user to input text that needs to be read aloud, thereby achieving control over the generation of video content.
[0112] In some embodiments, obtaining configuration information via the configuration interface includes: obtaining recorded or uploaded first media content via a second control; and determining reference audio based on the first media content.
[0113] In this way, the embodiments of the present disclosure can support the user to input reference audio for driving, thereby achieving control over the generation of video content.
[0114] In some embodiments, based on the first media content, determining the reference audio includes: in response to acquiring the first media content, presenting an audio editing interface for editing the first audio content of the first media content; and determining, via the audio editing interface, an audio segment of the first audio content as the reference audio.
[0115] In this way, embodiments of the present disclosure can support users in selecting audio clips, thereby improving the quality of generated video content.
[0116] In some embodiments, obtaining the reference audio via the second control includes: obtaining link information corresponding to the second media content via the second control; and determining the reference audio based on the second media content.
[0117] In this way, the embodiments of the present disclosure can support the user to specify reference audio through linking, thereby improving the efficiency of video generation.
[0118] In some embodiments, the configuration interface further includes a third control, and obtaining the configuration information via the configuration interface further includes: determining a target audio parameter via the third control, wherein the generated video content includes a second audio content corresponding to the target audio parameter.
[0119] In this way, embodiments of the present disclosure may further support configuration of audio properties, thereby improving the quality of generated video content.
[0120] In some embodiments, determining the audio parameter via the third control includes: providing a plurality of candidate audio parameters in the third control; and receiving a selection of a target audio parameter from among the plurality of candidate audio parameters.
[0121] In this way, embodiments of the present disclosure may provide preset audio parameters to reduce the user's learning cost.
[0122] In some embodiments, determining the audio parameter via the third control includes: acquiring recorded or uploaded second audio content via the third control; and determining the target audio parameter based on the second audio content.
[0123] In this way, embodiments of the present disclosure may support users to customize audio parameters, thereby improving the quality of video generation.
[0124] In some embodiments, process 400 further includes: in response to receiving a publishing request for video content, publishing the video content; and presenting a video generation entry in a viewing interface of the video content, wherein the video generation entry is configured to trigger a display configuration interface.
[0125] In this way, the embodiments of the present disclosure can support other users to generate videos more conveniently.
[0126] In some embodiments, the video generation portal is configured to trigger presentation of at least a portion of the configuration information in the configuration interface, wherein at least a portion of the configuration information is determined based on the publish request.
[0127] In this way, the embodiments of the present disclosure can support other users to reuse the configuration information generated by the video, thereby lowering the creation threshold.
[0128] In some embodiments, in response to the length of the reference audio exceeding a threshold, the first motion of the video content corresponds to a target segment of the reference audio, the target segment being determined from the reference audio based on audio information and / or text information of the reference audio.
[0129] In this way, embodiments of the present disclosure may improve the quality of audio clips used to drive video generation.
[0130] In some embodiments, the video content is generated based on the following process: determining an audio control signal based on a reference text or a reference audio; and providing the reference image and the audio control signal to a generative model to generate the video content.
[0131] In this way, embodiments of the present disclosure can achieve efficient video driving.
[0132] Example devices and equipment
[0133] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 5 A schematic structural block diagram of an example apparatus 500 for video generation according to some embodiments of the present disclosure is shown. The apparatus 500 may be implemented as or included in the electronic device 110. Each module / component in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0134] like Figure 5 As shown, the device 500 includes a presentation module 510, which is configured to present a configuration interface in response to receiving a video generation request, wherein the configuration interface includes a first control and a second control; an acquisition module 520, which is configured to acquire configuration information via the configuration interface, wherein the configuration information includes: a reference image acquired via the first control, and a reference text or reference audio acquired via the second control, wherein the reference image includes a preset object; and a providing module 530, which is configured to provide video content generated based on the configuration information, wherein the video content includes a first motion associated with a first area of the reference image and a second motion associated with a second area of the reference image, wherein the first area includes a facial area, and the first motion corresponds to the reference text or reference audio
[0135] In some embodiments, the facial region is associated with a preset object in the reference image, and the second region includes a limb region of the preset object.
[0136] In some embodiments, the acquisition module 520 is further configured to: acquire text content input in the second control as reference text.
[0137] In some embodiments, the acquisition module 520 is further configured to: acquire, via the second control, the recorded or uploaded first media content; and determine the reference audio based on the first media content.
[0138] In some embodiments, the acquisition module 520 is further configured to: in response to acquiring the first media content, present an audio editing interface for editing the first audio content of the first media content; and determine an audio segment of the first audio content as a reference audio via the audio editing interface.
[0139] In some embodiments, the acquisition module 520 is further configured to: acquire link information corresponding to the second media content via the second control; and determine reference audio based on the second media content.
[0140] In some embodiments, the configuration interface further includes a third control, and the acquisition module 520 is further configured to: determine a target audio parameter via the third control, wherein the generated video content includes a second audio content corresponding to the target audio parameter.
[0141] In some embodiments, the acquisition module 520 is further configured to: provide a plurality of candidate audio parameters in a third control; and receive a selection of a target audio parameter from among the plurality of candidate audio parameters.
[0142] In some embodiments, the acquisition module 520 is further configured to: acquire, via a third control, a second audio content that has been recorded or uploaded; and determine a target audio parameter based on the second audio content.
[0143] In some embodiments, the device 500 also includes a publishing module, which is configured to publish video content in response to receiving a publishing request for video content; and present a video generation entry in the viewing interface of the video content, and the video generation entry is configured to trigger a display configuration interface.
[0144] In some embodiments, the video generation portal is configured to trigger presentation of at least a portion of the configuration information in the configuration interface, wherein at least a portion of the configuration information is determined based on the publish request.
[0145] In some embodiments, in response to the length of the reference audio exceeding a threshold, the first motion of the video content corresponds to a target segment of the reference audio, the target segment being determined from the reference audio based on audio information and / or text information of the reference audio.
[0146] In some embodiments, the video content is generated based on the following process: determining an audio control signal based on a reference text or a reference audio; and providing the reference image and the audio control signal to a generative model to generate the video content.
[0147] The modules included in the device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the modules in the device 500 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0148] Figure 6 1 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 6 The electronic device 600 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 6 The electronic device 600 shown can be used to implement Figure 1 An electronic device 110.
[0149] like Figure 6 As shown, the electronic device 600 is in the form of a general electronic device. The components of the electronic device 600 may include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be an actual or virtual processor and is capable of performing various processes according to a program stored in the memory 620. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 600.
[0150] The electronic device 600 typically includes a plurality of computer storage media. Such media may be any accessible media that is accessible to the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 may be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 may be a removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which may be capable of being used to store information and / or data and may be accessed within the electronic device 600.
[0151] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 6 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 620 may include a computer program product 625 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.
[0152] The communication unit 640 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0153] The input device 650 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 660 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 600, or communicate with any device that allows the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0154] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0155] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0156] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0157] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0158] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0159] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.
Claims
1. A method for video generation, comprising: In response to receiving the video generation request, presenting a configuration interface, the configuration interface including a first control and a second control; Acquiring configuration information via the configuration interface, the configuration information comprising: a reference image acquired via the first control, and a reference text or reference audio acquired via the second control, wherein the reference image comprises a preset object; as well as Provide video content generated based on the configuration information, the video content including a first motion associated with a first area of the reference image and a second motion associated with a second area of the reference image, the first area including a facial area, and the first motion corresponding to the reference text or the reference audio. 2 . The method according to claim 1 , wherein the facial area is associated with a preset object in the reference image, and the second area includes a limb area of the preset object.
3. The method according to claim 1, wherein obtaining configuration information via the configuration interface comprises: The text content input in the second control is obtained as the reference text.
4. The method according to claim 1, wherein obtaining configuration information via the configuration interface comprises: Obtaining, via the second control, first media content that has been recorded or uploaded; as well as The reference audio is determined based on the first media content.
5. The method according to claim 4, wherein based on the first media content, determining the reference audio comprises: In response to acquiring the first media content, presenting an audio editing interface for editing first audio content of the first media content; as well as An audio segment of the first audio content is determined via the audio editing interface to serve as the reference audio.
6. The method according to claim 1, wherein obtaining the reference audio via the second control comprises: Acquiring link information corresponding to the second media content via the second control; as well as Based on the second media content, the reference audio is determined.
7. The method according to claim 1, wherein the configuration interface further comprises a third control, and obtaining configuration information via the configuration interface further comprises: Via the third control, a target audio parameter is determined, wherein the generated video content includes second audio content corresponding to the target audio parameter.
8. The method of claim 7, wherein determining, via the third control, an audio parameter comprises: providing a plurality of candidate audio parameters in the third control; as well as A selection of the target audio parameter from among the plurality of candidate audio parameters is received.
9. The method of claim 7, wherein determining, via the third control, an audio parameter comprises: Obtaining, via the third control, a second audio content that has been recorded or uploaded; as well as Based on the second audio content, the target audio parameter is determined.
10. The method according to claim 1, further comprising: In response to receiving a publishing request for the video content, publishing the video content; as well as A video generation entry is presented in the viewing interface of the video content, and the video generation entry is configured to trigger display of the configuration interface. 11 . The method of claim 10 , wherein the video generation portal is configured to trigger presentation of at least a portion of the configuration information in the configuration interface, wherein the at least a portion is determined based on the publishing request.
12. The method of claim 1, wherein: In response to the length of the reference audio exceeding a threshold, the first motion of the video content corresponds to a target segment of the reference audio, the target segment being determined from the reference audio based on audio information and / or text information of the reference audio.
13. The method of claim 1, wherein the video content is generated based on the following process: determining the audio control signal based on the reference text or the reference audio; and The reference image and the audio control signal are provided to a generative model to generate the video content.
14. An apparatus for video generation, comprising: A presentation module, configured to present a configuration interface in response to receiving a video generation request, wherein the configuration interface includes a first control and a second control; an acquisition module, configured to acquire configuration information via the configuration interface, the configuration information comprising: a reference image acquired via the first control, and a reference text or reference audio acquired via the second control, wherein the reference image comprises a preset object; as well as A providing module is configured to provide video content generated based on the configuration information, the video content including a first motion associated with a first area of the reference image and a second motion associated with a second area of the reference image, the first area including a facial area, and the first motion corresponding to the reference text or the reference audio.
15. An electronic device, comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 13 when executed by the at least one processing unit.
16. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 13 when executed by a processor.
Citation Information
Cited By
Method and device for generating video content, equipment and storage medium
CN120614508A