Method and device for generating video content, equipment and storage medium
By presenting candidate images in the interactive interface of virtual objects and determining video generation information based on user interaction operations, the problem of insufficient video generation quality in the existing technology is solved, and higher quality video content generation is achieved.
Patent Information
- Application Number
- CN202510847350.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-23
AI Technical Summary
In the prior art, how to improve the quality of video content generated based on text or images input by users is a focus of attention.
Through the interactive interface with the virtual object, the user input content is obtained, and the candidate images are presented in the interface. The video generation information is determined based on the interaction between the user and the candidate images, and finally high-quality video content is generated.
The quality of video generation information and content is improved, making the generated video more in line with user expectations, enhancing the interaction efficiency between users and virtual objects and the effect of video generation.
Smart Images

Figure CN120614508A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for generating video content. Background Art
[0002] With the development of computer technology, more and more platforms and applications are providing video generation services for users. For example, some platforms and applications can generate corresponding video content based on user-entered text or images. Improving the quality of generated media content is a key issue. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for generating video content is provided. The method comprises: obtaining user input content via an interactive interface with a virtual object; presenting first response content from the virtual object in the interactive interface, the first response content including at least one candidate image generated based on the input content; determining video generation information based on an interactive operation associated with the first response content, the video generation information including at least a reference image for generating a video; and providing second response content from the virtual object, the second response content including video content generated based on the video generation information.
[0004] In a second aspect of the present disclosure, a device for generating video content is provided. The device includes: an acquisition module configured to acquire user input content via an interactive interface with a virtual object; a first presentation module configured to present, in the interactive interface, first response content from the virtual object, the first response content including at least one candidate image generated based on the input content; a first determination module configured to determine video generation information based on an interactive operation associated with the first response content, the video generation information including at least a reference image for generating a video; and a provision module configured to provide second response content from the virtual object, the second response content including video content generated based on the video generation information.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the present disclosure, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions, which, when executed by a device, cause the device to perform the method of the first aspect.
[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram illustrating an example environment in which embodiments according to the present disclosure may be implemented;
[0011] Figures 2A to 2J shows an example interface according to some embodiments of the present disclosure;
[0012] Figure 3 A flowchart illustrating an example process for generating video content according to some embodiments of the present disclosure;
[0013] Figure 4 A schematic structural block diagram illustrating an example apparatus for generating video content according to some embodiments of the present disclosure is shown; and
[0014] Figure 5 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0015] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0016] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.
[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0018] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0019] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.
[0020] As mentioned above, with the development of computer technology, more and more platforms or applications are providing video generation services for users. For example, some platforms or applications can generate corresponding video content based on user-entered text or user-entered images. How to improve the quality of generated media content is a key issue.
[0021] An embodiment of the present disclosure proposes a scheme for generating video content. According to the scheme, the user's input content can be obtained via an interactive interface with a virtual object. Furthermore, a first response content from the virtual object can be presented in the interactive interface, and the first response content includes at least one candidate image generated based on the input content. In addition, video generation information can be determined based on an interactive operation associated with the first response content, and the video generation information includes at least a reference image for generating a video. Additionally, a second response content from the virtual object can be provided, and the second response content includes video content generated based on the video generation information.
[0022] Based on this approach, embodiments of the present disclosure can first provide at least one candidate image generated based on user input before providing video content from a virtual object. This allows the user's interaction with the at least one candidate image to improve the quality of the video generation information used for video generation. Thus, embodiments of the present disclosure can generate higher-quality video content based on higher-quality video generation information.
[0023] Therefore, the embodiments of the present disclosure can present at least one candidate image generated based on input content to the user during the video generation process, thereby improving the quality of video generation based on the user's interactive operation with the at least one candidate image.
[0024] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.
[0025] Sample Environment
[0026] Figure 1 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, example environment 100 may include electronic device 110 .
[0027] In this example environment 100, electronic device 110 may run an application 120 that supports the generation of video content. Application 120 may be any suitable type of application for generating video content, examples of which may include, but are not limited to, media applications, social applications, or other suitable applications. User 140 may interact with application 120 via electronic device 110 and / or its attached devices.
[0028] exist Figure 1 In the environment 100 , if the application 120 is active, the electronic device 110 may present an interface 150 for supporting the generation of video content through the application 120 .
[0029] In some embodiments, the electronic device 110 communicates with the server 130 to enable the provision of services for the application 120. The electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.).
[0030] The server 130 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The server 130 may include, for example, a computing system / server such as a mainframe, an edge computing node, a computing device in a cloud environment, etc. The server 130 may provide background services for the application 120 that supports content presentation in the electronic device 110.
[0031] A communication connection may be established between the server 130 and the electronic device 110. The communication connection may be established in a wired or wireless manner. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the server 130 and the electronic device 110 may implement signaling interaction through the communication connection between the two.
[0032] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0033] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.
[0034] Example Interaction
[0035] The following will refer to Figures 2A to 2J To describe some example interaction processes according to embodiments of the present disclosure. Figures 2A to 2J Example interfaces 200A to 200J according to some embodiments of the present disclosure are shown. Interfaces 200A to 200J may be, for example, Figure 1 The electronic device 110 is provided as shown.
[0036] In some embodiments, the electronic device 110 may present Figure 2A The interface 200A shown in FIG. 200A may include an interface for generating video content. Figure 2A As shown, the electronic device 110 can present an interface 200A for the user 140 to interact with a virtual object (e.g., virtual object 210), and generate video content through the interaction between the user 140 and the virtual object. In some scenarios, the interface 200A may include an interactive interface for the user 140 to interact with the virtual object. It should be understood that, on the one hand, the virtual object can have a virtual image for interaction. On the other hand, the virtual object can also use an appropriate model to drive its interactive behavior, and such a model can include any appropriate machine learning model, for example, a generative model. In the present disclosure, the various interactive processes and / or generation operations performed by the virtual object can be actually performed using an appropriate model corresponding to the virtual object.
[0037] In some scenarios, such an interactive interface may, for example, include a one-on-one chat interface between user 140 and a virtual object, or a group chat interface with at least user 140 and the virtual object as participants. In other scenarios, such a conversation interface may also include a conversational search interface associated with the virtual object, which may, for example, search for content based on a conversation between user 140 and the virtual object, and then present search results in the form of a reply message from the virtual object. In some other scenarios, such an interactive interface may also have other appropriate forms. For example, the interactive interface may correspond to an interactive space for a virtual object, which may, for example, support multiple modal conversational interactions between user 140 and the virtual object.
[0038] In some scenarios, user 140 can trigger input of some content to instruct the virtual object to generate corresponding video content based on the input content of user 140. For example, user 140 can trigger the sending of message 221 (for example, including displaying an image of a small ball) and message 222 (for example, including text such as "turn the image into multiple small balls and make it move"). In some examples, such messages 221 and 222 can be referred to as user input content. In some examples, the user's input content can include at least one of text, images, and the like. Thus, electronic device 110 can provide video content generated by the virtual object.
[0039] In order to improve the quality of the generated video content, in some embodiments of the present disclosure, the electronic device 110 may present first response content from the virtual object in the interactive interface upon obtaining the input content of the user 140, and such first response content includes at least one candidate image generated based on the input content of the user 140.
[0040] Furthermore, the electronic device 110 may determine video generation information based on the user 140's interactive operation on the content 231. Such video generation information may include at least a reference image for generating the video. Thus, the electronic device 110 may provide second response content from the virtual object. Such second response content may include video content generated based on such video generation information.
[0041] Based on this approach, the embodiment of the present disclosure can improve the quality of video generation information through interactive operations related to at least one candidate image, thereby improving the quality of generated video content based on such video generation information. Figure 2B To the attached Figure 2F The following describes in detail an example where at least one candidate image may include an image.
[0042] In some embodiments, according to Figure 2B In the interface 200B shown, the electronic device 110 can provide a virtual object reply content 231 (also known as a first response content) to the message 221 and the message 222. Such content 231 may include at least one candidate image (for example, image 232) generated based on the message 221 and the message 222.
[0043] As an example, such content 231 may include the reply text "I will add multiple small balls to this picture" to message 222. Furthermore, electronic device 110 may present image 232 generated based on messages 221 and 222. Such image 232 may include multiple small balls. Based on such image 232, user 140 may quickly view the visual effects of the video content to be generated. Furthermore, user 140 may interact with content 231 to trigger an interaction operation based on content 231 to determine video generation information. Such video generation information may include at least a reference image for generating the video.
[0044] As some examples, such an image 232 can be generated using a model. For example, the electronic device 110 can input the message 221 and the message 222 (or the prompt item generated based on the message 222) into the model to obtain the image 232 generated by the model. As an example, such a model can include any suitable image generation model, for example, a generative model, which can output an image that meets the prompt item based on the input image and prompt item. As an example, such a model can be deployed in the electronic device 110 and / or the server 130. As an example, such a model can also be deployed in other terminals or servers. Furthermore, the electronic device 110 can call such a model through a communication connection with other terminals or servers. Thus, the embodiments of the present disclosure can improve the quality of the image 232 and the image generation efficiency.
[0045] For example, user 140 may confirm image 232 to use image 232 as a reference image. In some examples, user 140 may also trigger a first regeneration request associated with image 232 to be sent to the virtual object. For example, user 140 may trigger a message such as "Adjust XXX image" or "Adjust XX image using XXX method" to be sent to the virtual object to trigger the first regeneration request to be sent to the virtual object.
[0046] Furthermore, electronic device 110 may present at least one updated image generated based on such a first regeneration request. Thus, user 140 may confirm the second image in such at least one updated image to determine the confirmed second image as the reference image. Alternatively or additionally, user 140 may also trigger a regeneration request associated with such at least one updated image to further trigger image regeneration, and so on. The embodiments of the present disclosure are not exhaustive in this regard.
[0047] In some embodiments, the virtual object can provide user 140 with more information for generating a video, further helping the user to efficiently adjust the video generation information. For example, such first response content can also include at least one candidate prompt item for generating a video. Thus, the determined video generation information can also include a target prompt item determined based on such at least one candidate prompt item.
[0048] For example, according to Figure 2B In the illustrated interface 200B, the message 222 may also include a prompt item 234 (also referred to as a candidate prompt item). Such a prompt item 234 may indicate the video effect of the video content to be generated. Thus, the user 140 can quickly understand the direction of the video content generation, thereby enabling rapid adjustment of the effect of the generated video content.
[0049] Further, user 140 can confirm such prompt item 234 to determine such prompt item 234 as the target prompt item. In some examples, prompt item 234 can indicate "let the ball move up and down". Thus, user 140 can confirm such a video effect of "let the ball move up and down" to trigger the generation of video content from the virtual object. As an example, user 140 can trigger the confirmation video effect by triggering the sending of a confirmation message. In some examples, user 140 may not perform an interactive operation as such a confirmation operation. For example, if no interactive operation from user 140 is received within the preset time period when message 231 is presented, electronic device 110 can confirm the generation of a video corresponding to the video effect of "let the ball move up and down".
[0050] Alternatively, user 140 may trigger editing of such prompt item 234 to determine the edited prompt item 234 as the target prompt item. For example, user 140 may trigger sending of a message indicating "Move the ball left and right" to determine "Move the ball left and right" as the target prompt item. For another example, user 140 may trigger sending of a message indicating "Move the ball forward and backward" to determine "Move the ball forward and backward" as the target prompt item.
[0051] Alternatively or additionally, if the first response content includes an image, the first response content may also include multiple candidate prompt items. Furthermore, user 140 may select a first prompt item from these multiple candidate prompt items to trigger the selection of the selected first prompt item as the target prompt item. Alternatively, user 140 may also trigger editing of a second prompt item from these multiple candidate prompt items to trigger the selection of the edited second prompt item as the target prompt item.
[0052] In some embodiments, user 140 can also trigger a second regeneration request associated with at least one such candidate prompt. Furthermore, electronic device 110 can present at least one updated prompt generated based on the second regeneration request. As an example, user 140 can trigger a second regeneration request associated with prompt 234. For example, the user can trigger the sending of a message indicating "re-adjust the ball's motion," "make the ball move in other ways," and so on. Thus, electronic device 110 can present corresponding response content from the virtual object to present at least one updated prompt generated based on the message "re-adjust the ball's motion," "make the ball move in other ways," and so on.
[0053] Thus, user 140 can confirm the third prompt item in the at least one updated prompt item to determine the confirmed second prompt item as the reference prompt item. Alternatively or additionally, user 140 can also trigger a regeneration request associated with the at least one updated prompt item to further trigger the regeneration of prompt items, and so on. The embodiments of the present disclosure are not exhaustive. Based on this approach, the embodiments of the present disclosure can support users to flexibly and efficiently adjust the candidate images and / or candidate prompt items included in the first response content, thereby improving the quality of video generation and improving user interaction efficiency.
[0054] In some embodiments, in order to improve the interaction efficiency between the virtual object and the user 140, before generating the image 232, the electronic device 110 may provide a message from the virtual object and present a placeholder element indicating the image generation status in such a message. Figure 2C In the illustrated interface 200C, the electronic device 110 may present a placeholder element 233 to indicate an image generation status of "image generating" and may present the progress of the image generation (e.g., as a percentage, etc.). Thus, embodiments of the present disclosure can support user interaction with a virtual object before generating the image 232.
[0055] In some examples, according to Figure 2D In the illustrated interface 200D, the user 140 may trigger sending a message 223 (e.g., "Generate video") to trigger confirmation of generating a video corresponding to the video effect of "making the ball move up and down" based on the image 232. As a result, the electronic device 110 may provide a message 235 (also referred to as second response content) from the virtual object, and such message 235 may include video content generated based on the video generation information.
[0056] In some examples, before the video generation is complete, the electronic device 110 may continue to present the image 232 in the message 235. Furthermore, when the video generation is complete, the electronic device 110 may replace the image 232 in the message 235 with the corresponding video. In this manner, embodiments of the present disclosure can help the user 140 more conveniently adjust the video effects of the video to be generated.
[0057] In some embodiments, during the video generation process, the user 140 can continue to interact with the virtual object. Further, when the video generation is completed, the electronic device 110 can provide a reminder message from the virtual object to indicate that such video generation is complete. Figure 2E In the interface 200E shown, the electronic device 110 may provide a reminder message 236 from the virtual object, and such a reminder message 236 may indicate that the task of generating a video based on the image 232 has been completed.
[0058] In some examples, such a reminder message 236 may include reference content 237 for message 222 (for example, a reference to “turn the image into multiple small balls and make it move”). Such reference content 237 can help the user 140 understand such a reminder message 236 more efficiently. Furthermore, such a reminder message 236 may include a viewing entry 238 (for example, “view generated content”). The user 140 can trigger the presentation of the generated video content by triggering the viewing entry 238. Thus, the embodiments of the present disclosure can support the user to quickly view the generated video content through the viewing entry in the reminder message during the interaction with the virtual object, thereby improving the interface interaction efficiency between the user and the interactive interface.
[0059] As an example, according to Figure 2F In the illustrated interface 200F, user 140 can trigger viewing entry 238 to slide a message in the conversation interface to present message 235. For example, electronic device 110 can slide such message 235 to a preset position in the conversation interface for presentation. In some examples, message 235 can include generated video content 240. Such video content 240 can be generated based on image 232 and have a "small ball moving up and down" video effect. In some examples, user 140 can trigger playback of such video content 240 in the interactive interface by triggering play control 241, thereby improving interface interaction efficiency.
[0060] In some examples, such video content 240 can be generated using a model. For example, the electronic device 110 can input the image 232 and the prompt item 234 (or other prompt items generated based on the prompt item 234) into the model to obtain the video content 240 generated by the model. As an example, such a model can include any suitable image generation model, for example, a generative model, which can output a video that meets the prompt item based on the input image and prompt item. As an example, such a model can be deployed in the electronic device 110 and / or the server 130. As an example, such a model can also be deployed in other terminals or servers. Then, the electronic device 110 can call such a model by being connected to the communication with other terminals or servers. Thus, the embodiments of the present disclosure can improve the quality of the video content 240.
[0061] In some embodiments, based on the message 221 and the message 222, the electronic device 110 may provide multiple candidate images from the virtual object to improve the richness of the interaction. Figure 2G To describe such at least one candidate image includes multiple specific implementations.
[0062] As an example, according to Figure 2G In the interface 200G shown, the electronic device 110 may provide a message 251 from the virtual object. Such a message 251 may be a response message from the virtual object to the message 221 and the message 222. Figure 2G As shown, such message 251 may include multiple generated candidate images (e.g., image 252-1, image 252-2, etc., which may be individually or collectively referred to as images 252). As an example, the generation process of such image 252 can refer to the generation process of image 232 described above, and will not be repeated here. It should be understood that the present disclosure is not intended to limit the number of multiple generated images.
[0063] As an example, such an image 252-1 may be an image of adding three small balls to the message 221, and such an image 252-2 may be an image of adding two small balls to the message 221. Furthermore, the user 140 may select a target image from the plurality of images 252 to trigger the generation of a video corresponding to the target image. In some examples, the electronic device 110 may generate a corresponding video based on a specified image (e.g., image 251-1, etc.) when the user 140 does not perform an interactive operation within a specified period of time (e.g., within a specified duration of presenting the message 251). Based on this approach, the embodiments of the present disclosure can improve the richness of video generation and can further improve the quality of video generation.
[0064] In some embodiments, the user 140 can select one or more images from such images 252 as one or more reference images for video generation. Furthermore, such one or more reference images can be used to determine one or more video frames of the video content. As an example, the user 140 can trigger the determination of image 252-1 as a reference image for user video generation. Furthermore, the electronic device 110 can provide video content from a virtual object with image 252-1 as one video frame. As an example, the user 140 can trigger the determination of image 252-1 and image 252-2 as multiple reference images for user video generation. Furthermore, the electronic device 110 can provide video content from a virtual object with image 252-1 and image 252-2 as video frames. As a result, the embodiments of the present disclosure can improve the quality of video generation.
[0065] In some embodiments, different images 252 may be associated with different prompts to associate with different video effects. Figure 2GAs shown, image 252-1 may be associated with prompt 253-1 "The ball moves up and down," and image 252-2 may be associated with prompt 253-2 "The ball moves left and right." In some examples, electronic device 110 may generate a video based on image 252-1 and prompt 253-1. In other examples, a video may be generated based on image 252-2 and prompt 253-2.
[0066] In some examples, electronic device 110 can adjust such a video generation process based on a message from user 140. As an example, user 140 can trigger the generation of a video based on image 252-1 and prompt item 253-2. As an example, user 140 can also trigger the generation of a video based on image 252-2 and prompt item 253-1. As an example, user 140 can trigger the sending of an adjustment message to customize such a video generation process. For example, user 140 can send a message "Add five small balls" to instruct to regenerate image 252. For another example, user 140 can send a message "Move the small ball in the first image back and forth" to trigger the generation of a video based on image 252-1 and prompt item "Move the small ball back and forth". As a result, the embodiments of the present disclosure can further improve the richness of the generated video.
[0067] In some embodiments, the electronic device 110 can provide multiple videos from the virtual object based on the message from the user 140, thereby, the embodiment of the present disclosure can further improve the quality of video generation. Figure 2H To the attached Figure 2J To further describe some embodiments of the present disclosure.
[0068] For example, according to Figure 2H In the illustrated interface 200H, the electronic device 110 may present a message 224 from the user 140 (e.g., "Show the scenes in the poem "XXX" by author 1, and provide a corresponding video for each line of the poem"). Based on this, the electronic device 110 may provide a video corresponding to each line of the poem "XXX." As an example, the user's input content may include the message 224. In some examples, the user's input content may also include extended information determined based on the message 224. For example, the user's input content may also include the full text of the poem "XXX."
[0069] For example, the poem "XXX" may include "aaaaa", "bbbbb", "ccccc", and "ddddd". Figure 2IIn the interface 200I shown, the electronic device 110 can present a message 261 (also referred to as the first response content) from the virtual object. As an example, the message 261 may include a first set of images generated based on "aaaaa", a second set of images generated based on "bbbbb", and so on. Thus, the electronic device 110 can determine the first video generation information based on the first set of images to generate the first video content (e.g., video A) corresponding to "aaaaa". Such video A can include at least a first reference image for video generation (e.g., an image in the first set of images determined by the user or an image that is regenerated and confirmed by the user after triggering, etc.).
[0070] Alternatively or additionally, the electronic device 110 may determine second video generation information based on the second group of images to generate second video content corresponding to "bbbbb" (e.g., video B), such video B may include at least a second reference image for video generation (e.g., an image in the second group of images determined by the user or an image triggered to be regenerated and confirmed by the user, etc.).
[0071] For ease of explanation, the embodiments of the present disclosure are described using the example of electronic device 110 message 261 including an image generated based on each line of the poem "XXX". The specific implementation of each line of poetry corresponding to multiple images can be referred to the embodiments of the present disclosure and will not be repeated here.
[0072] Based on this, according to Figure 2H In the illustrated interface 200H, such a message 261 may include an image 262-1 corresponding to "aaaaa", an image 262-2 corresponding to "bbbbb", etc. Furthermore, the electronic device 110 may provide a video (e.g., video A) generated based on the image 262-1, a video (e.g., video B) generated based on the image 262-2, etc.
[0073] As an example, the generation process of such image 262 can refer to the detailed description of the generation process of image 232 described above. Alternatively or additionally, the generation process of such videos A, B, etc. can refer to the detailed description of the generation process of video content 240 described above. The embodiments of the present disclosure are not further described here.
[0074] Therefore, the embodiments of the present disclosure can improve the richness of interaction with virtual objects.
[0075] In some embodiments, electronic device 110 may combine videos such as Video A and Video B corresponding to each line of the poem "XXX" to obtain a combined video corresponding to the poem "XXX." For example, electronic device 110 may determine second video generation information based on message 261. Such second video generation information may also be used to generate a combined video corresponding to the poem "XXX." Such a combined video may include at least such first video content (e.g., Video A or Video B, etc.).
[0076] As an example, such video synthesis can be implemented by any suitable video generation tool. For example, such video A, video B, etc. can be spliced to obtain such a merged video. For another example, the electronic device 110 can input such video A, video B, etc. into a video processing model (e.g., a generative model that can synthesize multiple input videos) to obtain the merged video output by the video processing model.
[0077] In some embodiments, such a merged video may also be provided based on a user's merge request. Specifically, electronic device 110 may trigger the merging of videos corresponding to each line of the poem "XXX" based on a merge request from user 140. For example, user 140 may trigger the sending of a message indicating "Merge videos," "Generate complete video," or the like to trigger such video merging.
[0078] As an example, according to Figure 2J In the interface 200J shown, the electronic device 110 can be associated with the message 261 and provide multiple interactive items (e.g., interactive item 281-1, interactive item 281-2, interactive item 281-3, interactive item 281-4, etc., which can be individually or collectively referred to as interactive items 281) to support the user 140 to conveniently interact with the virtual object. As an example, such interactive item 281 can indicate a request to merge multiple videos in the message 261. For example, the user 140 can select the interactive item 281-1 (e.g., "Generate complete video") to trigger a request to send a merged video to the virtual object. As a result, the embodiments of the present disclosure can further improve the richness of video generation.
[0079] It should be understood that the above-mentioned specific content items mentioned in the embodiments of the present disclosure (for example, the specific content of the image, the specific content of the video, the specific content of the prompt item, the specific content of the message, the number of specific objects, the image number, etc.) are only exemplary, and the present disclosure is not intended to limit this.
[0080] Based on this approach, embodiments of the present disclosure can first provide at least one candidate image generated based on user input before providing video content from a virtual object. This allows the user's interaction with the at least one candidate image to improve the quality of the video generation information used for video generation. Thus, embodiments of the present disclosure can generate higher-quality video content based on higher-quality video generation information.
[0081] Therefore, the embodiments of the present disclosure can present at least one candidate image generated based on input content to the user during the video generation process, thereby improving the quality of video generation based on the user's interactive operation with the at least one candidate image.
[0082] Example Process
[0083] Figure 3 FIG. 3 is a flow chart showing an example process 300 for generating video content according to some embodiments of the present disclosure. The process 300 may be implemented at the electronic device 110. Figure 1 Process 300 is described.
[0084] like Figure 3 As shown, in box 310, the electronic device 110 obtains user input content via an interactive interface with a virtual object.
[0085] In box 320, the electronic device 110 presents first response content from the virtual object in the interactive interface, where the first response content includes at least one candidate image generated based on the input content.
[0086] In block 330 , the electronic device 110 determines video generation information based on the interaction operation associated with the first response content, where the video generation information includes at least a reference image for generating a video.
[0087] In block 340 , the electronic device 110 provides second response content from the virtual object, the second response content including video content generated based on the video generation information.
[0088] In some embodiments, determining the video generation information based on the interaction operation associated with the first response content includes: in response to receiving a selection of a first image from the at least one candidate image, determining the first image as a reference image.
[0089] In this way, the embodiments of the present disclosure can support the user to flexibly select at least one candidate image included in the first response content to trigger the generation of a video, thereby improving the quality of video generation.
[0090] In some embodiments, determining video generation information based on an interactive operation associated with the first response content includes: receiving a first regeneration request associated with at least one candidate image; presenting at least one updated image generated based on the first regeneration request; and in response to receiving confirmation of a second image in the at least one updated image, determining the second image as a reference image.
[0091] In this way, embodiments of the present disclosure can support user-triggered image regeneration, thereby improving the quality of the image used as a reference for video generation. Further, embodiments of the present disclosure can improve the quality of video generation by user-triggered image regeneration rather than video regeneration.
[0092] In some embodiments, the first response content further includes at least one candidate prompt item for generating the video, and the video generation information further includes a target prompt item determined based on the at least one candidate prompt item.
[0093] In this way, the embodiments of the present disclosure can help users understand more information about video generation before the video is generated by presenting at least one candidate prompt item for generating a video, thereby further improving the quality of video generation.
[0094] In some embodiments, process 300 further includes: in response to receiving a selection of a first prompt item from at least one candidate prompt item, determining the first prompt item as a target prompt item; or based on an editing operation on a second prompt item from at least one candidate prompt item, determining the edited second prompt item as a target prompt item.
[0095] In this way, the embodiments of the present disclosure can help users efficiently adjust the video effects of video generation and further improve the quality of video generation through the user's interactive operation with at least one such candidate prompt item.
[0096] In some embodiments, process 300 further includes: receiving a second regeneration request associated with at least one candidate prompt; presenting at least one updated prompt generated based on the second regeneration request; and in response to receiving confirmation of a third prompt in the at least one updated prompt, determining the third prompt as a target prompt.
[0097] In this way, the embodiments of the present disclosure can improve the quality of prompt items referenced by video generation by supporting users to regenerate prompt items for video generation, thereby improving the quality of video generation.
[0098] In some embodiments, the video generation information includes a plurality of reference images determined based on the first response content, and the plurality of reference images are used to determine a plurality of video frames in the video content.
[0099] In this way, the embodiments of the present disclosure can further constrain the direction of video generation by referring to multiple reference images determined based on the first response content during the video generation process, thereby further improving the quality of video generation.
[0100] In some embodiments, the at least one candidate image is a first set of images generated based on a first portion of the input content, and process 300 further includes: presenting, in the first response content, a second set of images generated based on a second portion of the input content.
[0101] In this way, the embodiments of the present disclosure can support users and virtual objects to provide more images with a lower interaction frequency, thereby improving the interaction efficiency between users and virtual objects.
[0102] In some embodiments, the video generation information is first video generation information, the video content is first video content, and the reference image is a first reference image. Process 300 also includes: determining second video generation information based on the first response content, the second video generation information at least including a second reference image for video generation; and providing second video content generated based on the second video generation information.
[0103] In this way, the embodiments of the present disclosure can further improve the interaction efficiency between the user and the virtual object by supporting the user and the virtual object to provide multiple video contents in the first response content.
[0104] In some embodiments, providing the second video content generated based on the second video generation information includes: providing the second video content independent of the first video content; or providing the first video content, wherein the second video content is a video clip of the first video content.
[0105] In this way, the embodiments of the present disclosure can support flexible provision of multiple independent video contents or merged video contents corresponding to the user's input content, thereby improving the richness of video generation.
[0106] In some embodiments, providing second response content from the virtual object includes: in response to the video content being generated, presenting a reminder message in the interactive interface, the reminder message including a viewing entrance for the video content; and presenting the video content in response to triggering the viewing entrance.
[0107] In this way, the embodiments of the present disclosure can support users to quickly view the generated video content through the viewing entrance in the reminder message during the interaction with the virtual object, thereby improving the interface interaction efficiency between the user and the interactive interface.
[0108] Example devices and equipment
[0109] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4 1 shows a schematic structural block diagram of an example apparatus 400 for generating video content according to some embodiments of the present disclosure. Apparatus 400 may be implemented as or included in electronic device 110. Each module / component in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0110] like Figure 4 As shown, the device 400 includes an acquisition module 410, which is configured to obtain the user's input content via an interactive interface with the virtual object; a first presentation module 420, which is configured to present a first response content from the virtual object in the interactive interface, the first response content including at least one candidate image generated based on the input content; a first determination module 430, which is configured to determine video generation information based on an interactive operation associated with the first response content, the video generation information including at least a reference image for generating a video; and a providing module 440, which is configured to provide a second response content from the virtual object, the second response content including video content generated based on the video generation information.
[0111] In some embodiments, the first determination module 430 is further configured to: in response to receiving a selection of a first image from the at least one candidate image, determine the first image as a reference image.
[0112] In some embodiments, the first determination module 430 is further configured to: receive a first regeneration request associated with at least one candidate image; present at least one updated image generated based on the first regeneration request; and in response to receiving confirmation of the second image in the at least one updated image, determine the second image as a reference image.
[0113] In some embodiments, the first response content further includes at least one candidate prompt item for generating the video, and the video generation information further includes a target prompt item determined based on the at least one candidate prompt item.
[0114] In some embodiments, the device 400 also includes a second determination module, which is configured to: determine the first prompt item as a target prompt item in response to receiving a selection of a first prompt item among at least one candidate prompt item; or determine the edited second prompt item as a target prompt item based on an editing operation on a second prompt item among at least one candidate prompt item.
[0115] In some embodiments, the apparatus 400 further includes a third determination module configured to: receive a second regeneration request associated with at least one candidate prompt item; present at least one updated prompt item generated based on the second regeneration request; and in response to receiving confirmation of a third prompt item in the at least one updated prompt item, determine the third prompt item as a target prompt item.
[0116] In some embodiments, the video generation information includes a plurality of reference images determined based on the first response content, and the plurality of reference images are used to determine a plurality of video frames in the video content.
[0117] In some embodiments, at least one candidate image is a first group of images generated based on a first portion of the input content, and the apparatus 400 further includes a second presentation module configured to present, in the first response content, a second group of images generated based on a second portion of the input content.
[0118] In some embodiments, the video generation information is first video generation information, the video content is first video content, and the reference image is a first reference image. The device 400 also includes a fourth determination module, which is configured to: determine second video generation information based on the first response content, the second video generation information at least including a second reference image for video generation; and provide second video content generated based on the second video generation information.
[0119] In some embodiments, the second presentation module is further configured to: provide second video content independent of the first video content; or provide the first video content, wherein the second video content is a video clip in the first video content.
[0120] In some embodiments, the providing module 440 is further configured to: in response to the video content being generated, present a reminder message in the interactive interface, the reminder message including a viewing entry for the video content; and in response to triggering the viewing entry, present the video content.
[0121] The modules included in the device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the modules in the device 400 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0122] Figure 51 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. Figure 5 The illustrated electronic device 500 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to implement Figure 1 electronic device 110 or Figure 4 device 400.
[0123] like Figure 5 As shown, electronic device 500 is in the form of a general electronic device. Components of electronic device 500 may include, but are not limited to, at least one processor 510 or processing unit, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processor 510 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to increase the parallel processing capabilities of electronic device 500.
[0124] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.
[0125] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 5 As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. Memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0126] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0127] Input device 550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 560 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 500 may also communicate with one or more external devices (not shown) via communication unit 540 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with electronic device 500, or with any device that allows electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0128] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0129] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0130] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0131] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0132] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0133] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating video content, comprising: Obtaining user input through an interactive interface with a virtual object; In the interactive interface, presenting first response content from the virtual object, the first response content including at least one candidate image generated based on the input content; Determining video generation information based on an interactive operation associated with the first response content, the video generation information including at least a reference image for generating a video; as well as Second response content from the virtual object is provided, the second response content including video content generated based on the video generation information.
2. The method according to claim 1, wherein determining video generation information based on the interaction operation associated with the first response content comprises: In response to receiving a selection of a first image from the at least one candidate image, the first image is determined as the reference image.
3. The method according to claim 1 , wherein determining video generation information based on the interaction operation associated with the first response content comprises: receiving a first regeneration request associated with the at least one candidate image; presenting at least one updated image generated based on the first regeneration request; as well as In response to receiving confirmation of a second image in the at least one updated image, the second image is determined as the reference image.
4. The method according to claim 1, wherein the first response content further includes at least one candidate prompt item for generating a video, and the video generation information further includes a target prompt item determined based on the at least one candidate prompt item.
5. The method according to claim 4, further comprising: In response to receiving a selection of a first prompt item among the at least one candidate prompt item, determining the first prompt item as the target prompt item; or Based on an editing operation on a second prompt item among the at least one candidate prompt item, the edited second prompt item is determined as the target prompt item.
6. The method according to claim 5, further comprising: receiving a second regeneration request associated with the at least one candidate prompt item; presenting at least one update prompt item generated based on the second regeneration request; as well as In response to receiving confirmation of a third prompt item in the at least one updated prompt item, the third prompt item is determined as the target prompt item. 7 . The method according to claim 1 , wherein the video generation information includes a plurality of reference images determined based on the first response content, and the plurality of reference images are used to determine a plurality of video frames in the video content.
8. The method according to claim 1, wherein the at least one candidate image is a first group of images generated based on a first portion of the input content, the method further comprising: In the first response content, a second group of images generated based on the second part of the input content is presented.
9. The method according to claim 8, wherein the video generation information is first video generation information, the video content is first video content, and the reference image is a first reference image, the method further comprising: Determining second video generation information based on the first response content, where the second video generation information at least includes a second reference image for video generation; as well as Second video content generated based on the second video generation information is provided.
10. The method according to claim 9, wherein providing the second video content generated based on the second video generation information comprises: providing the second video content independent of the first video content; or The first video content is provided, and the second video content is a video clip in the first video content.
11. The method of claim 1 , wherein providing second response content from the virtual object comprises: In response to the video content being generated, presenting a reminder message in the interactive interface, the reminder message including a viewing entry for the video content; as well as In response to triggering the viewing portal, the video content is presented.
12. An apparatus for generating video content, comprising: An acquisition module is configured to acquire user input content via an interactive interface with a virtual object; A first presentation module is configured to present, in the interactive interface, first response content from the virtual object, where the first response content includes at least one candidate image generated based on the input content; A first determining module is configured to determine video generation information based on an interactive operation associated with the first response content, wherein the video generation information includes at least a reference image for generating a video; as well as A providing module is configured to provide second response content from the virtual object, wherein the second response content includes video content generated based on the video generation information.
13. An electronic device comprising: at least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 11 when executed by the at least one processor.
14. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 11.
15. A computer program product tangibly stored in a computer storage medium and comprising computer executable instructions which, when executed by a device, cause the device to perform the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method and device for generating video in text mode, equipment and medium
CN112153475A
Video generation method and device, readable medium and electronic equipment
CN117896592A
Method and system for generating intelligent picture or video of large model interaction
CN118334145A
Video generation method, computing device, storage medium and program product
CN119815142A
Video generation method and device, equipment and storage medium
CN120111319A
Cited By
Video processing method and device, equipment, storage medium and product
CN121665056A