Video generation method and device, and storage medium
By obtaining the target object image and alternative video description text, users are allowed to select video description text, which solves the problem that the prior art is difficult to generate rich and free videos, and achieves high-level video production and improved user experience.
Patent Information
- Application Number
- PCT/CN2024/138822
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-26
AI Technical Summary
The existing AIGC technology is difficult to generate rich and free videos, and cannot quickly and efficiently meet users' personalized video creation needs.
By obtaining the target object image, at least two alternative video description texts are acquired, and the user is allowed to select, determine the target video description text, and generate the video based on the image and video description text.
It realizes that the user controls the plot development before the video is generated, and generates videos with more free plots, improving the freedom of video production and user experience.
Smart Images

Figure CN2024138822_26062025_PF_FP_ABST
Abstract
Description
Video generation method, device and storage medium
[0001] This application claims priority to Chinese patent application No. 202311786494.6 filed on December 22, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby cited in their entirety as part of this application. Technical Field
[0002] The embodiments of the present disclosure relate to a video generation method, device, and storage medium. Background Art
[0003] Artificial Intelligence Generated Content (AIGC) is a field that applies artificial intelligence technology to content creation. With the continuous development of deep learning and generative models, AIGC has made significant progress in a variety of fields, from image generation to music composition, text writing, and video generation. AIGC technology has shown great potential and can significantly reduce the difficulty of many tasks.
[0004] Video production is one of AIGC's tasks, but currently AIGC can only generate videos with relatively simple plots, and cannot generate videos with richer and freer plots, and cannot quickly and efficiently meet users' personalized video creation needs. Summary of the Invention
[0005] The embodiments of the present disclosure provide a video generation method, device, and storage medium to achieve high-degree-of-freedom video production.
[0006] In a first aspect, an embodiment of the present disclosure provides a video generation method, comprising:
[0007] Acquire a target object image;
[0008] Obtaining at least two candidate video description texts based on the target object image and displaying them, wherein any video description text is used to describe at least one of attribute information, environment, and action of the target object;
[0009] In response to a selection instruction for the candidate video description text, determining a target video description text from the at least two candidate video description texts;
[0010] Video generation is performed based on the target object image and the target video description text.
[0011] In a second aspect, an embodiment of the present disclosure provides a video generation device, including:
[0012] an acquisition unit, configured to acquire an image of a target object;
[0013] a video description text determination unit, configured to obtain at least two candidate video description texts based on the target object image and display them, wherein any video description text is used to describe at least one of the attribute information, environment, and action of the target object;
[0014] a video description text selection unit, configured to determine a target video description text from the at least two candidate video description texts in response to a selection instruction for the candidate video description texts;
[0015] The video generation unit is used to generate a video according to the target object image and the target video description text.
[0016] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor and a memory;
[0017] The memory stores computer-executable instructions;
[0018] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the video generation method described in the first aspect and various possible designs of the first aspect.
[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video generation method described in the first aspect and various possible designs of the first aspect is implemented.
[0020] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, comprising computer-executable instructions. When a processor executes the computer-executable instructions, the video generation method as described in the first aspect and various possible designs of the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0022] FIG1 is a scene example diagram of a video generation method provided by an embodiment of the present disclosure;
[0023] FIG2 is a flow chart of a video generation method according to an embodiment of the present disclosure;
[0024] FIG3 is a schematic diagram of an interface provided by an embodiment of the present disclosure;
[0025] FIG4 is a schematic diagram of an interface provided by another embodiment of the present disclosure;
[0026] FIG5 is a flow chart of a video generation method provided by another embodiment of the present disclosure;
[0027] FIG6 is a structural block diagram of a video generation device provided by an embodiment of the present disclosure;
[0028] FIG7 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0030] Currently, AIGC can only generate videos with relatively simple plots, but cannot generate videos with richer and more free plots. It cannot quickly and efficiently meet users' personalized video creation needs. If you need to generate a video with a complex plot, you need to first write a complex plot script and then generate the video through AIGC. Its efficiency is relatively low and the user experience is poor.
[0031] In order to solve the above technical problems, the present disclosure provides a video generation method, which comprises obtaining a target object image; obtaining at least two alternative video description texts based on the target object image and displaying them, wherein any video description text is used to describe at least one of the attribute information, environment, and action of the target object; determining a target video description text from the at least two alternative video description texts in response to a selection instruction for the alternative video description text; and generating a video based on the target object image and the target video description text. In the present disclosure, based on the target object image, alternative video description texts are first obtained for user selection, and then video generation is automatically performed based on the target video description text selected by the user. This allows the user to control the development of the plot before the video is generated, thereby generating a video with a more flexible plot, improving the freedom of video production, while also increasing interaction with the user and improving the user experience.
[0032] The video generation method provided by the present disclosure can be applied to electronic devices such as terminal devices or servers. For example, if it is applied to a terminal device, it is necessary to deploy a video generation model and a video description text library on the terminal device in advance. After obtaining the target object image (for example, the user takes or uploads a target object image), at least two alternative video description texts can be obtained based on the target object image and the video description text library and displayed. The user can select the displayed alternative video description text to determine the target video description text. Then, the terminal device can call the video generation model based on the target object image and the target video description text to generate the video. Optionally, the last frame or multiple frames of the generated video clip can be further used as the target object image, and the above process can be repeated to generate the next video clip. By repeated repetition, multiple video clips can be obtained, and the multiple video clips can be spliced in sequence to obtain a complete target video.
[0033] If applied on a server, as shown in FIG1 , a video generation model and a video description text library need to be deployed on the server in advance. After the terminal device obtains the target object image (for example, the user takes or uploads a target object image), the terminal device can send the target object image to the server. The server can obtain at least two alternative video description texts based on the target object image and the video description text library and send them to the terminal device for display. The user can select the displayed alternative video description texts on the terminal device to determine the target video description text. Then, the terminal device sends a selection instruction for the alternative video description text to the server. The server can determine the target video description text according to the selection instruction. Further, based on the target object image and the target video description text, the video generation model can be called to generate the video and the generated video (or video clip) can be sent to the terminal device for playback on the terminal device. Optionally, the server can further use the last frame or multiple frames of the generated video clip as the target object image (which does not need to be sent by the terminal device again) and repeat the above process to generate the next video clip. By repeating the above process, multiple video clips can be obtained. By splicing the multiple video clips in sequence, a complete target video can be obtained and finally sent to the terminal device.
[0034] It should be noted that the user information and data involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0035] The video generation method disclosed herein will be described in detail below with reference to specific embodiments.
[0036] Referring to FIG2 , FIG2 is a flow chart of a video generation method provided by an embodiment of the present disclosure. The method of this embodiment can be applied to electronic devices such as terminal devices or servers, and the video generation method includes:
[0037] S201: Acquire a target object image.
[0038] In this embodiment, the target object may be a real person, a real animal, a cartoon character, a cartoon animal, or other possible objects, and the target object image is an image including the target object.
[0039] Optionally, initially for the first video segment to be generated (or when only one video segment needs to be generated), the target object image may be a target object image photographed or uploaded by a user, or a target object image obtained through other means.
[0040] Optionally, multiple video clips can also be generated in this embodiment. In order to ensure the coherence between the video clips, the last frame or multiple frames of the previously generated video clip (including the target object) can be used as the target object image, and the video generation method can be continued to generate the next video clip.
[0041] S202: Obtain at least two candidate video description texts according to the target object image and display them, wherein any video description text is used to describe at least one of the attribute information, environment, and action of the target object.
[0042] In this embodiment, the video description text is used to describe one or more of the attribute information (such as role, dress, etc.), environment, action, etc. of the target object in the video, as input to the video generation model. For example, the video description text can be "on a meadow, blue sky, white clouds, a young man, a ranger, riding a white horse, with a beard and a cloak."
[0043] Optionally, in this embodiment, based on the target object image and a preset video description text library, or at least two alternative video description texts, the video description text library can be pre-configured and can serve as the basis for obtaining the alternative video description text. Optionally, the video description text library can include a complete preset video description text; or, the video description text library can also include different types of description word sets, for example, including but not limited to an object attribute information description word set, an environment description word set, an action description word set, etc., wherein the object attribute information description word set includes multiple preset object attribute information description words, the environment description word set includes multiple preset environment description words, and the action description word set includes multiple preset action description words. Alternative description words can be selected from different types of description word sets respectively, and different types of alternative description words can be combined to generate alternative video description texts. Of course, the alternative video description texts can also be obtained by other means, which are not limited in this embodiment.
[0044] Among them, optionally, in the case where a video generation model is used for video generation in the following S204, if the video generation model is any possible AI model such as a large language model (LLM) or an AIGC (generative artificial intelligence) model, the video description text can also be called a prompt word (Prompt), which is a structured text used as input for the large language model or the AIGC model.
[0045] Furthermore, when obtaining alternative video description text for the target object image, the obtained alternative video description text needs to match the target object image. For example, if the light in the target object image is brighter, the environment in which the target object is located in the alternative video description text can also be a relatively bright environment. For example, if the makeup of the target object in the target object image is more delicate, the target object in the alternative video description text can be some characters with delicate makeup, and can also be equipped with delicate costumes, and so on.
[0046] Therefore, based on the target object image and the video description text library, at least two alternative video description texts can be obtained and displayed for user selection. Optionally, the at least two alternative video description texts can be selected as controls. In addition, considering that the alternative video description texts may be long, identification information of the alternative video description texts can be displayed during display. The identification information is relatively brief and is used to identify the alternative video description text or summarize the alternative video description text. For example, the identification information of the alternative video description text is "fighting with soldiers", while the actual alternative video description text describes many characters, environments, actions, etc. involved in the battle process.
[0047] S203: In response to an instruction to select the candidate video description texts, determine a target video description text from the at least two candidate video description texts.
[0048] In this embodiment, the alternative video description texts are displayed in the interface, and the user can trigger the required alternative video description text, where the triggering method includes but is not limited to clicking, sliding, dragging, etc., to select the target video description text from the alternative video description texts.
[0049] S204: Generate a video based on the target object image and the target video description text.
[0050] In this embodiment, any method that can generate a video based on an image and video description text, such as a deep learning model, can be used.
[0051] Optionally, in this embodiment, a video generation model may be used to implement the video generation. The video generation model is activated based on the target object image and the target video description text to perform video generation.
[0052] The video generation model can be used to automatically generate videos. Its input is an image and a video description text. Some training data (including the image, video description text, and corresponding video) can be obtained in advance to train the video generation model. Optionally, the video generation model is a large language model, AIGC model, or any other deep learning model that can understand the meaning of the video description text and process the image based on the video description text to generate the video.
[0053] In this embodiment, after obtaining the target object image and the target video description text, the video generation model can be called, and the target object image and the target video description text can be input into the video generation model for processing, and the video generation model automatically generates a video (or video clip).
[0054] Furthermore, the generated video (or video clip) can be played to be displayed to the user in real time, so that the user can judge whether it meets his or her needs and make a decision on the next plot development.
[0055] Optionally, in this embodiment, steps S201-S204 may be executed repeatedly to generate multiple video segments, which are then sequentially spliced together (e.g., each generated video segment is spliced after the previously generated video segment) to generate a target video. To ensure continuity between the multiple video segments, the last frame or frames of the previously generated video segment may be used as the target object image when acquiring the target object image.
[0056] The video generation method provided in this embodiment obtains a target object image; based on the target object image, obtains and displays at least two alternative video description texts, wherein any video description text is used to describe at least one of the target object's attribute information, environment, and action; in response to a selection instruction for the alternative video description text, determines a target video description text from the at least two alternative video description texts; and generates a video based on the target object image and the target video description text. Based on the target object image, this embodiment first obtains alternative video description texts for user selection, and then automatically generates a video based on the target video description text selected by the user. This allows the user to control the development of the plot before the video is generated, thereby generating a video with a more flexible plot, improving the freedom of video production, and also increasing interaction with the user, thereby improving the user experience.
[0057] Based on any of the above embodiments, initially, for the first video clip to be generated (that is, the target object appears in the first video clip as the starting part of the plot), the user can shoot or upload an image of the target object (such as a photo, etc.). Of course, the target object image can also be obtained through other means, which is not limited in this embodiment.
[0058] Furthermore, obtaining at least two candidate video description texts according to the target object image includes:
[0059] At least two candidate video description texts for a first video segment to be generated are obtained according to the target object image and a preset video description text library.
[0060] In this embodiment, since the first video clip to be generated initially lacks a preceding plot, the plot requires the appearance of a target object. Based on the target object image captured or uploaded by the user and the video description text library, at least two candidate video description texts for the target object's appearance are generated. These candidate video description texts, i.e., the at least two candidate video description texts for the first video clip to be generated, can include one or more of the target object's attribute information (e.g., role, attire, etc.), the environment in which it appears, and its actions. Furthermore, they can include the transition from the target object image to its appearance. After obtaining the at least two candidate video description texts for the first video clip to be generated, the at least two candidate video description texts (or their identifiers) for the first video clip to be generated can be displayed in an interface. For example, as shown in FIG3 , these can include "Become a Horseback Ranger" and "Become a Cyberpunk Robot," respectively displayed as controls. The user can trigger these controls by clicking, dragging, or other means to select the target video description text for the first video clip to be generated.
[0061] Optionally, when obtaining at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library, at least one of the following methods may be used:
[0062] Method 1:
[0063] 1) obtaining image features of the target object image;
[0064] 2) obtaining a third correlation index between the image feature of the target object image and a preset video description text corresponding to the first video clip in the video description text library, wherein the video description text library includes a plurality of preset video description texts corresponding to the first video clip;
[0065] 3) Selecting at least two preset video description texts corresponding to the first video segment according to the third relevance index, and determining them as at least two candidate video description texts for the first video segment to be generated.
[0066] In this embodiment, the target video description text library may be pre-configured with multiple complete preset video description texts corresponding to the first video clip. Each preset video description text may include one or more items of object attribute information, environment, action, etc. of the determined appearance scene. When it is necessary to obtain candidate video description texts for the first video clip to be generated for the target object image, image features of the target object image may be extracted, wherein the method for extracting image features may adopt any known method. Then, a third correlation index between the image features of the target object image and each preset video description text corresponding to the first video clip in the video description text library is obtained. Based on the magnitude of the third correlation index, at least two preset video description texts with the largest third correlation index are selected from the preset video description texts corresponding to the first video clip as candidate video description texts. The third correlation index between the image features of the target object image and each preset video description text corresponding to the first video clip in the video description text library may be obtained using any known method for obtaining the correlation between image features and text, which is not limited here.
[0067] Method 2:
[0068] 1) obtaining image features of the target object image;
[0069] 2) obtaining image features of the target object image and matching them with image features of historical object images in the video description text library to determine a matching historical object image, wherein the video description text library includes target video description text of the historical object image;
[0070] 3) The target video description text of the first video segment corresponding to the matched historical object image is determined as the candidate video description text of the first video segment to be generated.
[0071] In this embodiment, the target video description text library stores the processed target video description text (or alternative video description text) of the historical object image, including the target video description text (or alternative video description text) of the first video segment corresponding to the historical object image. When it is necessary to obtain the alternative video description text of the first video segment to be generated for the target object image, image features can be extracted from the target object image, wherein the method for extracting the image features can adopt any known method, and then the image features of the target object image are obtained and matched with the image features of the historical object image to determine the historical object image that matches the target object image, that is, the target object image and the historical object image have a high degree of similarity. The target video description text (or alternative video description text) of the first video segment corresponding to the historical object image can be referenced, that is, the target video description text (or alternative video description text) of the first video segment corresponding to the matched historical object image can be determined as the alternative video description text of the first video segment to be generated for the target object image. The image features of the target object image and the image features of the historical object image can be matched using any known image matching method, which is not limited here.
[0072] Method 3:
[0073] 1) obtaining image features of the target object image;
[0074] 2) obtaining a fourth correlation index between an image feature of the target object image and each preset descriptive word in a set of different types of descriptive words corresponding to a first video clip in the video description text library, wherein the set of different types of descriptive words includes at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set;
[0075] 3) Selecting different types of preset description words corresponding to the first video clip according to the fourth relevance index, determining them as different types of candidate description words, combining the different types of candidate description words, and generating at least two candidate video description texts for the first video clip to be generated.
[0076] In this embodiment, the target video description text library may include different types of descriptor sets corresponding to the first video clip, wherein the different types of descriptor sets include, but are not limited to, at least one of an object attribute information descriptor set, an environment descriptor set, and an action descriptor set. The object attribute information descriptor set includes multiple preset object attribute information descriptors, the environment descriptor set includes multiple preset environment descriptors, and the action descriptor set includes multiple preset action descriptors. When it is necessary to obtain candidate video description text for the first video clip to be generated for the target object image, image features of the target object image may be extracted, wherein the method for extracting the image features may adopt any known method. Then, a fourth correlation index between the image features of the target object image and each preset descriptor in the different types of descriptor sets corresponding to the first video clip in the video description text library is obtained. Based on the magnitude of the fourth correlation index, at least two preset descriptors with the largest fourth correlation index are selected from the preset descriptor sets of each type corresponding to the first video clip as candidate descriptors of each type. The candidate descriptors of the different types are then combined to generate at least two candidate video description texts for the first video clip to be generated. Among them, the fourth correlation index between the image features of the target object image and the preset descriptive words in the set of different types of descriptive words corresponding to the first video clip in the video description text library can be obtained by any known method of obtaining the correlation between image features and text (words), and there is no limitation here.
[0077] Based on the above embodiment, after the video segments have been generated, subsequent video segments can be generated to continue the plot. To ensure the coherence between the video segments, the last frame or frames of the previously generated video segment (including the target object) can be used as the target object image to continue the video generation method.
[0078] Specifically, after obtaining the last frame or frames of the previously generated video clip, the last frame or frames can be used as the target object image. Based on the target object image and a preset video description text library, at least two candidate video description texts can be obtained, which may specifically include:
[0079] At least two candidate video description texts for the turning event are obtained according to the target object image, the target video description text of the previously generated video clip, and a preset video description text library.
[0080] In this embodiment, in order to ensure the continuation of the plot, it is necessary to obtain at least two alternative video description texts of the turning event. The turning event is a turning point relative to the previously generated video clip, that is, at least one of the object attribute information, environment, and action has changed relative to the previously generated video clip. By providing the user with at least two alternative video description texts of the turning event, the user can select the turning event to control the direction of the plot development.
[0081] When obtaining alternative video description texts for a turning point event, in addition to relying on the target object image, since the target video description text of the previously generated video clip also contains a large amount of information that affects the subsequent plot development, it is necessary to combine the target video description text of the previously generated video clip. Therefore, in this embodiment, at least two alternative video description texts for a turning point event can be obtained based on the target object image, the target video description text of the previously generated video clip, and the preset video description text library.
[0082] Optionally, during the playback of the previously generated video clip, at least two alternative video description texts of the turning event can be obtained based on the target object image, the target video description text of the previously generated video clip and the preset video description text library, and after the playback of the previously generated video clip is completed, at least two alternative video description texts of the turning event can be displayed for the user to choose to determine the next plot direction. For example, as shown in Figure 4, after the playback of the previously generated video clip is completed (it can be displayed that the last frame has been played), at least two alternative video description texts (or their logos) of the turning event can be displayed in the interface, such as "fighting with soldiers" and "shaking hands with the old man", etc., which are displayed in the form of controls respectively. The user can trigger them by clicking, dragging, etc. to select the target video description text of the turning event.
[0083] Optionally, when obtaining at least two candidate video description texts for a turning point event based on the target object image, the target video description text of the previously generated video clip, and a preset video description text library, at least one of the following methods may be used:
[0084] Method 1:
[0085] 1) obtaining image features of the target object image;
[0086] 2) obtaining an image feature of the target object image and a first correlation index between a target video description text of a previously generated video clip and a preset video description text of a turning point event in the video description text library, wherein the video description text library includes a plurality of preset video description texts of a turning point event;
[0087] 3) Selecting at least two preset video description texts of the turning event according to the first correlation index to determine them as at least two candidate video description texts of the turning event.
[0088] In this embodiment, the target video description text library can be pre-configured with multiple complete preset video description texts for turning events. Each preset video description text can include one or more items of the object attribute information, environment, action, etc. of the already determined turning event. When it is necessary to obtain alternative video description texts for the turning event for the target object image, image features can be extracted from the target object image, wherein the method for extracting image features can adopt any known method. Then, the image features of the target object image and the first correlation index between the target video description text of the previously generated video clip and the preset video description texts of the turning event in the video description text library are obtained. Then, based on the magnitude of the first correlation index, at least two preset video description texts with the largest first correlation index are selected from the preset video description texts of the turning event as the alternative video description texts. The method for obtaining the image features of the target object image and the first correlation index between the target video description text of the previously generated video clip and the preset video description texts of the turning event in the video description text library can adopt any known method for obtaining the correlation between image features and text, which is not limited here.
[0089] Method 2:
[0090] 1) obtaining image features of the target object image;
[0091] 2) obtaining a second correlation index between the image features of the target object image and the target video description text of a previously generated video clip and each preset descriptive word in a set of different types of descriptive words for transition events in the video description text library, wherein the set of different types of descriptive words includes at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set;
[0092] 3) Selecting different types of preset description words for the turning event according to the second relevance index, determining them as different types of candidate description words, combining the different types of candidate description words, and generating at least two candidate video description texts for the turning event.
[0093] In this embodiment, the target video description text library may include a set of different types of descriptive words for the turning event (which may be the same as or different from the first set of different types of descriptive words for the first video clip), wherein the set of different types of descriptive words includes, but is not limited to, at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set, wherein the object attribute information descriptive word set includes multiple preset object attribute information descriptive words, the environment descriptive word set includes multiple preset environment descriptive words, and the action descriptive word set includes multiple preset action descriptive words. When it is necessary to obtain an alternative video description text for the turning event for the target object image, image features of the target object image may be extracted, wherein the method for extracting the image features may adopt any known method, and then the image features of the target object image and the second correlation index between each preset descriptive word in the set of different types of descriptive words for the turning event in the video description text library are obtained. Then, based on the magnitude of the second correlation index, at least two preset descriptive words with the largest second correlation index are selected from the set of preset descriptive words for each type of the turning event as the candidate descriptive words for each type. The candidate descriptive words of different types are combined to generate at least two alternative video description texts for the turning event. Among them, the second correlation index between the image features of the target object image and the target video description text of the previously generated video clip and each preset descriptive word in the different types of descriptive word sets of the turning events in the video description text library is obtained. Any known method for obtaining the correlation between image features and text (words) can be used, and there is no limitation here.
[0094] On the basis of the above embodiment, as shown in FIG5 , initially, the user can shoot or upload the target object image, and determine the alternative video description texts of the first video segment to be generated, such as appearance prompt1 and appearance prompt2, based on the target object image and in combination with the video description text library (prompt library), and display them. The user can select the target video description text of the first video segment to be generated: target appearance prompt, and then based on the target object image shot or uploaded by the user and the target video description text of the first video segment to be generated, call the video generation model to generate the first video segment (that is, the video segment of the appearance scene), and play it; the last frame (or multiple frames) of the first video segment can also be obtained and determined as the target object image, and based on the target object image and the target video description text of the first video segment (target appearance prompt), and in combination with the video description text library (prompt library), the target object image and the target video description text of the first video segment (target appearance prompt) are generated. The library) determines the alternative video description texts of the turning event 1, such as the turning prompt 1 and the turning prompt 2, and displays them. The user can select the target video description text of the turning event 1: the target turning prompt, and then based on the target object image and the target video description text (target turning prompt) of the turning event, call the video generation model to generate a video clip of the turning event 1 (which can be spliced after the first video clip) and play it; the last frame (or multiple frames) of the video clip of the turning event 1 can be further obtained to determine it as the target object image, and on the basis of the target object image and the target video description text of the turning event 1, the alternative video description texts of the turning event 2, the target video description text, and the generation of the video clip can be continued. The above process is repeated to obtain multiple video clips, and the multiple video clips are spliced in sequence to obtain a target video with a coherent plot and rich diversity.
[0095] On the basis of any of the above embodiments, the camera movement mode may be controlled so that the generated video (or video clip) moves in the determined target camera movement mode.
[0096] Optionally, in response to an instruction to select an alternative video description text, the target camera movement mode corresponding to the target video description text can also be determined. That is, in this embodiment, different alternative video description texts can correspond to different camera movement modes. When the user selects an alternative video description text as the target video description text, the camera movement mode corresponding to the alternative video description text can also be determined as the target camera movement mode corresponding to the target video description text.
[0097] Furthermore, in any of the above embodiments, when generating a video based on the target object image and the target video description text, the video can be generated based on the target object image, the target video description text and the target camera movement method. Optionally, the target camera movement method can also be input into the video generation model as an input data, or the target camera movement method can also be added to the target video description text and input into the video generation model together.
[0098] Optionally, when at least two alternative video description texts are displayed in the interface, a specific arrangement can be adopted, such as left-right arrangement. When selecting, the user can trigger the alternative video description text displayed in the interface, such as clicking, sliding, dragging, etc., and the trigger instructions for the alternative video description texts in different positions can determine different camera movement methods, that is, according to the position of the target video description text in the interface, the target camera movement method corresponding to the target video description text is determined. For example, if the user selects the alternative video description text on the left as the target video description text, the target camera movement method can be determined as the lens moving to the left first and then returning to the middle, that is, in the video, the picture first turns left and then returns to the front; if the user selects the alternative video description text on the right as the target video description text, the target camera movement method can be determined as the lens moving to the right first and then returning to the middle, that is, in the video, the picture first turns right and then returns to the front, so that it can be distinguished in the video that the user has selected different alternative video description texts, and the user can feel that the video responds to the user's operation, increasing interactivity and improving the interactive experience.
[0099] Optionally, in the first video, the camera movement may be fixed, for example, the picture only moves backward to show the whole body of the target object, so that the target object appears to be moving towards the audience in front of the screen, that is, presenting an appearance effect.
[0100] Corresponding to the video generation method described in the preceding embodiment, FIG6 is a block diagram of a video generation device provided in an embodiment of the present disclosure. For ease of illustration, only the portions relevant to the embodiment of the present disclosure are shown. Referring to FIG6 , the video generation device 600 includes an acquisition unit 601, a video description text determination unit 602, a video description text selection unit 603, and a video generation unit 604.
[0101] The acquisition unit 601 is used to acquire an image of a target object;
[0102] The video description text determining unit 602 is configured to obtain at least two candidate video description texts based on the target object image and display them, wherein any video description text is used to describe at least one of the attribute information, environment, and action of the target object;
[0103] The video description text selection unit 603 is configured to determine a target video description text from the at least two candidate video description texts in response to a selection instruction for the candidate video description texts;
[0104] The video generating unit 604 is configured to generate a video according to the target object image and the target video description text.
[0105] In one or more embodiments of the present disclosure, when acquiring the target object image, the acquisition unit is configured to:
[0106] Obtaining the last frame or frames of a previously generated video clip and determining them as the target object image;
[0107] Accordingly, after generating a video according to the target object image and the target video description text, the video generation unit is further configured to:
[0108] The generated video segment is spliced with the previously generated video segment.
[0109] In one or more embodiments of the present disclosure, when acquiring at least two candidate video description texts based on the target object image, the video description text determination unit is configured to:
[0110] At least two candidate video description texts for the turning event are obtained according to the target object image, the target video description text of the previously generated video clip, and a preset video description text library.
[0111] In one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for a turning point event based on the target object image, the target video description text of a previously generated video clip, and a preset video description text library, it is configured to:
[0112] Acquiring image features of the target object image;
[0113] Obtaining an image feature of the target object image and a first correlation index between a target video description text of a previously generated video clip and a preset video description text of a turning point event in the video description text library, wherein the video description text library includes a plurality of preset video description texts of a turning point event;
[0114] At least two preset video description texts of the turning event are selected according to the first correlation index and determined as at least two candidate video description texts of the turning event.
[0115] In one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for a turning point event based on the target object image, the target video description text of a previously generated video clip, and a preset video description text library, it is configured to:
[0116] Acquiring image features of the target object image;
[0117] Obtaining a second correlation index between an image feature of the target object image and a target video description text of a previously generated video clip and each preset descriptive word in a set of different types of descriptive words for transition events in the video description text library, wherein the set of different types of descriptive words includes at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set;
[0118] Different types of preset description words for the turning event are selected according to the second correlation index, determined as different types of candidate description words, and the different types of candidate description words are combined to generate at least two candidate video description texts for the turning event.
[0119] In one or more embodiments of the present disclosure, when acquiring the target object image, the acquisition unit is configured to:
[0120] For the first video clip to be generated, obtain the target object image taken or uploaded by the user;
[0121] Accordingly, obtaining at least two candidate video description texts according to the target object image includes:
[0122] At least two candidate video description texts for the first video segment to be generated are obtained according to the target object image and a preset video description text library.
[0123] In one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library, it is configured to:
[0124] Acquiring image features of the target object image;
[0125] Obtaining a third correlation index between an image feature of the target object image and a preset video description text corresponding to a first video clip in the video description text library, wherein the video description text library includes a plurality of preset video description texts corresponding to the first video clip;
[0126] At least two preset video description texts corresponding to the first video segment are selected according to the third correlation index, and are determined as the at least two candidate video description texts for the first video segment to be generated.
[0127] In one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library, it is configured to:
[0128] Acquiring image features of the target object image;
[0129] Acquire image features of the target object image and match them with image features of historical object images in the video description text library to determine a matching historical object image, wherein the video description text library includes target video description text of the historical object image;
[0130] The target video description text of the first video segment corresponding to the matched historical object image is determined as the candidate video description text of the first video segment to be generated.
[0131] In one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library, it is configured to:
[0132] Acquiring image features of the target object image;
[0133] Obtaining a fourth correlation index between an image feature of the target object image and each preset descriptive word in a set of different types of descriptive words corresponding to a first video clip in the video description text library, wherein the set of different types of descriptive words includes at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set;
[0134] Different types of preset description words corresponding to the first video clip are selected according to the fourth correlation index, determined as different types of candidate description words, and the different types of candidate description words are combined to generate at least two candidate video description texts for the first video clip to be generated.
[0135] In one or more embodiments of the present disclosure, when the video generation unit generates a video based on the target object image and the target video description text, it is configured to:
[0136] According to the target object image and the target video description text, a video generation model is called to generate the video.
[0137] In one or more embodiments of the present disclosure, the video description text selection unit 603, in response to an instruction to select an alternative video description text, is further configured to:
[0138] Determining a target camera movement method corresponding to the target video description text;
[0139] Accordingly, when generating a video based on the target object image and the target video description text, the video generating unit 604 is configured to:
[0140] Video generation is performed according to the target object image, the target video description text, and the target camera movement method.
[0141] In one or more embodiments of the present disclosure, the instruction to select the candidate video description text is a trigger instruction for the candidate video description text displayed in the interface; accordingly, when determining the target camera movement mode corresponding to the target video description text, the video description text selection unit 603 is configured to:
[0142] According to the position of the target video description text in the interface, a target camera movement mode corresponding to the target video description text is determined.
[0143] The device provided in this embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0144] Referring to FIG7 , a schematic diagram of the structure of an electronic device 700 suitable for implementing an embodiment of the present disclosure is shown. The electronic device 700 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG7 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0145] As shown in FIG7 , the electronic device 700 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the electronic device 700 are also stored in the RAM 703. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0146] Typically, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 7 shows an electronic device 700 having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0147] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0148] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0149] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0150] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0151] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0153] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0154] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0155] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0156] In a first aspect, according to one or more embodiments of the present disclosure, a video generation method is provided, comprising:
[0157] Acquire a target object image;
[0158] Obtaining at least two candidate video description texts based on the target object image and displaying them, wherein any video description text is used to describe at least one of attribute information, environment, and action of the target object;
[0159] In response to a selection instruction for the candidate video description text, determining a target video description text from the at least two candidate video description texts;
[0160] Video generation is performed based on the target object image and the target video description text.
[0161] According to one or more embodiments of the present disclosure, acquiring the target object image includes:
[0162] Obtaining the last frame or frames of a previously generated video clip and determining them as the target object image;
[0163] Correspondingly, after generating a video based on the target object image and the target video description text, the method further includes:
[0164] The generated video segment is spliced with the previously generated video segment.
[0165] According to one or more embodiments of the present disclosure, obtaining at least two candidate video description texts based on the target object image includes:
[0166] At least two candidate video description texts for the turning event are obtained according to the target object image, the target video description text of the previously generated video clip, and a preset video description text library.
[0167] According to one or more embodiments of the present disclosure, obtaining at least two candidate video description texts for a turning point event based on the target object image, the target video description text of a previously generated video clip, and a preset video description text library includes:
[0168] Acquiring image features of the target object image;
[0169] Obtaining an image feature of the target object image and a first correlation index between a target video description text of a previously generated video clip and a preset video description text of a turning point event in the video description text library, wherein the video description text library includes a plurality of preset video description texts of a turning point event;
[0170] At least two preset video description texts of the turning event are selected according to the first correlation index and determined as at least two candidate video description texts of the turning event.
[0171] According to one or more embodiments of the present disclosure, obtaining at least two candidate video description texts for a turning point event based on the target object image, the target video description text of a previously generated video clip, and a preset video description text library includes:
[0172] Acquiring image features of the target object image;
[0173] Obtaining a second correlation index between an image feature of the target object image and a target video description text of a previously generated video clip and each preset descriptive word in a set of different types of descriptive words for transition events in the video description text library, wherein the set of different types of descriptive words includes at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set;
[0174] Different types of preset description words for the turning event are selected according to the second correlation index, determined as different types of candidate description words, and the different types of candidate description words are combined to generate at least two candidate video description texts for the turning event.
[0175] According to one or more embodiments of the present disclosure, acquiring the target object image includes:
[0176] For the first video clip to be generated, obtain the target object image taken or uploaded by the user;
[0177] Accordingly, obtaining at least two candidate video description texts according to the target object image includes:
[0178] At least two candidate video description texts for the first video segment to be generated are obtained according to the target object image and a preset video description text library.
[0179] According to one or more embodiments of the present disclosure, obtaining at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library includes:
[0180] Acquiring image features of the target object image;
[0181] Obtaining a third correlation index between an image feature of the target object image and a preset video description text corresponding to a first video clip in the video description text library, wherein the video description text library includes a plurality of preset video description texts corresponding to the first video clip;
[0182] At least two preset video description texts corresponding to the first video segment are selected according to the third correlation index, and are determined as the at least two candidate video description texts for the first video segment to be generated.
[0183] According to one or more embodiments of the present disclosure, obtaining at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library includes:
[0184] Acquiring image features of the target object image;
[0185] Acquire image features of the target object image and match them with image features of historical object images in the video description text library to determine a matching historical object image, wherein the video description text library includes target video description text of the historical object image;
[0186] The target video description text of the first video segment corresponding to the matched historical object image is determined as the candidate video description text of the first video segment to be generated.
[0187] According to one or more embodiments of the present disclosure, obtaining at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library includes:
[0188] Acquiring image features of the target object image;
[0189] Obtaining a fourth correlation index between an image feature of the target object image and each preset descriptive word in a set of different types of descriptive words corresponding to a first video clip in the video description text library, wherein the set of different types of descriptive words includes at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set;
[0190] Different types of preset description words corresponding to the first video clip are selected according to the fourth correlation index, determined as different types of candidate description words, and the different types of candidate description words are combined to generate at least two candidate video description texts for the first video clip to be generated.
[0191] According to one or more embodiments of the present disclosure, generating a video based on the target object image and the target video description text includes:
[0192] According to the target object image and the target video description text, a video generation model is called to generate the video.
[0193] According to one or more embodiments of the present disclosure, in response to an instruction to select an alternative video description text, the method further includes:
[0194] Determining a target camera movement method corresponding to the target video description text;
[0195] Accordingly, the video generation according to the target object image and the target video description text includes:
[0196] Video generation is performed according to the target object image, the target video description text, and the target camera movement method.
[0197] According to one or more embodiments of the present disclosure, the instruction to select the alternative video description text is a trigger instruction for the alternative video description text displayed in the interface; accordingly, determining the target camera movement mode corresponding to the target video description text includes:
[0198] According to the position of the target video description text in the interface, a target camera movement mode corresponding to the target video description text is determined.
[0199] In a second aspect, according to one or more embodiments of the present disclosure, a video generating device is provided, including:
[0200] an acquisition unit, configured to acquire an image of a target object;
[0201] a video description text determination unit, configured to obtain at least two candidate video description texts based on the target object image and display them, wherein any video description text is used to describe at least one of the attribute information, environment, and action of the target object;
[0202] a video description text selection unit, configured to determine a target video description text from the at least two candidate video description texts in response to a selection instruction for the candidate video description texts;
[0203] The video generation unit is used to generate a video according to the target object image and the target video description text.
[0204] According to one or more embodiments of the present disclosure, when acquiring the target object image, the acquisition unit is configured to:
[0205] Obtaining the last frame or frames of a previously generated video clip and determining them as the target object image;
[0206] Accordingly, after generating a video according to the target object image and the target video description text, the video generation unit is further configured to:
[0207] The generated video segment is spliced with the previously generated video segment.
[0208] According to one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts based on the target object image, it is configured to:
[0209] At least two candidate video description texts for the turning event are obtained according to the target object image, the target video description text of the previously generated video clip, and a preset video description text library.
[0210] According to one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for a turning point event based on the target object image, the target video description text of a previously generated video clip, and a preset video description text library, it is configured to:
[0211] Acquiring image features of the target object image;
[0212] Obtaining an image feature of the target object image and a first correlation index between a target video description text of a previously generated video clip and a preset video description text of a turning point event in the video description text library, wherein the video description text library includes a plurality of preset video description texts of a turning point event;
[0213] At least two preset video description texts of the turning event are selected according to the first correlation index and determined as at least two candidate video description texts of the turning event.
[0214] According to one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for a turning point event based on the target object image, the target video description text of a previously generated video clip, and a preset video description text library, it is configured to:
[0215] Acquiring image features of the target object image;
[0216] Obtaining a second correlation index between an image feature of the target object image and a target video description text of a previously generated video clip and each preset descriptive word in a set of different types of descriptive words for transition events in the video description text library, wherein the set of different types of descriptive words includes at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set;
[0217] Different types of preset description words for the turning event are selected according to the second correlation index, determined as different types of candidate description words, and the different types of candidate description words are combined to generate at least two candidate video description texts for the turning event.
[0218] According to one or more embodiments of the present disclosure, when acquiring the target object image, the acquisition unit is configured to:
[0219] For the first video clip to be generated, obtain the target object image taken or uploaded by the user;
[0220] Accordingly, obtaining at least two candidate video description texts according to the target object image includes:
[0221] At least two candidate video description texts for the first video segment to be generated are obtained according to the target object image and a preset video description text library.
[0222] According to one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library, it is configured to:
[0223] Acquiring image features of the target object image;
[0224] Obtaining a third correlation index between an image feature of the target object image and a preset video description text corresponding to a first video clip in the video description text library, wherein the video description text library includes a plurality of preset video description texts corresponding to the first video clip;
[0225] At least two preset video description texts corresponding to the first video segment are selected according to the third correlation index, and are determined as the at least two candidate video description texts for the first video segment to be generated.
[0226] According to one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library, it is configured to:
[0227] Acquiring image features of the target object image;
[0228] Acquire image features of the target object image and match them with image features of historical object images in the video description text library to determine a matching historical object image, wherein the video description text library includes target video description text of the historical object image;
[0229] The target video description text of the first video segment corresponding to the matched historical object image is determined as the candidate video description text of the first video segment to be generated.
[0230] According to one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two candidate video description texts for the first video segment to be generated based on the target object image and a preset video description text library, it is configured to:
[0231] Acquiring image features of the target object image;
[0232] Obtaining a fourth correlation index between an image feature of the target object image and each preset descriptive word in a set of different types of descriptive words corresponding to a first video clip in the video description text library, wherein the set of different types of descriptive words includes at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set;
[0233] Different types of preset description words corresponding to the first video clip are selected according to the fourth correlation index, determined as different types of candidate description words, and the different types of candidate description words are combined to generate at least two candidate video description texts for the first video clip to be generated.
[0234] According to one or more embodiments of the present disclosure, when the video generation unit generates a video based on the target object image and the target video description text, it is configured to:
[0235] According to the target object image and the target video description text, a video generation model is called to generate the video.
[0236] According to one or more embodiments of the present disclosure, the video description text selection unit, in response to an instruction to select an alternative video description text, is further configured to:
[0237] Determining a target camera movement method corresponding to the target video description text;
[0238] Accordingly, when the video generation unit generates a video according to the target object image and the target video description text, it is configured to:
[0239] Video generation is performed according to the target object image, the target video description text, and the target camera movement method.
[0240] According to one or more embodiments of the present disclosure, the instruction to select the alternative video description text is a trigger instruction for the alternative video description text displayed in the interface; accordingly, when determining the target camera movement mode corresponding to the target video description text, the video description text selection unit is configured to:
[0241] According to the position of the target video description text in the interface, a target camera movement mode corresponding to the target video description text is determined.
[0242] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;
[0243] The memory stores computer-executable instructions;
[0244] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the video generation method described in the first aspect and various possible designs of the first aspect.
[0245] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video generation method described in the first aspect and various possible designs of the first aspect is implemented.
[0246] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising computer-executable instructions. When a processor executes the computer-executable instructions, the video generation method as described in the first aspect and various possible designs of the first aspect is implemented.
[0247] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0248] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0249] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A video generation method, comprising: Acquire a target object image; According to the target object image, at least two candidate video description texts are obtained and displayed, wherein any video description text is used to describe at least one of attribute information, environment, and action of the target object; In response to a selection instruction for the alternative video description text, the target video description text is determined from the at least two alternative video description texts; Video generation is performed based on the target object image and the target video description text.
2. The method according to claim 1, wherein: The step of acquiring the target object image comprises: Obtain the last frame or multiple frames of the previous generated video clip and determine it as the target object image; Accordingly, after video generation is performed based on the target object image and the target video description text, it further includes: The generated video clip is spliced with the previous generated video clip.
3. The method according to claim 2, wherein: The acquisition of at least two alternative video description texts according to the target object image, including: At least two alternative video description texts for the turning event are obtained according to the target object image, the target video description text of the previous generated video clip, and the preset video description text library.
4. The method according to claim 3, wherein: The step of obtaining at least two candidate video description texts for the transition event according to the target object image, the target video description text of the previously generated video clip, and a preset video description text library includes: Acquire image features of the target object image; Acquire an image feature of the target object image and a first correlation index between a target video description text of a previously generated video clip and a preset video description text of a transition event in the video description text library, wherein the video description text library includes a plurality of preset video description texts of a transition event; At least two preset video description texts of the turning event are selected according to the first relevance indicator, and at least two alternative video description texts of the turning event are determined as the at least two alternative video description texts of the turning event.
5. The method according to claim 3, wherein: The step of obtaining at least two candidate video description texts for the transition event according to the target object image, the target video description text of the previously generated video clip, and a preset video description text library includes: Acquire image features of the target object image; Obtaining a second correlation index between the image feature of the target object image and the target video description text of a previously generated video clip and each preset descriptive word in a different type of descriptive word set of a transition event in the video description text library, wherein the different type of descriptive word set includes at least one item of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set; Different types of preset description words for the turning event are selected according to the second correlation index, determined as different types of candidate description words, and the different types of candidate description words are combined to generate at least two candidate video description texts for the turning event.
6. The method according to claim 1, wherein: The step of acquiring the target object image comprises: For the first video clip to be generated, obtain the target object image captured or uploaded by the user; Accordingly, the acquisition of at least two alternative video description texts according to the target object image, including: According to the target object image and the preset video description text library, at least two alternative video description texts of the first video clip to be generated are obtained.
7. The method according to claim 6, wherein: The at least two alternative video description texts of the first video clip to be generated are obtained according to the target object image and the preset video description text library, including: Acquire image features of the target object image; Acquire a third correlation index between the image feature of the target object image and a preset video description text corresponding to a first video clip in the video description text library, wherein the video description text library includes a plurality of preset video description texts corresponding to the first video clip; At least two preset video description texts corresponding to the first video clip are selected according to the third correlation index, and at least two alternative video description texts of the first video clip to be generated are determined.
8. The method according to claim 6, wherein: The at least two alternative video description texts of the first video clip to be generated are obtained according to the target object image and the preset video description text library, including: Acquire image features of the target object image; Acquire image features of the target object image and match them with image features of historical object images in the video description text library to determine a matching historical object image, wherein the video description text library includes a target video description text of the historical object image; The target video description text of the first video clip corresponding to the matching historical object image is determined as an alternative video description text of the first video clip to be generated.
9. The method according to claim 6, wherein: The at least two alternative video description texts of the first video clip to be generated are obtained according to the target object image and the preset video description text library, including: Acquire image features of the target object image; Obtaining a fourth correlation index between an image feature of the target object image and each preset descriptive word in a different type descriptive word set corresponding to a first video clip in the video description text library, wherein the different type descriptive word set includes at least one of an object attribute information descriptive word set, an environment descriptive word set, and an action descriptive word set; Different types of preset description words corresponding to the first video clip are selected according to the fourth correlation index, determined as different types of candidate description words, and the different types of candidate description words are combined to generate at least two candidate video description texts for the first video clip to be generated.
10. The method according to any one of claims 1 to 9, wherein: The video generation is performed based on the target object image and the target video description text, including: According to the target object image and the target video description text, a video generation model is called to perform video generation.
11. The method according to any one of claims 1 to 10, wherein: In response to a selection instruction for the alternative video description text, it is also included: Determine the target mirror mode corresponding to the target video description text; Accordingly, the video generation is performed based on the target object image and the target video description text, including: Video generation is performed based on the target object image, the target video description text, and the target mirroring method.
12. The method according to claim 10, wherein: The selection instruction for the candidate video description text is a trigger instruction for the candidate video description text displayed in the interface; accordingly, determining the target camera movement mode corresponding to the target video description text includes: According to the position of the target video description text in the interface, the target mirror mode corresponding to the target video description text is determined.
13. A video generating device, comprising: The acquisition unit is configured to acquire the target object image; A video description text determination unit is configured to obtain at least two candidate video description texts according to the target object image and display them, wherein any video description text is used to describe at least one of the attribute information, the environment, and the action of the target object; The video description text selection unit is configured to determine the target video description text from the at least two alternative video description texts in response to a selection instruction for the alternative video description text; The video generation unit is configured to perform video generation based on the target object image and the target video description text.
14. An electronic device comprising: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes computer-executing instructions stored in the memory, causing the at least one processor to perform the method of any of claims 1-12.
15. A computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the method according to any one of claims 1 to 12 is implemented.
16. A computer program product, comprising a computer-executing instructions, implementing the method of any one of claims 1-12 when a processor executes the computer-executing instructions.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and storage medium
CN113821690A
Video generation method and device, equipment, medium and product
CN114501064A
Video generation method and related device
CN116916112A
Video generation method and device, equipment and storage medium
CN117177025A
Video generating method and electronic device
WO2023005194A1
Cited By
Video generation method and device, electronic equipment and storage medium
CN121509771A