Video generation method and device and storage medium

By obtaining the target object image and alternative video description text, allowing users to select and generate videos, the problem that existing AIGC technology is difficult to generate rich and free videos is solved, and high-degree of freedom of video production and improved user experience is achieved.

CN120201249APending Publication Date: 2025-06-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311786494.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-22
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing AIGC technology is difficult to generate rich and free videos, and cannot quickly and efficiently meet users' personalized video creation needs.

Method used

By obtaining the target object image, at least two alternative video description texts are acquired, and the user is allowed to select, determine the target video description text, and then generate the video based on the image and video description text.

Benefits of technology

It realizes the development of the plot controlled by the user before the video is generated, generates videos with more free plots, improves the freedom of video production, increases user interaction and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201249A_ABST
    Figure CN120201249A_ABST
Patent Text Reader

Abstract

Embodiments of the invention provide a video generation method and device, and a storage medium. The method comprises the steps of obtaining a target object image; obtaining at least two alternative video description texts according to the target object image, and displaying the alternative video description texts; in response to a selection instruction for alternative video description texts, determining a target video description text from the at least two alternative video description texts; and generating a video according to the target object image and the target video description text. According to the embodiment of the invention, based on the target object image, the alternative video description text is firstly acquired for the user to select, and then the video is automatically generated based on the target video description text selected by the user, so that the user can control the development of a plot before the video is generated, thereby generating a video with a freer plot, and improving the user experience. The degree of freedom of video production is improved, the interaction with the user is increased, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer and network communication technologies, and in particular, to a video generation method, device, and storage medium. Background Art

[0002] AIGC (Artificial Intelligence Generated Content) is a field that applies artificial intelligence technology to content creation. With the continuous development of deep learning and generative models, AIGC has made remarkable progress in multiple fields. From image generation to music creation, to text writing and video generation, it has shown great potential. In terms of AIGC technology, it can greatly reduce the difficulty of many tasks.

[0003] Video production is a task of AIGC. However, currently, AIGC can only generate videos with relatively simple plots, and cannot generate videos with richer and more free plots, thus unable to quickly and efficiently meet the personalized video creation needs of users. Summary of the Invention

[0004] Embodiments of the present disclosure provide a video generation method, device, and storage medium to achieve highly free video production.

[0005] In a first aspect, embodiments of the present disclosure provide a video generation method, including:

[0006] Obtain a target object image;

[0007] According to the target object image, obtain at least two alternative video description texts and display them, where any video description text is used to describe at least one of the attribute information, the environment where the target object is located, and the actions of the target object;

[0008] In response to a selection instruction for the alternative video description texts, determine a target video description text from the at least two alternative video description texts;

[0009] Generate a video according to the target object image and the target video description text.

[0010] In a second aspect, embodiments of the present disclosure provide a video generation device, including:

[0011] An obtaining unit, configured to obtain a target object image;

[0012] A video description text determination unit, configured to obtain at least two alternative video description texts according to the target object image and display them, where any video description text is used to describe at least one of the attribute information, the environment where the target object is located, and the actions of the target object;

[0013] A video description text selection unit, configured to determine a target video description text from the at least two alternative video description texts in response to a selection instruction for the alternative video description texts;

[0014] A video generation unit, configured to generate a video according to the target object image and the target video description text.

[0015] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor and a memory;

[0016] The memory stores computer-executable instructions;

[0017] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video generation method described in the first aspect and various possible designs of the first aspect as above.

[0018] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video generation method described in the first aspect and various possible designs of the first aspect as above is implemented.

[0019] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including computer-executable instructions. When a processor executes the computer-executable instructions, the video generation method described in the first aspect and various possible designs of the first aspect as above is implemented.

[0020] The video generation method, device, and storage medium provided by the embodiments of the present disclosure obtain a target object image; according to the target object image, obtain at least two alternative video description texts and display them, where any video description text is used to describe at least one of the attribute information, the environment where the target object is located, and the actions of the target object; in response to a selection instruction for the alternative video description texts, determine a target video description text from the at least two alternative video description texts; and generate a video according to the target object image and the target video description text. The embodiments of the present disclosure first obtain alternative video description texts for the user to select based on the target object image, and then automatically generate a video based on the target video description text selected by the user, which can realize the user's control of the development of the plot before video generation, so that a video with a more free plot can be generated, improving the freedom of video production, increasing the interaction with the user, and improving the user experience. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a scene example diagram of a video generation method provided by an embodiment of the present disclosure;

[0023] Figure 2 It is a schematic flowchart of a video generation method provided by an embodiment of the present disclosure;

[0024] Figure 3 It is a schematic diagram of an interface provided by an embodiment of the present disclosure;

[0025] Figure 4 It is a schematic diagram of an interface provided by another embodiment of the present disclosure;

[0026] Figure 5 It is a schematic flowchart of a video generation method provided by another embodiment of the present disclosure;

[0027] Figure 6 It is a block diagram of the structure of a video generation device provided by an embodiment of the present disclosure;

[0028] Figure 7 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0030] Currently, AIGC can only generate videos with relatively simple plots, and cannot generate videos with richer and more free plots, and cannot quickly and efficiently meet the personalized video creation needs of users. If you want to generate videos with complex plots, you need to first write complex plot scripts and then generate videos through AIGC, and its efficiency is relatively low and the user experience is poor.

[0031] To solve the above technical problems, the present disclosure provides a video generation method, which includes obtaining an image of a target object; obtaining at least two alternative video description texts based on the image of the target object and displaying them, where any video description text is used to describe at least one of the attribute information, the environment where the target object is located, and the actions of the target object; in response to a selection instruction for the alternative video description texts, determining a target video description text from the at least two alternative video description texts; and generating a video based on the image of the target object and the target video description text. In the present disclosure, based on the image of the target object, alternative video description texts are first obtained for the user to select, and then a video is automatically generated based on the target video description text selected by the user. It is possible to realize that the user controls the development of the plot before video generation, so that a video with a more free plot can be generated, the degree of freedom of video production is improved, and at the same time, the interaction with the user is increased, improving the user experience.

[0032] The video generation method provided by the present disclosure can be applied to electronic devices such as terminal devices or servers. For example, if it is applied to a terminal device, a video generation model and a video description text library need to be pre-deployed on the terminal device. After obtaining an image of a target object (such as a user taking or uploading an image of a target object), at least two alternative video description texts can be obtained based on the image of the target object and the video description text library and displayed. The user can select from the displayed alternative video description texts to determine the target video description text. Further, the terminal device can call the video generation model based on the image of the target object and the target video description text to generate a video. Optionally, the last frame or multiple frames of the generated video segment can be used as the image of the target object, and the above process can be repeated to generate the next video segment; by continuously repeating, multiple video segments can be obtained, and the multiple video segments can be sequentially spliced to obtain a complete target video.

[0033] If it is applied to a server, then as Figure 1As shown in the figure, it is necessary to pre-deploy a video generation model and a video description text library on the server. After the terminal device obtains the target object image (for example, the user takes or uploads a target object image), the terminal device can send the target object image to the server. The server can obtain at least two alternative video description texts based on the target object image and the video description text library, and send them to the terminal device for display. The user can select the alternative video description texts displayed on the terminal device to determine the target video description text. Then, the terminal device sends a selection instruction for the alternative video description texts to the server. The server can determine the target video description text according to the selection instruction. Further, according to the target object image and the target video description text, the server can call the video generation model to generate a video, and send the generated video (or video segment) to the terminal device for playback. Optionally, the server can further use the last frame or multiple frames of the generated video segment as the target object image (it may not be necessary to be sent by the terminal device again), and repeat the above process to generate the next video segment; by continuously repeating, multiple video segments can be obtained, and the multiple video segments can be spliced in sequence to obtain a complete target video, which is finally sent to the terminal device.

[0034] It should be noted that the user information and data involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0035] The video generation method of the present disclosure will be introduced in detail below in combination with specific embodiments.

[0036] Reference Figure 2 , Figure 2 is a schematic flowchart of a video generation method provided by an embodiment of the present disclosure. The method of this embodiment can be applied to electronic devices such as terminal devices or servers. The video generation method includes:

[0037] S201. Obtain a target object image.

[0038] In this embodiment, the target object can be a real person, a real animal, a cartoon character, a cartoon animal or other possible objects, and the target object image is an image including the target object.

[0039] Optionally, at the beginning, for the first video segment to be generated (or in the case where only one video segment needs to be generated), the target object image can be a target object image taken or uploaded by the user, or a target object image obtained through other means.

[0040] Optionally, in this embodiment, multiple video segments can also be generated. To ensure the coherent connection between video segments, the last frame or multiple frames of the previously generated video segment (including the target object) can be used as the target object image, and the video generation method can be continued to generate the next video segment.

[0041] S202. Obtain at least two alternative video description texts based on the target object image and display them, where any video description text is used to describe at least one of the attribute information, the environment where the target object is located, and the actions of the target object.

[0042] In this embodiment, the video description text is used to describe one or more of the attribute information (such as the character, dressing, etc.), the environment where the target object is located, and the actions of the target object in the video, so as to be used as the input of the video generation model. For example, the video description text can be "On a grassland, blue sky, white clouds, a young man, a ranger, riding a white horse, with a beard, wearing a cloak".

[0043] Optionally, in this embodiment, based on the target object image and a preset video description text library, or at least two alternative video description texts, the video description text library can be pre-configured and used as the basis for obtaining alternative video description texts. Optionally, the video description text library can include complete preset video description texts; or, the video description text library can also include different types of description word sets, such as but not limited to an object attribute information description word set, an environment description word set, an action description word set, etc. The object attribute information description word set includes multiple preset object attribute information description words, the environment description word set includes multiple preset environment description words, and the action description word set includes multiple preset action description words. Alternative description words can be selected from different types of description word sets respectively, and the alternative description words of different types can be combined to generate alternative video description texts. Of course, the alternative video description texts can also be obtained by other means, which is not limited in this embodiment.

[0044] Among them, optionally, in the case of using a video generation model to generate a video in S204 below, if the video generation model is any possible AI model such as a Large Language Model (LLM) or an AIGC (Generative Artificial Intelligence) model, the video description text can also be called a Prompt. A Prompt is a structured text used as the input of a large language model or an AIGC model.

[0045] Further, when obtaining alternative video description texts for the target object image, the obtained alternative video description texts need to match the target object image. For example, if the light in the target object image is relatively bright, the environment where the target object is located in the alternative video description text can also be a relatively bright environment. Another example is that if the makeup of the target object in the target object image is relatively delicate, the target object in the alternative video description text can be some characters with delicate makeup and can also be equipped with delicate costumes, and so on.

[0046] Therefore, at least two alternative video description texts can be obtained based on the video description text library according to the target object image and displayed for the user to select. Optionally, the at least two alternative video description texts can be selected in the form of controls. In addition, considering that the alternative video description texts may be relatively long, when displaying, the identification information of the alternative video description texts can be displayed. The identification information is relatively short and is used to identify the alternative video description texts or summarize the alternative video description texts. For example, the identification information of the alternative video description text is "fighting with soldiers", while the actual alternative video description text describes many characters, environments, actions, etc. involved in the fighting process.

[0047] S203. In response to the selection instruction for the alternative video description text, determine the target video description text from the at least two alternative video description texts.

[0048] In this embodiment, the alternative video description texts are displayed on the interface, and the user can trigger the required alternative video description texts. The triggering methods include but are not limited to clicking, swiping, dragging, etc., to select the target video description text from the alternative video description texts.

[0049] S204. Generate a video according to the target object image and the target video description text.

[0050] In this embodiment, any method capable of generating a video based on an image and a video description text can be adopted, such as a deep learning model, etc.

[0051] Optionally, in this embodiment, a video generation model can be adopted to realize the generation of a video by mobilizing the video generation model according to the target object image and the target video description text.

[0052] Among them, the video generation model can be used to automatically generate a video. Its input is a picture and a video description text. Some training data (including pictures, video description texts, and corresponding videos) can be obtained in advance to train the video generation model. Optionally, the video generation model is a large language model, an AIGC model, or any other deep learning model, which can understand the meaning of the video description text and realize the generation of a video by processing the picture according to the video description text.

[0053] In this embodiment, after obtaining the target object image and the target video description text, a video generation model can be called to input the target object image and the target video description text into the video generation model for processing, and the video generation model automatically generates a video (or a video clip).

[0054] Furthermore, the generated video (or video clip) can be played to be presented to the user in real time, so as to facilitate the user to judge whether it meets their own needs and to facilitate the user to make decisions on the development of the next plot.

[0055] Optionally, in this embodiment, the above S201-S204 can be repeatedly executed to generate multiple video clips, and then the multiple video clips are sequentially spliced (for example, each generated video clip is spliced after the previously generated video clip), to generate a target video. In order to ensure the coherence between multiple video clips, when obtaining the target object image, the last frame or multiple frames of the previously generated video clip can be used as the target object image.

[0056] The video generation method provided in this embodiment includes: obtaining a target object image; obtaining at least two alternative video description texts according to the target object image and displaying them, where any video description text is used to describe at least one of the attribute information, the environment, and the actions of the target object; in response to a selection instruction for the alternative video description texts, determining a target video description text from the at least two alternative video description texts; and performing video generation according to the target object image and the target video description text. Based on the target object image, this embodiment first obtains alternative video description texts for the user to select, and then automatically performs video generation based on the target video description text selected by the user, which can realize the control of the plot development by the user before video generation, so as to generate a video with a more free plot, improve the freedom of video production, and also increase the interaction with the user and improve the user experience.

[0057] Based on any of the above embodiments, initially, for the first video clip to be generated (that is, the target object appears in the first video clip as the starting part of the plot), the user can take or upload a target object image (such as a photo), and of course, the target object image can also be obtained through other means, which is not limited in this embodiment.

[0058] Furthermore, the obtaining at least two alternative video description texts according to the target object image includes:[[]]

[0059] Obtaining at least two alternative video description texts for the first video clip to be generated according to the target object image and a preset video description text library.

[0060] In this embodiment, since there is no previous plot for the first video segment to be generated initially, the plot needs the target object to appear at this time. According to the target object image captured or uploaded by the user and the video description text library, at least two alternative video description texts when the target object appears can be generated first, that is, at least two alternative video description texts for the first video segment to be generated. The alternative video description texts for the first video segment to be generated may include one or more of the attribute information (such as character, dressing, etc.), the environment where the target object is located, actions, etc. when the target object appears. In addition, it may include the transition process from the target object image to the appearance. After obtaining at least two alternative video description texts for the first video segment to be generated, the at least two alternative video description texts (or their identifiers) for the first video segment to be generated can be displayed on the interface. For example Figure 3 as shown, it may include "becoming a ranger on horseback" and "becoming a cyberpunk robot", which are displayed in the form of controls respectively. The user can trigger them by clicking, dragging, etc. to select the target video description text for the first video segment to be generated.

[0061] Optionally, when obtaining at least two alternative video description texts for the first video segment to be generated according to the target object image and the preset video description text library, at least one of the following methods can be used to implement:

[0062] Method 1:

[0063] 1) Obtain the image features of the target object image;

[0064] 2) Obtain the third relevance index between the image features of the target object image and the preset video description text corresponding to the first video segment in the video description text library, where the video description text library includes multiple preset video description texts corresponding to the first video segment;

[0065] 3) Select at least two preset video description texts corresponding to the first video segment according to the third relevance index, and determine them as at least two alternative video description texts for the first video segment to be generated.

[0066] In this embodiment, a plurality of complete preset video description texts corresponding to the first video segment can be pre-configured in the target video description text library. Each preset video description text can include one or more of the object attribute information, environment, actions, etc. of the appearance scene that have been determined. When it is necessary to obtain the alternative video description text of the first video segment to be generated for the target object image, image features can be extracted from the target object image. Any known method can be used for the method of extracting image features. Then, a third correlation index between the image features of the target object image and each preset video description text corresponding to the first video segment in the video description text library is obtained. Based on the magnitude of the third correlation index, at least two preset video description texts with the largest third correlation index are selected from each preset video description text corresponding to the first video segment as the alternative video description texts. Among them, any known method for obtaining the correlation between image features and text can be used to obtain the third correlation index between the image features of the target object image and each preset video description text corresponding to the first video segment in the video description text library, and there is no limitation here.

[0067] Method 2:

[0068] 1) Obtain the image features of the target object image;

[0069] 2) Match the image features of the target object image with the image features of the historical object images in the video description text library to determine the matching historical object images, where the target video description texts of the historical object images are included in the video description text library;

[0070] 3) Determine the target video description text of the first video segment corresponding to the matching historical object image as the alternative video description text of the first video segment to be generated.

[0071] In this embodiment, the target video description text library stores the target video description text (or alternative video description text) of the processed historical object images, including storing the target video description text (or alternative video description text) of the first video segment corresponding to the historical object images. When it is necessary to obtain the alternative video description text of the first video segment to be generated for the target object image, image features can be extracted from the target object image. Any known method can be used for the method of extracting image features. Then, the image features of the target object image are obtained and matched with the image features of the historical object images to determine the historical object image that matches the target object image. That is, the target object image and the historical object image have a high similarity. The target video description text (or alternative video description text) of the first video segment corresponding to the historical object image can be referred to. That is, the target video description text (or alternative video description text) of the first video segment corresponding to the matched historical object image can be determined as the alternative video description text of the first video segment to be generated for the target object image. The matching of the image features of the target object image and the image features of the historical object image can adopt any known image matching method, which is not limited here.

[0072] Method 3:

[0073] 1) Obtain the image features of the target object image;

[0074] 2) Obtain the fourth correlation index between the image features of the target object image and each preset descriptor in the set of different type descriptors corresponding to the first video segment in the video description text library, where the set of different type descriptors includes at least one of the set of object attribute information descriptors, the set of environment descriptors, and the set of action descriptors;

[0075] 3) Select different type preset descriptors corresponding to the first video segment according to the fourth correlation index, determine them as different type alternative descriptors, and combine the different type alternative descriptors to generate at least two alternative video description texts of the first video segment to be generated.

[0076] In this embodiment, the target video description text library may include sets of descriptive words of different types corresponding to the first video segment, where the sets of descriptive words of different types include at least one of a set of object attribute information descriptive words, a set of environment descriptive words, and a set of action descriptive words. The set of object attribute information descriptive words includes multiple preset object attribute information descriptive words, the set of environment descriptive words includes multiple preset environment descriptive words, and the set of action descriptive words includes multiple preset action descriptive words. When it is necessary to obtain alternative video description texts for the first video segment to be generated for the target object image, image features can be extracted from the target object image, where any known method can be used for the method of extracting image features. Then, a fourth correlation index between the image features of the target object image and each preset descriptive word in the sets of descriptive words of different types corresponding to the first video segment in the video description text library is obtained. Based on the magnitude of the fourth correlation index, at least two preset descriptive words with the largest fourth correlation index are respectively selected from each type of preset descriptive word set corresponding to the first video segment as alternative descriptive words of each type, and the alternative descriptive words of different types are combined to generate at least two alternative video description texts for the first video segment to be generated. Among them, for obtaining the fourth correlation index between the image features of the target object image and each preset descriptive word in the sets of descriptive words of different types corresponding to the first video segment in the video description text library, any known method for obtaining the correlation between image features and text (words) can be used, and no limitation is imposed here.

[0077] Based on the above embodiment, on the basis of the generated video segment, the generation of subsequent video segments can continue to make the plot continue. To ensure the coherent connection between video segments, the last frame or multiple frames of the previously generated video segment (including the target object) can be used as the target object image, and the video generation method can be continued to be executed.

[0078] Specifically, after obtaining the last frame or multiple frames of the previously generated video segment, it can be used as the target object image. According to the target object image and the preset video description text library, at least two alternative video description texts can be obtained, which may specifically include:

[0079] According to the target object image, the target video description text of the previously generated video segment, and the preset video description text library, at least two alternative video description texts for the turning event are obtained.

[0080] In this embodiment, in order to ensure the continuity of the plot, at least two alternative video description texts of the turning event need to be obtained. The turning event is a turning point relative to the previously generated video segment, that is, at least one of the object attribute information, environment, and action changes relative to the previously generated video segment. By providing at least two alternative video description texts of the turning event to the user, the user can select the turning event to control the development direction of the plot.

[0081] When obtaining the alternative video description text of the turning event, in addition to relying on the target object image, since the target video description text of the previously generated video segment also contains a large amount of information that affects the subsequent plot development, it is necessary to combine the target video description text of the previously generated video segment. Therefore, in this embodiment, at least two alternative video description texts of the turning event can be obtained according to the target object image, the target video description text of the previously generated video segment, and the preset video description text library.

[0082] Optionally, during the playback of the previously generated video segment, at least two alternative video description texts of the turning event can be obtained according to the target object image, the target video description text of the previously generated video segment, and the preset video description text library, and after the previously generated video segment is played, at least two alternative video description texts of the turning event are displayed for the user to select to determine the next plot direction. For example Figure 4 As shown, after the previously generated video segment is played (it can be shown that the last frame has been played), at least two alternative video description texts (or their identifiers) of the turning event can be displayed on the interface, such as "fighting with soldiers" and "shaking hands with the old man", etc., which are displayed in the form of controls respectively, and the user can trigger them by clicking, dragging, etc. to select the target video description text of the turning event.

[0083] Optionally, when obtaining at least two alternative video description texts of the turning event according to the target object image, the target video description text of the previously generated video segment, and the preset video description text library, at least one of the following methods can be used to implement:

[0084] Method 1:

[0085] 1) Obtain the image features of the target object image;

[0086] 2) Obtain the first relevance index between the image features of the target object image, the target video description text of the previously generated video segment, and the preset video description text of the turning event in the video description text library, where the video description text library includes multiple preset video description texts of the turning event;

[0087] 3) Select at least two preset video description texts of the turning event according to the first relevance index, and determine them as at least two alternative video description texts of the turning event.

[0088] In this embodiment, a plurality of complete preset video description texts of the turning event can be pre-configured in the target video description text library. Each preset video description text can include one or more of the object attribute information, environment, actions, etc. of the determined turning event. When it is necessary to obtain alternative video description texts of the turning event for the target object image, image features can be extracted from the target object image. The method for extracting image features can adopt any known method. Then, obtain the image features of the target object image and the first relevance index between the target video description text of the previous generated video segment and each preset video description text of the turning event in the video description text library. Then, based on the magnitude of the first relevance index, select at least two preset video description texts with the largest first relevance index from each preset video description text of the turning event as alternative video description texts. Among them, to obtain the first relevance index between the image features of the target object image and the target video description text of the previous generated video segment and each preset video description text of the turning event in the video description text library, any known method for obtaining the relevance between image features and text can be adopted, and no limitation is imposed here.

[0089] Method 2:

[0090] 1) Obtain the image features of the target object image;

[0091] 2) Obtain the second relevance index between the image features of the target object image and the target video description text of the previous generated video segment and each preset description word in different type description word sets of the turning event in the video description text library, where the different type description word sets include at least one of the object attribute information description word set, the environment description word set, and the action description word set;

[0092] 3) Select different type preset description words of the turning event according to the second relevance index, determine them as different type alternative description words, and combine the different type alternative description words to generate at least two alternative video description texts of the turning event.

[0093] In this embodiment, the target video description text library may include sets of different types of description words for turning events (which may be the same as or different from the sets of different types of description words for the first video segment). The sets of different types of description words include at least one of a set of object attribute information description words, a set of environment description words, and a set of action description words. The set of object attribute information description words includes multiple preset object attribute information description words, the set of environment description words includes multiple preset environment description words, and the set of action description words includes multiple preset action description words. When it is necessary to obtain alternative video description texts for turning events for the target object image, image features can be extracted from the target object image. Any known method can be used for extracting the image features. Then, the second relevance index between the image features of the target object image, the target video description text of the previous generated video segment, and each preset description word in the sets of different types of description words for turning events in the video description text library is obtained. Based on the magnitudes of the second relevance indices, at least two preset description words with the largest second relevance index are respectively selected from each type of preset description word set for turning events as alternative description words for each type. After combining the alternative description words of different types, at least two alternative video description texts for turning events are generated. Among them, any known method for obtaining the relevance between image features and text (words) can be used to obtain the second relevance index between the image features of the target object image, the target video description text of the previous generated video segment, and each preset description word in the sets of different types of description words for turning events in the video description text library, and there is no limitation here.

[0094] Based on the above embodiments, as Figure 5As shown, initially, the user can take or upload an image of the target object. Based on the image of the target object and in combination with the video description text library (prompt library), determine the alternative video description texts for the first video segment to be generated, such as the entrance prompt1 and entrance prompt2, and display them. The user can select the target video description text for the first video segment to be generated: the target entrance prompt. Then, based on the image of the target object taken or uploaded by the user and the target video description text of the first video segment to be generated, call the video generation model to generate the first video segment (i.e., the video segment of the entrance scene) and play it. It is also possible to obtain the last frame (or multiple frames) of the first video segment, determine it as the image of the target object, and based on this image of the target object and the target video description text of the first video segment (the target entrance prompt), and in combination with the video description text library (prompt library), determine the alternative video description texts for turning event 1, such as turning prompt1 and turning prompt2, and display them. The user can select the target video description text for turning event 1: the target turning prompt. Then, based on this image of the target object and the target video description text of the turning event (the target turning prompt), call the video generation model to generate the video segment of turning event 1 (which can be spliced after the first video segment) and play it. It is also possible to continue to obtain the last frame (or multiple frames) of the video segment of turning event 1, determine it as the image of the target object, and continue to determine the alternative video description texts for turning event 2, select the target video description text, and generate the video segment on the basis of the image of the target object and the target video description text of turning event 1. By repeatedly executing the above process, multiple video segments can be obtained, and the multiple video segments can be spliced in sequence to obtain a target video with a coherent and rich plot.

[0095] Based on any of the above embodiments, the camera movement method can also be controlled so that the generated video (or video segment) is camera-moved in the determined target camera movement method.

[0096] Optionally, when responding to the selection instruction for the alternative video description text, the target camera movement method corresponding to the target video description text can also be determined. That is, in this embodiment, different alternative video description texts can correspond to different camera movement methods. When the user selects an alternative video description text as the target video description text, the camera movement method corresponding to this alternative video description text is also determined as the target camera movement method corresponding to the target video description text.

[0097] Further, in any of the above embodiments, when generating a video based on the target object image and the target video description text, the video can be generated according to the target object image, the target video description text, and the target camera movement mode. Optionally, the target camera movement mode can also be input as an input data into the video generation model, or the target camera movement mode can be added to the target video description text and then input into the video generation model together.

[0098] Optionally, when at least two alternative video description texts are displayed on the interface, a specific arrangement method can be adopted, such as arranging them left and right. When the user makes a selection, the user can trigger the alternative video description texts displayed on the interface, such as clicking, swiping, dragging, etc. For the trigger instructions for alternative video description texts in different positions, different camera movement modes can be determined. That is, according to the position of the target video description text on the interface, the target camera movement mode corresponding to the target video description text can be determined. For example, if the user selects the alternative video description text on the left as the target video description text, it can be determined that the target camera movement mode is that the camera first moves to the left and then returns to the middle, that is, in the video, the picture first turns left and then returns to the front; if the user selects the alternative video description text on the right as the target video description text, it can be determined that the target camera movement mode is that the camera first moves to the right and then returns to the middle, that is, in the video, the picture first turns right and then returns to the front. Thus, it can be distinguished in the video that the user has selected different alternative video description texts, and it also enables the user to feel that the video responds to the user's operations, increasing the interactivity and improving the interaction experience.

[0099] Optionally, in the first segment of the video, the camera movement mode can also adopt a fixed camera movement mode. For example, the picture only moves backward to show the whole body of the target object, making the target object present an effect of moving forward towards the audience in front of the screen, that is, presenting an appearance effect.

[0100] Corresponding to the video generation method in the above embodiment Figure 6 is a structural block diagram of a video generation device provided by an embodiment of the present disclosure. For ease of description, only parts related to the embodiments of the present disclosure are shown. Referring to Figure 6 the video generation device 600 includes: an acquisition unit 601, a video description text determination unit 602, a video description text selection unit 603, and a video generation unit 604.

[0101] Among them, the acquisition unit 601 is configured to acquire a target object image;

[0102] The video description text determination unit 602 is configured to obtain at least two alternative video description texts according to the target object image and display them, where any video description text is used to describe at least one of the attribute information, the environment where the target object is located, and the actions of the target object;

[0103] A video description text selection unit 603, configured to determine a target video description text from the at least two alternative video description texts in response to a selection instruction for an alternative video description text;

[0104] A video generation unit 604, configured to generate a video according to the target object image and the target video description text.

[0105] In one or more embodiments of the present disclosure, when obtaining the target object image, the obtaining unit is configured to:

[0106] Obtain the last frame or multiple frames of the previous generated video segment and determine them as the target object image;

[0107] Correspondingly, after generating a video according to the target object image and the target video description text, the video generation unit is further configured to:

[0108] Stitch the generated video segment with the previous generated video segment.

[0109] In one or more embodiments of the present disclosure, when obtaining at least two alternative video description texts according to the target object image, the video description text determination unit is configured to:

[0110] Obtain at least two alternative video description texts for turning events according to the target object image, the target video description text of the previous generated video segment, and a preset video description text library.

[0111] In one or more embodiments of the present disclosure, when obtaining at least two alternative video description texts for turning events according to the target object image, the target video description text of the previous generated video segment, and a preset video description text library, the video description text determination unit is configured to:

[0112] Obtain the image feature of the target object image;

[0113] Obtain the image feature of the target object image and a first relevance index between the target video description text of the previous generated video segment and the preset video description texts of turning events in the video description text library, where the video description text library includes multiple preset video description texts of turning events;

[0114] Select at least two preset video description texts of turning events according to the first relevance index and determine them as at least two alternative video description texts of turning events.

[0115] In one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two alternative video description texts of a turning event according to the target object image, the target video description text of the previous generated video segment, and a preset video description text library, it is configured to:

[0116] Obtain the image features of the target object image;

[0117] Obtain the second relevance index between the image features of the target object image, the target video description text of the previous generated video segment, and each preset description word in the set of different type description words of the turning event in the video description text library, where the set of different type description words includes at least one of the set of object attribute information description words, the set of environment description words, and the set of action description words;

[0118] Select preset description words of different types of the turning event according to the second relevance index, determine them as alternative description words of different types, and combine the alternative description words of different types to generate at least two alternative video description texts of the turning event.

[0119] In one or more embodiments of the present disclosure, when the obtaining unit obtains the target object image, it is configured to:

[0120] For the first video segment to be generated, obtain the target object image captured or uploaded by the user;

[0121] Correspondingly, obtaining at least two alternative video description texts according to the target object image includes:

[0122] Obtain at least two alternative video description texts of the first video segment to be generated according to the target object image and a preset video description text library.

[0123] In one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two alternative video description texts of the first video segment to be generated according to the target object image and a preset video description text library, it is configured to:

[0124] Obtain the image features of the target object image;

[0125] Obtain the third relevance index between the image features of the target object image and the preset video description text corresponding to the first video segment in the video description text library, where the video description text library includes multiple preset video description texts corresponding to the first video segment;

[0126] Select at least two preset video description texts corresponding to the first video segment according to the third relevance index, and determine them as at least two alternative video description texts for the first video segment to be generated.

[0127] In one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two alternative video description texts for the first video segment to be generated according to the target object image and the preset video description text library, it is used for:

[0128] Obtain the image features of the target object image;

[0129] Match the image features of the target object image with the image features of the historical object images in the video description text library to determine the matching historical object images, where the target video description texts of the historical object images are included in the video description text library;

[0130] Determine the target video description text of the first video segment corresponding to the matching historical object image as the alternative video description text for the first video segment to be generated.

[0131] In one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two alternative video description texts for the first video segment to be generated according to the target object image and the preset video description text library, it is used for:

[0132] Obtain the image features of the target object image;

[0133] Obtain the fourth relevance index between the image features of the target object image and each preset description word in the set of different type description words corresponding to the first video segment in the video description text library, where the set of different type description words includes at least one of the set of object attribute information description words, the set of environment description words, and the set of action description words;

[0134] Select different type preset description words corresponding to the first video segment according to the fourth relevance index, determine them as different type alternative description words, and combine the different type alternative description words to generate at least two alternative video description texts for the first video segment to be generated.

[0135] In one or more embodiments of the present disclosure, when the video generation unit performs video generation according to the target object image and the target video description text, it is used for:

[0136] Call a video generation model according to the target object image and the target video description text to perform video generation.

[0137] In one or more embodiments of the present disclosure, when the video description text selection unit 603 responds to a selection instruction for an alternative video description text, it is further configured to:

[0138] Determine the target camera movement mode corresponding to the target video description text;

[0139] Correspondingly, when the video generation unit 604 generates a video according to the target object image and the target video description text, it is configured to:

[0140] Generate a video according to the target object image, the target video description text, and the target camera movement mode.

[0141] In one or more embodiments of the present disclosure, the selection instruction for the alternative video description text is a trigger instruction for the alternative video description text displayed in the interface; correspondingly, when the video description text selection unit 603 determines the target camera movement mode corresponding to the target video description text, it is configured to:

[0142] Determine the target camera movement mode corresponding to the target video description text according to the position of the target video description text in the interface.

[0143] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principles and technical effects are similar, which will not be elaborated here in this embodiment.

[0144] Refer to Figure 7 , which shows a schematic structural diagram of an electronic device 700 suitable for implementing the embodiments of the present disclosure. The electronic device 700 can be a terminal device or a server. Among them, the terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0145] Such as Figure 7As shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0146] Generally, the following devices may be connected to the I / O interface 705: an input device 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 708 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 709. The communication device 709 may allow the electronic device 700 to communicate with other devices wirelessly or wirelesly to exchange data. Although Figure 7 the electronic device 700 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.

[0147] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0148] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0149] The above computer-readable medium can be included in the above electronic device; it can also exist separately and not be assembled into the electronic device.

[0150] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.

[0151] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, execute as a stand-alone software package, partly on the user's computer and partly on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0153] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0154] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0155] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0156] In a first aspect, according to one or more embodiments of the present disclosure, there is provided a video generation method, including:

[0157] Obtaining a target object image;

[0158] According to the target object image, obtaining at least two alternative video description texts and displaying them, where any of the video description texts is used to describe at least one of the attribute information, the environment where the target object is located, and the actions of the target object;

[0159] In response to a selection instruction for the alternative video description texts, determining a target video description text from the at least two alternative video description texts;

[0160] According to the target object image and the target video description text, performing video generation.

[0161] According to one or more embodiments of the present disclosure, the obtaining of the target object image includes:

[0162] Obtaining the last frame or multiple frames of the previous generated video segment and determining them as the target object image;

[0163] Correspondingly, after performing video generation according to the target object image and the target video description text, it further includes:

[0164] Splicing the generated video segment with the previous generated video segment.

[0165] According to one or more embodiments of the present disclosure, the obtaining of at least two alternative video description texts according to the target object image includes:

[0166] Obtain at least two alternative video description texts for the turning event according to the target object image, the target video description text of the previous generated video segment, and a preset video description text library.

[0167] According to one or more embodiments of the present disclosure, the obtaining at least two alternative video description texts for the turning event according to the target object image, the target video description text of the previous generated video segment, and a preset video description text library includes:

[0168] Obtain the image features of the target object image;

[0169] Obtain a first relevance index between the image features of the target object image, the target video description text of the previous generated video segment, and the preset video description texts of the turning event in the video description text library, where the video description text library includes multiple preset video description texts of the turning event;

[0170] Select at least two preset video description texts of the turning event according to the first relevance index, and determine them as at least two alternative video description texts of the turning event.

[0171] According to one or more embodiments of the present disclosure, the obtaining at least two alternative video description texts for the turning event according to the target object image, the target video description text of the previous generated video segment, and a preset video description text library includes:

[0172] Obtain the image features of the target object image;

[0173] Obtain a second relevance index between the image features of the target object image, the target video description text of the previous generated video segment, and each preset description word in the set of different type description words of the turning event in the video description text library, where the set of different type description words includes at least one of a set of object attribute information description words, a set of environment description words, and a set of action description words;

[0174] Select preset description words of different types of the turning event according to the second relevance index, determine them as alternative description words of different types, and combine the alternative description words of different types to generate at least two alternative video description texts of the turning event.

[0175] According to one or more embodiments of the present disclosure, the obtaining the target object image includes:

[0176] For the first video segment to be generated, obtain the target object image captured or uploaded by the user;

[0177] Correspondingly, the obtaining at least two alternative video description texts according to the target object image includes:

[0178] Obtain at least two alternative video description texts for the first video segment to be generated according to the target object image and a preset video description text library.

[0179] According to one or more embodiments of the present disclosure, the obtaining at least two alternative video description texts for the first video segment to be generated according to the target object image and a preset video description text library includes:

[0180] Obtain the image features of the target object image;

[0181] Obtain a third relevance index between the image features of the target object image and a preset video description text corresponding to the first video segment in the video description text library, where the video description text library includes multiple preset video description texts corresponding to the first video segment;

[0182] Select at least two preset video description texts corresponding to the first video segment according to the third relevance index, and determine them as at least two alternative video description texts for the first video segment to be generated.

[0183] According to one or more embodiments of the present disclosure, the obtaining at least two alternative video description texts for the first video segment to be generated according to the target object image and a preset video description text library includes:

[0184] Obtain the image features of the target object image;

[0185] Match the image features of the target object image with the image features of historical object images in the video description text library to determine the matched historical object images, where the video description text library includes target video description texts of the historical object images;

[0186] Determine the target video description text of the first video segment corresponding to the matched historical object image as the alternative video description text for the first video segment to be generated.

[0187] According to one or more embodiments of the present disclosure, the obtaining at least two alternative video description texts for the first video segment to be generated according to the target object image and a preset video description text library includes:

[0188] Obtain the image features of the target object image;

[0189] Obtain a fourth correlation index between the image features of the target object image and each preset descriptor in the set of different types of descriptors corresponding to the first video segment in the video description text library, where the set of different types of descriptors includes at least one of a set of object attribute information descriptors, a set of environment descriptors, and a set of action descriptors;

[0190] Select different types of preset descriptors corresponding to the first video segment according to the fourth correlation index, determine them as different types of alternative descriptors, and combine the different types of alternative descriptors to generate at least two alternative video description texts for the to-be-generated first video segment.

[0191] According to one or more embodiments of the present disclosure, the video generation based on the target object image and the target video description text includes:

[0192] Call a video generation model according to the target object image and the target video description text to perform video generation.

[0193] According to one or more embodiments of the present disclosure, when responding to a selection instruction for an alternative video description text, it further includes:

[0194] Determine the target camera movement mode corresponding to the target video description text;

[0195] Correspondingly, the video generation based on the target object image and the target video description text includes:

[0196] Perform video generation according to the target object image, the target video description text, and the target camera movement mode.

[0197] According to one or more embodiments of the present disclosure, the selection instruction for the alternative video description text is a trigger instruction for the alternative video description text displayed in the interface; correspondingly, the determination of the target camera movement mode corresponding to the target video description text includes:

[0198] Determine the target camera movement mode corresponding to the target video description text according to the position of the target video description text in the interface.

[0199] In a second aspect, according to one or more embodiments of the present disclosure, there is provided a video generation device, including:

[0200] An acquisition unit for acquiring a target object image;

[0201] A video description text determination unit for acquiring at least two alternative video description texts according to the target object image and displaying them, where any video description text is used to describe at least one of the attribute information, the environment, and the actions of the target object;

[0202] A video description text selection unit, configured to determine a target video description text from the at least two alternative video description texts in response to a selection instruction for the alternative video description texts;

[0203] A video generation unit, configured to generate a video according to the target object image and the target video description text.

[0204] According to one or more embodiments of the present disclosure, when acquiring the target object image, the acquisition unit is configured to:

[0205] Acquire the last frame or multiple frames of the previous generated video segment, and determine them as the target object image;

[0206] Correspondingly, after generating a video according to the target object image and the target video description text, the video generation unit is further configured to:

[0207] Splice the generated video segment with the previous generated video segment.

[0208] According to one or more embodiments of the present disclosure, when acquiring at least two alternative video description texts according to the target object image, the video description text determination unit is configured to:

[0209] Acquire at least two alternative video description texts of the turning event according to the target object image, the target video description text of the previous generated video segment, and a preset video description text library.

[0210] According to one or more embodiments of the present disclosure, when acquiring at least two alternative video description texts of the turning event according to the target object image, the target video description text of the previous generated video segment, and a preset video description text library, the video description text determination unit is configured to:

[0211] Acquire the image feature of the target object image;

[0212] Acquire a first relevance index between the image feature of the target object image, the target video description text of the previous generated video segment, and the preset video description text of the turning event in the video description text library, where the video description text library includes multiple preset video description texts of the turning event;

[0213] Select at least two preset video description texts of the turning event according to the first relevance index, and determine them as at least two alternative video description texts of the turning event.

[0214] According to one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two alternative video description texts of the turning event according to the target object image, the target video description text of the previous generated video segment, and a preset video description text library, it is used for:

[0215] Obtain the image features of the target object image;

[0216] Obtain the second relevance index between the image features of the target object image, the target video description text of the previous generated video segment, and each preset description word in the set of different type description words of the turning event in the video description text library, where the set of different type description words includes at least one of the set of object attribute information description words, the set of environment description words, and the set of action description words;

[0217] Select preset description words of different types of the turning event according to the second relevance index, determine them as alternative description words of different types, and combine the alternative description words of different types to generate at least two alternative video description texts of the turning event.

[0218] According to one or more embodiments of the present disclosure, when the obtaining unit obtains the target object image, it is used for:

[0219] For the first video segment to be generated, obtain the target object image taken or uploaded by the user;

[0220] Correspondingly, obtaining at least two alternative video description texts according to the target object image includes:

[0221] According to the target object image and a preset video description text library, obtain at least two alternative video description texts of the first video segment to be generated.

[0222] According to one or more embodiments of the present disclosure, when the video description text determination unit obtains at least two alternative video description texts of the first video segment to be generated according to the target object image and a preset video description text library, it is used for:

[0223] Obtain the image features of the target object image;

[0224] Obtain the third relevance index between the image features of the target object image and the preset video description text corresponding to the first video segment in the video description text library, where the video description text library includes multiple preset video description texts corresponding to the first video segment;

[0225] Select at least two preset video description texts corresponding to the first video segment according to the third relevance index, and determine them as at least two alternative video description texts for the to-be-generated first video segment.

[0226] According to one or more embodiments of the present disclosure, when obtaining at least two alternative video description texts for the to-be-generated first video segment based on the target object image and a preset video description text library, the video description text determination unit is configured to:

[0227] Obtain the image features of the target object image;

[0228] Match the image features of the target object image with the image features of historical object images in the video description text library to determine the matched historical object images, where the video description text library includes target video description texts of the historical object images;

[0229] Determine the target video description text of the first video segment corresponding to the matched historical object image as the alternative video description text for the to-be-generated first video segment.

[0230] According to one or more embodiments of the present disclosure, when obtaining at least two alternative video description texts for the to-be-generated first video segment based on the target object image and a preset video description text library, the video description text determination unit is configured to:

[0231] Obtain the image features of the target object image;

[0232] Obtain the fourth relevance index between the image features of the target object image and each preset description word in the set of different type description words corresponding to the first video segment in the video description text library, where the set of different type description words includes at least one of a set of object attribute information description words, a set of environment description words, and a set of action description words;

[0233] Select different type preset description words corresponding to the first video segment according to the fourth relevance index, determine them as different type alternative description words, and combine the different type alternative description words to generate at least two alternative video description texts for the to-be-generated first video segment.

[0234] According to one or more embodiments of the present disclosure, when performing video generation based on the target object image and the target video description text, the video generation unit is configured to:

[0235] Call a video generation model according to the target object image and the target video description text to perform video generation.

[0236] According to one or more embodiments of the present disclosure, when the video description text selection unit responds to a selection instruction for alternative video description texts, it is further configured to:

[0237] Determine the target camera movement mode corresponding to the target video description text;

[0238] Correspondingly, when the video generation unit generates a video according to the target object image and the target video description text, it is configured to:

[0239] Generate a video according to the target object image, the target video description text, and the target camera movement mode.

[0240] According to one or more embodiments of the present disclosure, the selection instruction for the alternative video description text is a trigger instruction for the alternative video description text displayed in the interface; correspondingly, when the video description text selection unit determines the target camera movement mode corresponding to the target video description text, it is configured to:

[0241] Determine the target camera movement mode corresponding to the target video description text according to the position of the target video description text in the interface.

[0242] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one processor and a memory;

[0243] The memory stores computer-executable instructions;

[0244] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video generation method described in the first aspect above and various possible designs of the first aspect.

[0245] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium, in which computer-executable instructions are stored, and when a processor executes the computer-executable instructions, the video generation method described in the first aspect above and various possible designs of the first aspect is implemented.

[0246] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided a computer program product, including computer-executable instructions, and when a processor executes the computer-executable instructions, the video generation method described in the first aspect above and various possible designs of the first aspect is implemented.

[0247] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0248] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0249] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video generation method, characterized in that, Including: Obtain an image of a target object; Based on the image of the target object, obtain at least two alternative video description texts and display them, where any video description text is used to describe at least one of the attribute information, the environment, and the actions of the target object; In response to a selection instruction for the alternative video description text, determine a target video description text from the at least two alternative video description texts; Generate a video based on the image of the target object and the target video description text.

2. The method according to claim 1, wherein The obtaining of the image of the target object includes: Obtain the last frame or multiple frames of the previous generated video segment and determine them as the image of the target object; Correspondingly, after generating the video based on the image of the target object and the target video description text, it further includes: Stitch the generated video segment with the previous generated video segment.

3. The method according to claim 2, characterized in that, The obtaining of at least two alternative video description texts based on the image of the target object includes: Based on the image of the target object, the target video description text of the previous generated video segment, and a preset video description text library, obtain at least two alternative video description texts for a turning event.

4. The method according to claim 3, characterized in that, The obtaining of at least two alternative video description texts for a turning event based on the image of the target object, the target video description text of the previous generated video segment, and a preset video description text library includes: Obtain the image features of the image of the target object; Obtain a first relevance index between the image features of the image of the target object, the target video description text of the previous generated video segment, and the preset video description texts of the turning events in the video description text library, where the video description text library includes multiple preset video description texts of the turning events; Select at least two preset video description texts of the turning event according to the first relevance index and determine them as at least two alternative video description texts of the turning event.

5. The method according to claim 3, wherein The obtaining of at least two alternative video description texts for a turning event based on the image of the target object, the target video description text of the previous generated video segment, and a preset video description text library includes: Obtain the image features of the image of the target object; Obtain a second relevance index between the image features of the image of the target object, the target video description text of the previous generated video segment, and each preset description word in the set of different type description words of the turning events in the video description text library, where the set of different type description words includes at least one of the set of object attribute information description words, the set of environment description words, and the set of action description words; Select preset description words of different types of the turning event according to the second relevance index, determine them as alternative description words of different types, and combine the alternative description words of different types to generate at least two alternative video description texts of the turning event.

6. The method according to claim 1, wherein The obtaining of the image of the target object includes: For the first video segment to be generated, obtain the image of the target object captured or uploaded by the user; Correspondingly, the obtaining of at least two alternative video description texts based on the image of the target object includes: Obtain at least two alternative video description texts for the first video segment to be generated according to the target object image and a preset video description text library.

7. The method according to claim 6, characterized in that, The obtaining at least two alternative video description texts for the first video segment to be generated according to the target object image and a preset video description text library includes: Obtain the image features of the target object image; Obtain a third relevance index between the image features of the target object image and the preset video description texts corresponding to the first video segment in the video description text library, where the video description text library includes multiple preset video description texts corresponding to the first video segment; Select at least two preset video description texts corresponding to the first video segment according to the third relevance index, and determine them as at least two alternative video description texts for the first video segment to be generated.

8. The method according to claim 6, wherein The obtaining at least two alternative video description texts for the first video segment to be generated according to the target object image and a preset video description text library includes: Obtain the image features of the target object image; Match the image features of the target object image with the image features of the historical object images in the video description text library to determine the matched historical object images, where the video description text library includes the target video description texts of the historical object images; Determine the target video description texts of the first video segment corresponding to the matched historical object images as the alternative video description texts for the first video segment to be generated.

9. The method according to claim 6, wherein The obtaining at least two alternative video description texts for the first video segment to be generated according to the target object image and a preset video description text library includes: Obtain the image features of the target object image; Obtain a fourth relevance index between the image features of the target object image and each preset description word in the set of different type description words corresponding to the first video segment in the video description text library, where the set of different type description words includes at least one of a set of object attribute information description words, a set of environment description words, and a set of action description words; Select different type preset description words corresponding to the first video segment according to the fourth relevance index, determine them as different type alternative description words, and combine the different type alternative description words to generate at least two alternative video description texts for the first video segment to be generated.

10. The method according to any one of claims 1-9, characterized in that, The performing video generation according to the target object image and the target video description text includes: Call a video generation model according to the target object image and the target video description text to perform video generation.

11. The method according to any one of claims 1-9, characterized in that, When responding to a selection instruction for the alternative video description text, it further includes: Determine the target camera movement method corresponding to the target video description text; Correspondingly, the performing video generation according to the target object image and the target video description text includes: Perform video generation according to the target object image, the target video description text, and the target camera movement method.

12. The method according to claim 10, wherein The selection instruction for the alternative video description text is a trigger instruction for the alternative video description text displayed in the interface; correspondingly, determining the target camera movement mode corresponding to the target video description text includes: Determining the target camera movement mode corresponding to the target video description text according to the position of the target video description text in the interface.

13. A video generation device, characterized in that, Including: An acquisition unit for acquiring an image of a target object; A video description text determination unit for acquiring at least two alternative video description texts according to the target object image and displaying them, where any video description text is used to describe at least one of the attribute information, the environment where the target object is located, and the actions of the target object; A video description text selection unit for determining a target video description text from the at least two alternative video description texts in response to a selection instruction for the alternative video description text; A video generation unit for generating a video according to the target object image and the target video description text.

14. An electronic device, characterized in that, Including: At least one processor and a memory; The memory stores computer execution instructions; The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the method according to any one of claims 1-12.

15. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the processor executes the computer execution instructions, the method according to any one of claims 1-12 is implemented.

16. A computer program product, characterized in that, Including computer execution instructions, and when the processor executes the computer execution instructions, the method according to any one of claims 1-12 is implemented.