Video generation method and device, electronic equipment, storage medium and program product
By selecting image materials from a cloud-stored image library and generating stylized animated videos, the problem of low efficiency in manual operation by users in existing technologies is solved, achieving efficient and fun stylized animated video generation and improving user experience.
Patent Information
- Application Number
- CN202511284894.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-25
AI Technical Summary
In existing technologies, generating stylized animated videos requires users to manually upload images and perform multiple steps, which is inefficient and cannot automatically generate stylized animated videos.
The process involves selecting image materials from a cloud-stored image library and generating stylized animation videos based on stylized animation templates. This includes determining the stylized animation template, selecting image materials based on the image identification information in the image library, responding to user operations to determine the target image, and generating the stylized animation video.
It simplifies user operations, improves the efficiency and fun of generating stylized animated videos, lowers the barrier to entry, and enhances the user experience and the sense of surprise in the interaction process.
Smart Images

Figure CN121012973A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to the fields of cloud storage, cloud computing, and cloud services, and particularly to a video generation method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] Existing technologies for generating stylized animated videos from images require users to first upload images, and then convert the uploaded images into stylized static images based on user-input prompts and operations. Each step depends on user operations, resulting in low efficiency and the inability to generate corresponding stylized animated videos for users. Summary of the Invention
[0003] This disclosure presents a video generation method, apparatus, electronic device, storage medium, and program product.
[0004] According to a first aspect of this disclosure, a video generation method is provided, comprising: determining a stylized animation template in response to an operation on a stylized animation template in a first page, wherein the first page is a page displaying a preset video template in an image library stored in a cloud; selecting at least one image from the image library as image material based on the image identification information in the image library, and displaying the image material in a second page, wherein the second page is a subpage of the first page, and the image material contains only a target object; determining a target image from the image material in response to an operation on the image material in the second page; and generating a stylized animation video of the target object based on the target image and the stylized animation template.
[0005] According to a second aspect of this disclosure, a video generation apparatus is provided, comprising: a template determining module configured to determine a stylized animation template in response to an operation on a stylized animation template in a first page, wherein the first page is a page displaying a preset video template in an image library stored in the cloud; a material determining module configured to select at least one image from the image library as image material based on the identification information of the images in the image library, and display the image material in a second page, wherein the second page is a subpage of the first page, and the image material contains only a target object; an image determining module configured to determine a target image from the image material in response to an operation on the image material in the second page; and a video generation module configured to generate a stylized animation video of the target object based on the target image and the stylized animation template.
[0006] According to a third aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first aspect.
[0007] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform a method as described in any implementation of the first aspect.
[0008] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as described in any implementation of the first aspect.
[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is an exemplary system architecture diagram to which this disclosure can be applied; Figure 2 This is a flowchart of an embodiment of the video generation method according to the present disclosure; Figure 3 This is a flowchart of another embodiment of the video generation method according to the present disclosure; Figure 4 This is a flowchart of yet another embodiment of the video generation method according to the present disclosure; Figure 5 This is a schematic diagram of the second page provided in an embodiment of this disclosure; Figure 6 This is a flowchart of yet another embodiment of the video generation method according to the present disclosure; Figure 7 This is a schematic diagram of the structure of an embodiment of the video generation apparatus according to the present disclosure; Figure 8 This is a block diagram of an electronic device used to implement the video generation method of the embodiments of this disclosure. Detailed Implementation
[0011] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0012] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0013] Figure 1 An exemplary system frame 100 is shown, to which embodiments of the video generation method or video generation apparatus of this disclosure may be applied.
[0014] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0015] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include cloud storage applications and instant messaging applications.
[0016] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.
[0017] Server 105 can provide various services through its built-in applications, taking cloud storage applications as an example. Users can operate on the cloud storage page through the cloud storage applications on terminal devices 101, 102, and 103, and send operation requests to server 105. Server 105 can receive operation requests and run the cloud storage applications for processing, performing the following steps: In response to an operation on a stylized animation template on the first page, determine the stylized animation template, where the first page is a page displaying preset video templates in the cloud storage's image library; select at least one image from the image library as image material based on the image identification information in the image library, and display the image material on a second page, where the second page is a subpage of the first page, and the image material only contains the target object; In response to an operation on the image material on the second page, determine the target image from the image material; Based on the target image and the stylized animation template, generate a stylized animation video of the target object. Terminal devices 101, 102, and 103 can receive the animation video and run the cloud storage applications to display the animation video to the user.
[0018] It should be noted that the video generation method provided in this embodiment is generally executed by server 105, and correspondingly, the video generation device is generally located in server 105.
[0019] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0020] Continue to refer to Figure 2 The diagram illustrates a flow 200 of an embodiment of a video generation method according to the present disclosure. The video generation method includes the following steps: Step 201: In response to the operation on the stylized animation template in the first page, determine the stylized animation template.
[0021] In this embodiment, the execution entity of the video generation method (e.g. Figure 1 The server 105 shown will determine the stylized animation template when determining the operation for the stylized animation template on the first page, where the first page is a page displaying a preset video template in the cloud storage image library. Users can, for example... Figure 1The cloud storage application on terminal devices 101, 102, or 103 can manipulate the image library icon on the cloud storage page. The executing entity can respond to the manipulation of the image library icon and open the image library page. The user can then continue to manipulate the creative function icons on the image library page using the cloud storage application on terminal devices 101, 102, or 103. The executing entity can respond to the manipulation of the creative function icons and open the first page, which displays various preset video templates. The user can then continue to manipulate the stylized animation templates on the first page using the cloud storage application on terminal devices 101, 102, or 103. The executing entity can respond to the manipulation of the stylized animation templates and confirm that a stylized animation template has been selected.
[0022] In this embodiment, cloud storage is an online storage system based on Internet technology that stores data on cloud servers and provides services such as cross-device access, management, and sharing. Cloud storage can provide large-capacity storage space for storing various types of files, such as documents, images, and videos. The image library is a module within the cloud storage specifically designed for the storage and intelligent management of multimedia files such as images and videos. For example, the image library can be a photo album within the cloud storage, providing functions such as automatic backup, AI (Artificial Intelligence) classification, and duplicate detection. Various types of video templates can be preset in the creative community of the image library, allowing users to create various types of videos from images stored in the image library using these templates. For example, preset video templates may include old photo revival templates, growth trajectory templates, stylized animation templates, etc., and the embodiments disclosed herein are not limited to these. The stylized animation template is a template used to convert images into corresponding stylized animated videos. For example, the stylization could be a blind box style, where a blind box refers to a toy box whose specific product style cannot be known in advance and has a random attribute. Accordingly, the blind box style can include features such as toy boxes. The stylization in the stylized animation template can also refer to other forms of style, and the embodiments of this disclosure do not limit this. The user's operation through the cloud storage application on the terminal device can be a click operation or a voice operation, and the embodiments of this disclosure do not limit this.
[0023] Step 202: Select at least one image from the image library as image material based on the image identification information in the image library, and display the image material on the second page.
[0024] Building upon step 201, this step aims to have the executing entity retrieve at least one image from the image library as image material based on the image identification information in the image library, and display the image material on the second page. The second page is a subpage of the first page, and the image material contains only the target object. After determining the stylized animation template, the executing entity can statistically analyze the identification information of the images stored in the image library, and then select suitable image material from the images stored in the image library based on this identification information. The selected image material is then displayed on the second page, recommending the selected image material to the user. Specifically, the images in the image library will have identification information indicating the identity of the object in the image. For example, the identification information could be "myself," "dad," "mom," "baby," etc. Suppose the user wants to generate a blind box-style personal avatar animation video, then when selecting image material, the image with the identification information "myself" will be selected as the image material. "Containing only the target object" here means that for the same type of object, the image material contains only one object of that type. For example, when the image is a person image, the selected image material contains only one person; when the image is an animal image, the selected image material contains only one animal.
[0025] The second page can be the page opened after the user interacts with the stylized animation template on the first page, and the executing entity determines that the stylized animation template has been selected. Image materials can be static images of target objects in the scene, such as people or animals; for example, image materials can be photos stored in a cloud storage album. The executing entity selects a static image from the image library as image material based on the image identification information, or it can select multiple static images from the image library as image materials; the embodiments disclosed herein do not limit this selection.
[0026] Step 203: In response to the operation on the image material in the second page, determine the target image from the image material.
[0027] Building upon step 202, this step aims to have the aforementioned executing entity determine the target image from the image assets in response to an operation on the image assets displayed on the second page. After acquiring and displaying the image assets, the user can, for example... Figure 1 The cloud storage application on terminal devices 101, 102, or 103 shown operates on the image materials on the second page. The executing entity can respond to the operation on the target image in the image materials and determine that the target image is selected. When the image material is a single static image, the target image is that static image. When the image material consists of multiple static images, the target image can be any one of the multiple static images.
[0028] Step 204: Generate a stylized animated video of the target object based on the target image and the stylized animation template.
[0029] Building upon step 203, this step aims to generate a stylized animated video of the target object from the target image and the stylized animation template by the aforementioned executing entity. After determining the target image, the executing entity can process the target image using the stylized animation template to generate an animated video of the target object corresponding to the template. Optionally, the executing entity can process the target image using a graph-based video model corresponding to the stylized animation template to generate an animated video of the target object corresponding to the template; the executing entity can also process the target image using other graph-based video algorithms corresponding to the stylized animation template to generate an animated video of the target object corresponding to the template. The embodiments of this disclosure do not limit the implementation method of generating a stylized animated video based on the target image and the stylized animation template.
[0030] The video generation method provided by this disclosure firstly determines a stylized animation template in response to an operation on a first page, wherein the first page is a page in a cloud-stored image library that displays preset video templates; then, at least one image is selected from the image library as image material based on the image identification information in the image library, and the image material is displayed on a second page, wherein the second page is a subpage of the first page, and the image material contains only the target object; subsequently, in response to an operation on the image material on the second page, a target image is determined from the image material; finally, a stylized animation video of the target object is generated based on the target image and the stylized animation template. This method can intelligently select images from cloud-stored image libraries and make personalized recommendations to users, simplifying the process of selecting and acquiring images. Furthermore, it can directly generate stylized animated videos from original images using stylized animation templates, thus simplifying the generation process and user operations from original images to stylized animated videos. This avoids cumbersome user procedures, making the user experience simple and efficient, lowering the barrier to entry, and providing a fun and engaging experience. The generated stylized animated videos also enhance the sense of surprise during user interaction, thereby increasing user engagement and satisfaction, and improving the overall user experience.
[0031] Furthermore, the acquisition, storage, use, processing, transportation, provision, and disclosure of user personal information (such as images in the image library involved in this disclosure) in the technical solutions disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0032] Continue to refer to Figure 3 , Figure 3 A flow 300 of another embodiment of the video generation method according to the present disclosure is shown. The video generation method includes the following steps: Step 301: In response to the operation on the stylized animation template in the first page, determine the stylized animation template.
[0033] In this embodiment, the execution entity of the video generation method (e.g. Figure 1 The server 105 shown will determine the stylized animation template in response to the operation on the stylized animation template in the first page. Step 301 is basically the same as step 201 in the previous embodiment. For the specific implementation, please refer to the previous description of step 201, which will not be repeated here.
[0034] Step 302: Select multiple images from the image library as candidate images based on the image identification information in the image library.
[0035] In this embodiment, the executing entity statistically analyzes the identification information of images stored in the image library, and then selects suitable images as candidate images from the images stored in the image library based on this identification information. Specifically, the images in the image library will have identification information to indicate the identity of the object in the image. For example, the identification information can be "myself," "dad," "mom," "baby," etc. Suppose the user wants to generate a blind box-style personal avatar animation video, then all images with the identification information "myself" will be selected as candidate images.
[0036] In some optional implementations of this embodiment, step 302 includes: statistically analyzing the identifiers of images in the image library, determining target identifier information from the statistically obtained identifier information, and using multiple images corresponding to the target identifier information as candidate images.
[0037] In this implementation, the execution entity statistically analyzes the identifiers of each image in the image library. Since the image identifier represents the identity of the object in the image, the identifier can be "myself," "father," "mother," etc. Then, the execution entity determines the target identifier from all the statistically obtained identifiers. The default target identifier is "myself," but it can also be determined based on the number of images corresponding to each identifier; this embodiment does not impose a specific limitation. Finally, the execution entity selects multiple images corresponding to the target identifier as candidate images. This allows for intelligent image selection based on the image identifiers and recommendations to the user.
[0038] In some optional implementations of this embodiment, determining the target identifier information from the statistically obtained identifier information includes: determining the number of images corresponding to the identifier information for the statistically obtained identifier information; and selecting the identifier information with the most images as the target identifier information.
[0039] In this implementation, for all the statistically obtained identification information, the aforementioned execution entity counts the number of images in the image library corresponding to each identification information, and then selects the identification information with the largest number of corresponding images as the target identification information. Thus, by determining the target identification information based on the number of images, images that better meet the user's needs can be recommended.
[0040] Step 303: Select at least one image from the candidate images as image material based on the image features of the candidate images and the preset feature threshold.
[0041] In this embodiment, the execution entity selects at least one image from the candidate images as image material based on the image features of the candidate images and a preset feature threshold. Here, the image features of the candidate images are determined first.
[0042] In some optional implementations of this embodiment, the image features include at least one of the following: image time, image type, image resolution, image size, and image aesthetics. The image time, image type, image resolution, and image size can be directly obtained from image information stored in an image library. The image aesthetics can be determined by evaluating the image using a preset algorithm model. For example, the image aesthetics can be scored using a preset algorithm model, and the aesthetics of the image can be evaluated based on the aesthetics score. Thus, by filtering candidate images using the above-mentioned image features, image materials are obtained, thereby ensuring the quality of the selected images, which in turn ensures the effect of the subsequently generated animation video, and makes the selected images more in line with the user's requirements.
[0043] Furthermore, the aforementioned execution entity pre-sets feature thresholds for each feature in the image features. The execution entity can compare the feature information of each feature with the corresponding feature threshold to filter candidate images and obtain image materials. Specifically, filtering images stored in the image library based on image feature information and preset feature thresholds can result in one or multiple images. When multiple images are selected, all of them can be used as image materials and displayed on the second page, or only a portion of the selected images can be used as image materials and displayed on the second page. For example, the number of image materials can be preset, and a preset number of images can be selected from the filtered images as image materials for display on the second page. The method for selecting the preset number of images from the filtered images can be random selection or sorting the images according to their feature information, selecting the preset number of images at the top of the sort order as image materials. This disclosure does not limit this approach. Therefore, selecting high-quality images for personalized recommendations to users ensures the quality of the selected images, thereby guaranteeing the effect of the subsequently generated animation video and making the selected images more in line with user requirements.
[0044] In some optional implementations of this embodiment, step 303 includes: filtering candidate images step by step according to the preset priority of each feature in the image features based on image features and preset feature thresholds; sorting the images obtained by step-by-step filtering based on the aesthetic information in the image features, and selecting a preset number of images at the top of the sort as image materials.
[0045] In this implementation, the execution entity pre-sets priorities and thresholds for each feature in the image features and user features. The execution entity can then compare the corresponding feature information with the corresponding feature thresholds step by step according to the priority of each feature, thereby filtering the image step by step.
[0046] After filtering the images stored in the image library step by step, the executing entity can sort the filtered images according to the aesthetic information in the feature information, and select a preset number of images at the top of the sort as image materials.
[0047] For example, the priority of image features is: image time > image type > image resolution > image size > image aesthetics. Based on the priority of image features, the most recently uploaded images from the image library are selected first, followed by images containing only a single person (the target object). Then, low-quality images such as those with poor clarity, excessively small size, and blurry images of people are filtered out. In an optional example, the filtered image categories are: 22, 23, 5, 799, 25, 24, 30, 600, 601, 602, 603, 604, 605, 606, 607, 610, 611, 700, 701, 702, 703, 1300, 1301, 1304, 1312, 1030, 1031, 1032, and 10001—image types that do not depict people. Here, number 22 is a screenshot, 23 is a cartoon image, and so on. Long images (width / height > 2 or height / width > 2) are filtered out. Images with a resolution < 6000 are filtered out. Images with an aesthetic score < 5600 are filtered out. This ensures the quality of the selected images, thereby guaranteeing the quality of the subsequently generated animated video.
[0048] Step 304: In response to the operation on the image material in the second page, determine the target image from the image material.
[0049] Step 304 is basically the same as step 203 in the aforementioned embodiment. For the specific implementation method, please refer to the aforementioned description of step 203, which will not be repeated here.
[0050] Step 305: In response to the operation of the video generation control on the second page, input the target image into the image-generated video model corresponding to the stylized animation template, and output the stylized animation video of the target object.
[0051] In this embodiment, the execution entity, in response to an operation on the video generation control on the second page, inputs the target image into the image-generated video model corresponding to the stylized animation template and outputs a stylized animation video of the target object. The image-generated video model is trained based on sample images and their corresponding stylized animation videos. The second page also provides a video generation control. Users can operate the video generation control on the second page through cloud storage applications on their terminal devices. The execution entity, in response to the operation of the video generation control, inputs the target image into the image-generated video model corresponding to the stylized animation template and outputs an animation video of the target object corresponding to the template. By setting a video generation control on the second page and utilizing user operations on the control to generate stylized animation videos, one-click generation of stylized video animations can be achieved, ensuring the stability of the generated stylized animation videos.
[0052] from Figure 3 It can be seen from this that, with Figure 2 Compared to the corresponding embodiments, the video generation method in this embodiment emphasizes the steps of selecting image materials and generating stylized animated videos of target objects. By filtering images stored in the image library based on image identification information, images that better meet user needs can be recommended to users in a personalized way, thereby bringing users a convenient and efficient user experience.
[0053] Continue to refer to Figure 4 , Figure 4 A flow 400 of another embodiment of the video generation method according to this disclosure is shown. The video generation method includes the following steps: Step 401: In response to the operation on the stylized animation template in the first page, determine the stylized animation template.
[0054] Step 402: Select at least one image from the image library as image material based on the image identification information in the image library, and display the image material on the second page.
[0055] Steps 401-402 are basically the same as steps 201-202 in the aforementioned embodiments. For specific implementation methods, please refer to the aforementioned description of steps 201-202, which will not be repeated here.
[0056] Step 403: In response to the operation on the image loading bit in the second page, display the image stored in the image library in the third page.
[0057] In this embodiment, the execution entity of the video generation method (e.g. Figure 1 The server 105 shown responds to an operation on the image loading position of the second page by displaying images stored in the image library on a third page, where the third page is a subpage of the second page. An image loading position is also provided on the second page. When the image materials displayed on the second page do not meet the user's requirements, the user can operate on the image loading position on the second page through a cloud storage application on the terminal device. The executing entity can respond to the operation on the image loading position by opening the third page and displaying images stored in the image library on the third page. The user can select the desired image from the images provided on the third page. Figure 5 As shown, seven images 511 are displayed as image materials on the second page 510, and the image loading position 512 follows the seven images as image materials on the second page 510.
[0058] Step 404: In response to the operation on the image displayed on the third page, determine the target image and display the target image in the image loading bit.
[0059] In this embodiment, the execution entity, in response to an operation on the image displayed on the third page, determines the target image and displays the target image in the image loading position. After the image stored in the image library is displayed on the third page, the user can operate on the image on the third page through a cloud storage application on the terminal device. The execution entity can, in response to the operation on the image, determine that the image is selected, and the selected image is the target image, and display the target image in the image loading position on the second page.
[0060] Step 405: In response to the operation of the video generation control on the second page, input the target image into the image-generated video model corresponding to the stylized animation template, and output the stylized animation video of the target object.
[0061] In this embodiment, the execution entity, in response to an operation on the video generation control on the second page, inputs the target image into the image-generated video model corresponding to the stylized animation template, and outputs a stylized animation video of the target object. The image-generated video model is trained based on sample images and their corresponding stylized animation videos. The second page also provides a video generation control. Users can operate the video generation control on the second page through cloud storage applications on their terminal devices. The execution entity, in response to the operation of the video generation control, inputs the target image into the image-generated video model corresponding to the stylized animation template, and outputs an animation video of the target object corresponding to the template. Figure 5 As shown, below the display positions of image material 511 and image loading position 512 on the second page 510, there is a video generation control 513.
[0062] from Figure 4 It can be seen from this that, with Figure 2 Compared to the corresponding embodiments, the video generation method in this embodiment emphasizes the step of determining the target image through the image loading bit. By setting the image loading bit in the second page, the third page that displays images stored in the image library is opened through the image loading bit. This method can provide users with more choices when the image materials provided on the second page do not meet the user's requirements, and can improve the flexibility of using stylized animation templates.
[0063] Continue to refer to Figure 6 , Figure 6 A flow 600 of another embodiment of the video generation method according to the present disclosure is shown. The video generation method includes the following steps: Step 601: In response to the operation on the stylized animation template in the first page, determine the stylized animation template.
[0064] Step 602: Select at least one image from the image library as image material based on the image identification information in the image library, and display the image material on the second page.
[0065] Step 603: In response to an operation on the image assets in the second page, determine the target image from the image assets.
[0066] Step 604: Generate a stylized animated video of the target object based on the target image and the stylized animation template.
[0067] Steps 601-604 are basically the same as steps 201-204 in the aforementioned embodiments. For specific implementation methods, please refer to the aforementioned description of steps 201-204, which will not be repeated here.
[0068] Step 605: Determine the key action frame positions of the target object in the stylized animation video.
[0069] In this embodiment, the execution entity of the video generation method (e.g. Figure 1 The server 105 shown determines the key action frame positions of the target object in the stylized animation video. When generating the animation video, the executing entity can mark the positions of the key action frames of the target object in the animation video. After generating the stylized animation video of the target object, the executing entity can determine the positions of the key action frames of the target object in the stylized animation video based on the marked positions.
[0070] Step 606: Based on the key action frame positions, extract image frames from the stylized animation video to obtain a stylized image of the target object.
[0071] In this embodiment, the execution entity extracts image frames from the stylized animation video based on the key action frame positions to obtain a stylized image of the target object. After determining the key action frame positions of the target object in the stylized animation video, the execution entity can extract image frames from the stylized animation video according to the key action frame positions, and use the extracted image frames as the stylized image of the target object. The execution entity can extract image frames from all key actions in the stylized animation video according to the key action frame positions to obtain a stylized image of the target object, or it can extract only some of the key action image frames to obtain a stylized image of the target object; the embodiments of this disclosure do not limit this approach.
[0072] In some optional implementations of this embodiment, the stylized animated video is an animated video with a toy box style; and in the stylized image, a 3D animated image of the target object is displayed next to the toy box. Here, the toy box style refers to the blind box style. A blind box is a toy box where the specific product style cannot be known in advance, possessing a random attribute. Accordingly, the blind box style can include features such as a toy box. This combines the fun concept of "drawing blind boxes" with the user's historical image recall, further enhancing the novelty and surprise of the user experience.
[0073] from Figure 6 It can be seen from this that, with Figure 2 Compared to the corresponding embodiments, the video generation method in this embodiment emphasizes the step of determining the stylized image. After the stylized animation video of the target object is generated, the positions of key action frames in the animation video are determined, and image frames are extracted from the animation video based on the positions of the key action frames to obtain the stylized image of the target object. This method can expand the application scope, meet the diverse needs of users, and provide users with richer and more diverse usage options.
[0074] Further reference Figure 7 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a video generation apparatus, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0075] like Figure 7 As shown, the video generation device 700 of this embodiment includes: a template determination module 701, a material determination module 702, an image determination module 703, and a video generation module 704. The template determination module 701 is configured to determine a stylized animation template in response to an operation on a stylized animation template in a first page, wherein the first page is a page in a cloud-stored image library displaying preset video templates; the material determination module 702 is configured to select at least one image from the image library as image material based on the image identification information in the image library, and display the image material on a second page, wherein the second page is a subpage of the first page, and the image material only contains the target object; the image determination module 703 is configured to determine a target image from the image material in response to an operation on the image material in the second page; and the video generation module 704 is configured to generate a stylized animation video of the target object based on the target image and the stylized animation template.
[0076] In this embodiment, the specific processing of the template determination module 701, the material determination module 702, the image determination module 703, and the video generation module 704 in the video generation device 700, and the resulting technical effects, can be referred to respectively. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.
[0077] In some optional implementations of this embodiment, the material determination module 702 includes: a candidate determination submodule, configured to select multiple images from the image library as candidate images based on the image identification information in the image library; and a material determination submodule, configured to select at least one image from the candidate images as image material based on the image features of the candidate images and a preset feature threshold.
[0078] In some optional implementations of this embodiment, the candidate determination submodule includes: a statistics unit configured to statistically analyze the identifiers of images in the image library and determine target identifier information from the statistically obtained identifier information; and a determination unit configured to use multiple images corresponding to the target identifier information as candidate images.
[0079] In some optional implementations of this embodiment, the statistical unit is further configured to: determine the number of images corresponding to the statistically obtained identification information; and select the identification information with the most images as the target identification information.
[0080] In some optional implementations of this embodiment, the image features include at least one of the following: image time, image type, image resolution, image size, and image aesthetics.
[0081] In some optional implementations of this embodiment, the material determination submodule is further configured to: filter candidate images step by step according to the preset priority of each feature in the image features based on image features and preset feature thresholds; sort the images obtained by step-by-step filtering based on the aesthetic information in the image features, and select a preset number of images at the top of the sort as image materials.
[0082] In some optional implementations of this embodiment, the video generation module 704 is further configured to: in response to an operation on the video generation control on the second page, input the target image into the image-generated video model corresponding to the stylized animation template, and output the stylized animation video of the target object.
[0083] In some optional implementations of this embodiment, the video generation apparatus 700 further includes: a display module configured to display an image stored in an image library on a third page in response to an operation on an image loading bit in a second page, wherein the third page is a subpage of the second page; and a loading module configured to determine the target image in response to an operation on the image displayed on the third page, and display the target image on the image loading bit.
[0084] In some optional implementations of this embodiment, the video generation apparatus 700 further includes: a position determination module configured to determine the position of key action frames of a target object in a stylized animation video; and an extraction module configured to extract image frames from the stylized animation video based on the key action frame positions to obtain a stylized image of the target object.
[0085] In some optional implementations of this embodiment, the stylized animated video is an animated video with a toy box style; and in the stylized image, a three-dimensional animated image of the target object is displayed next to the toy box.
[0086] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0087] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0088] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0089] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0090] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0091] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as video generation methods. For example, in some embodiments, the video generation method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the video generation method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the video generation method by any other suitable means (e.g., by means of firmware).
[0092] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0093] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0094] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0095] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0096] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0097] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0098] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0099] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A video generation method, comprising: In response to an operation on a stylized animation template in the first page, the stylized animation template is determined, wherein the first page is a page that displays a preset video template in an image library stored in the cloud; At least one image is selected from the image library as image material based on the image identification information in the image library, and the image material is displayed on the second page, wherein the second page is a subpage of the first page, and the image material contains only the target object; In response to an operation on image assets in the second page, a target image is determined from the image assets; Based on the target image and the stylized animation template, a stylized animation video of the target object is generated.
2. The method according to claim 1, wherein, The step of selecting at least one image from the image library as image material based on the image identification information in the image library includes: Multiple images are selected as candidate images from the image library based on the image identification information in the image library; Based on the image features of the candidate images and a preset feature threshold, at least one image is selected from the candidate images as the image material.
3. The method according to claim 2, wherein, The step of selecting multiple images as candidate images from the image library based on the image identification information in the image library includes: The identifiers of images in the image library are statistically analyzed, and the target identifier information is determined from the statistically obtained identifier information; Multiple images corresponding to the target identification information are used as candidate images.
4. The method according to claim 3, wherein, Determining the target identifier information from the statistically obtained identifier information includes: For the statistically obtained identification information, determine the number of images corresponding to the identification information; The identifier information with the largest number of images is selected as the target identifier information.
5. The method according to claim 2, wherein, The image features include at least one of the following: image time, image type, image resolution, image size, and image aesthetics.
6. The method according to claim 5, wherein, The step of selecting at least one image from the candidate images as the image material based on the image features of the candidate images and a preset feature threshold includes: Based on the image features and preset feature thresholds, the candidate images are filtered step by step according to the preset priority of each feature in the image features; Based on the aesthetic information in the image features, the images obtained by filtering step by step are sorted, and a preset number of images at the top of the sort are selected as the image materials.
7. The method according to claim 1, wherein, The step of generating a stylized animation video of the target object based on the target image and the stylized animation template includes: In response to an operation on the video generation control in the second page, the target image is input into the image-generated video model corresponding to the stylized animation template, and the stylized animation video of the target object is output.
8. The method according to claim 7, further comprising: In response to an operation on the image loading bit in the second page, an image stored in the image library is displayed on a third page, wherein the third page is a subpage of the second page; In response to an operation on the image displayed on the third page, the target image is determined and displayed at the image loading position.
9. The method according to any one of claims 1 to 8, further comprising: Determine the key action frame positions of the target object in the stylized animation video; Based on the key action frame positions, image frames are extracted from the stylized animation video to obtain the stylized image of the target object.
10. The method according to claim 9, wherein, The stylized animated video is an animated video with a toy box style; and In the stylized image, a three-dimensional animated image of the target object is displayed next to the toy box.
11. A video generation apparatus, comprising: The template determination module is configured to determine the stylized animation template in response to an operation on the stylized animation template in the first page, wherein the first page is a page that displays a preset video template in an image library stored in the cloud; The material determination module is configured to select at least one image from the image library as image material based on the image identification information in the image library, and display the image material on a second page, wherein the second page is a subpage of the first page, and the image material contains only the target object; An image determination module is configured to determine a target image from the image material in response to an operation on the image material in the second page; The video generation module is configured to generate a stylized animated video of the target object based on the target image and the stylized animation template.
12. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of any one of claims 1-10.
13. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method of any one of claims 1-10.
14. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.