Video generation method and apparatus, and device and medium
By using interface switching and video generation technologies, the limitations of existing technologies in character motion transfer have been solved, achieving efficient video generation effects and improved user experience.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-12-15
- Publication Date
- 2026-07-23
AI Technical Summary
Existing technologies for character motion transfer in animation, film production, virtual reality, or augmented reality applications require specialized equipment and software, making them difficult to promote and limiting their application.
By providing preset interfaces in the interface, switching interfaces to determine the original image and reference video, generating the target video based on the reference video, and making the objects in the target video move in a certain way.
It enables efficient motion transfer of preset category objects, improving video generation quality and user experience.
Smart Images

Figure CN2025142379_23072026_PF_FP_ABST
Abstract
Description
Methods, apparatus, equipment and media for video generation
[0001] Cross-references to related applications
[0002] This application claims the benefit of Chinese patent application No. 202510089889.3, filed on January 20, 2025, which is incorporated herein by reference. Technical Field
[0003] The embodiments disclosed herein relate to the field of media processing technology, and more particularly to a method, apparatus, device, and medium for generating video. Background Technology
[0004] With the continuous development and improvement of network technology and digital media technology, digital media such as images and videos are increasingly used in people's lives, providing entertainment services. In addition, digital media is also widely used in advertising, film and television, and new media art, bringing numerous conveniences to the development of art. Currently, in animation, film production, virtual reality (VR), or augmented reality (AR) applications, it is often necessary to analyze the movements of characters in videos or images and apply them to another object or environment to achieve the transfer of character movements. Summary of the Invention
[0005] Embodiments of this disclosure describe a method, apparatus, device, and medium for generating video.
[0006] According to a first aspect, a method for generating a video is provided, the method comprising:
[0007] A preset interface is provided in the first interface;
[0008] After detecting a first trigger operation for the preset interface, the first interface is switched to the second interface;
[0009] The second interface determines the original image and the reference video; the original image includes a first object of a preset category; the reference video includes a second moving object of the preset category;
[0010] In response to a second trigger operation on the second interface, a target video corresponding to the original image is generated based on the reference video, so that the first object in the target video moves in the manner of the second object.
[0011] According to a second aspect, a video generation apparatus is provided, the apparatus comprising:
[0012] The unit is configured to provide a preset interface in the first interface;
[0013] The switching unit is configured to switch the first interface to the second interface after detecting a first trigger operation for the preset interface;
[0014] The first determining unit is configured to determine an original image and a reference video through the second interface; the original image includes a first object of a preset category; the reference video includes a second moving object of the preset category;
[0015] The generation unit is configured to, in response to a second trigger operation on the second interface, generate a target video corresponding to the original image based on the reference video, such that the first object in the target video moves in the manner of the second object.
[0016] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform any of the methods described in the first aspect.
[0017] According to a fourth aspect, an electronic device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in any one of the first aspects.
[0018] According to a fifth aspect, a computer program product is provided, wherein the computer program product includes a computer program, which, when executed by a computer device, causes the computer device to perform any of the methods described in the first aspect.
[0019] According to a sixth aspect, a computer program is provided, wherein when the computer program is executed by a computer device, the computer device performs any of the methods described in the first aspect.
[0020] According to an embodiment of this disclosure, a video generation scheme is provided. A preset interface is provided in a first interface. Upon detecting a first trigger operation on the preset interface, the first interface is switched to a second interface. An original image and a reference video are determined through the second interface. The original image includes a first object of a preset category, and the reference video includes a second object of a preset category undergoing motion. In response to a second trigger operation on the second interface, a target video corresponding to the original image is generated based on the reference video, such that the first object in the target video moves according to the motion pattern of the second object. Attached Figure Description
[0021] Figure 1 is a schematic diagram of a video generation scenario according to an exemplary embodiment of the present disclosure;
[0022] Figure 2 is a schematic diagram of an exemplary system architecture applying an embodiment of the present disclosure;
[0023] Figure 3 is a flowchart illustrating a video generation method according to an exemplary embodiment of this disclosure;
[0024] Figure 4a is a schematic diagram of a video generation scenario according to an exemplary embodiment of the present disclosure;
[0025] Figure 4b is a schematic diagram of another video generation scenario according to an exemplary embodiment of the present disclosure;
[0026] Figure 4c is a schematic diagram of another video generation scenario according to an exemplary embodiment of the present disclosure;
[0027] Figure 4d is a schematic diagram of another video generation scenario according to an exemplary embodiment of the present disclosure;
[0028] Figure 4e is a schematic diagram of another video generation scenario according to an exemplary embodiment of the present disclosure;
[0029] Figure 4f is a schematic diagram of another video generation scenario according to an exemplary embodiment of the present disclosure;
[0030] Figure 5 is a block diagram of a video generation apparatus according to an exemplary embodiment of the present disclosure;
[0031] Figure 6 is a schematic block diagram of an electronic device provided in some embodiments of this disclosure. Detailed Implementation
[0032] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0033] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations of the technical solutions disclosed herein, based on the prompt message.
[0034] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0035] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0036] The technical solutions provided in this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the relevant invention and not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0037] With the continuous development and improvement of network technology and digital media technology, digital media such as images and videos are increasingly used in people's lives, providing entertainment services. Furthermore, digital media is also widely used in advertising, film and television, and new media art, bringing numerous conveniences to the development of art. Currently, in animation, film production, virtual reality (VR), or augmented reality (AR) applications, it is often necessary to analyze the movements of characters in videos or images and apply them to another object or environment to achieve the transfer of character movements. In related technologies, various specialized equipment and software are generally required, operated by professionals, to complete the transfer of character movements. Therefore, this method has significant limitations and is difficult to promote widely. Currently, a video generation solution is needed.
[0038] This disclosure provides a video generation scheme. By providing a preset interface in a first interface, upon detecting a first trigger operation on the preset interface, the first interface is switched to a second interface. The second interface determines an original image and a reference video. The original image includes a first object of a preset category, and the reference video includes a second object of a preset category undergoing motion. In response to a second trigger operation on the second interface, a target video corresponding to the original image is generated based on the reference video, causing the first object in the target video to move according to the motion pattern of the second object. This achieves the transfer of the motion pattern of the second object in the reference video to the first object in the original image, resulting in a target video corresponding to the original image. This efficiently completes the motion transfer of preset category objects, improves video generation quality, and enhances the user experience.
[0039] Referring to Figure 1, it is a schematic diagram of a video generation scenario according to an exemplary embodiment.
[0040] As shown in Figure 1, for example, image A includes person a, and video B includes person b in motion. The user wants to generate video C based on image A and video B, such that video C includes person a moving in the same manner as person b, and the background in video C is the same as the background in image A.
[0041] Specifically, firstly, the user can input image A into the media processing client via a terminal device and select video B as a reference video through the interface provided by the media processing client. The media processing client can utilize a pre-trained network model to extract the character features Z1 corresponding to person a in image A, the background features Z2 corresponding to the background, and the pre-stored keypoint features Z3 corresponding to person b in video B. Since video B consists of multiple video frames, the keypoint features Z3 can include the keypoint features corresponding to person b in each frame of video B. For example, if video B consists of m video frames, the keypoint features Z3 can include the keypoint features of person b in each of the m video frames. Additionally, the media processing client can also obtain a noisy image D based on random noise.
[0042] Next, the media processing client can input the extracted character features Z1, background features Z2, character keypoint features Z3, and noisy image D into the diffusion model M. The diffusion model M then performs multiple iterations of denoising based on the character features Z1, background features Z2, and character keypoint features Z3, resulting in an image sequence L. Image sequence L includes multiple image frames, each corresponding to a video frame in video B. For example, if video B includes m video frames B1, B2, B3…Bm, then an image sequence L consisting of m image frames can be generated. Image sequence L includes image frames L1, L2, L3…Lm, where image frame L1 is generated based on video frame B1 and corresponds to video frame B1, image frame L2 is generated based on video frame B2 and corresponds to video frame B2, and so on… image frame Lm is generated based on video frame Bm and corresponds to video frame Bm. In addition, the generated image frame includes person a, whose pose is similar to or the same as that of person b in the corresponding video frame, and the background, texture, etc. of the generated image frame are similar to or the same as those of image A.
[0043] Finally, video C can be generated based on image sequence L. Specifically, video B can include image components (i.e., multiple video frames) and audio components. The audio component of video B can be directly obtained and combined with image sequence L generated based on the image components of video B to obtain video C. That is, the image frames included in image sequence L are used as the image components of video C, and the audio component of video C is obtained based on the audio component of video B. The two components are then combined to obtain video C. Video C can include a moving person a, and the movement of person a is the same as the movement of person b in video B.
[0044] It should be noted that the embodiment in Figure 1 describes the example of a media processing client directly obtaining video C based on image A and video B. In other embodiments, the media processing client can also transmit information such as image A and video B to a media processing server deployed on the service platform via the network. The media processing server then generates video C based on image A and video B and transmits video C to the media processing client via the network to provide the video to the user. See Figure 2 for details.
[0045] Figure 2 is a schematic diagram of an exemplary system architecture applying an embodiment of this disclosure.
[0046] As shown in Figure 2, the system architecture 200 may include terminal devices 202, a network 203, and a server 204. It should be understood that the number or type of terminal devices, networks, and servers in Figure 2 is merely illustrative. Depending on implementation needs, any number or type of terminal devices, networks, and servers may be included.
[0047] Network 203 is a medium used to provide communication links between terminal devices and servers. Network 203 can include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0048] The terminal device 202 is equipped with a media processing client. The terminal device 202 can interact with the server via network 203 to receive or send requests or information. The terminal device 202 can be various electronic devices, including but not limited to smartphones, tablets, laptops, desktop computers, and smart wearable devices.
[0049] Server 204 is equipped with a media processing server. Server 204 can store, analyze, and process received data, and can also send control commands or requests to terminal devices or other servers. The server can provide media processing services in response to user service requests. It is understood that a single server can provide one or more services, and the same service can be provided by multiple servers.
[0050] Based on the system architecture shown in Figure 2, in this embodiment of the disclosure, user 201 can input an original image including a first object through terminal device 202, and select a reference video including a moving second object from a pool of candidate videos through terminal device 202. Next, terminal device 202 can transmit the original image and reference video information to server 204 via network 203. After receiving the original image and reference video information, server 204 can generate a target video corresponding to the original image based on the reference video, causing the first object in the target video to move in the manner of the second object. Finally, server 204 can return the target video to terminal device 202 via network 203, allowing user 201 to view and save the target video through terminal device 202.
[0051] The present disclosure will now be described in detail with reference to specific embodiments.
[0052] Figure 3 is a flowchart illustrating a video generation method according to an exemplary embodiment. This method can be applied to a media processing client or a media processing server. In this embodiment, the media processing client is installed on a terminal device, which may include, but is not limited to, mobile terminal devices such as smartphones, smart wearable devices, tablets, laptops, and desktop computers. The media processing server is deployed in a service platform, which can be any device, server, or device cluster with computing and processing capabilities. The method may include the following steps:
[0053] As shown in Figure 3, in step 301, a preset interface is provided in the first interface, and in step 302, after detecting the first trigger operation for the preset interface, the first interface is switched to the second interface.
[0054] In this embodiment, the first interface can be an image-related interface, such as an interface for displaying, editing, managing, or interacting with images. A preset interface within the first interface can be an interface for switching the first interface to a second interface. Users can trigger this preset interface to switch the first interface to the second interface, allowing them to perform operations for generating the target video through the second interface.
[0055] In one implementation, as shown in Figure 4a, the first interface 401 can be the image processing function entry interface in the media processing client, through which the user can select the desired image processing function. For example, the user can enter the image editing interface by triggering interface 402, or enter the interface for generating a video based on the uploaded image and the video selected by the user by triggering interface 403. Interface 403 can be a preset interface in the first interface 401 of this embodiment. It should be noted that in Figure 4a, the name of interface 403 is "Action Migration," which is only an example of a preset interface. In fact, the preset interface can also be set to other names, and this embodiment does not limit the specific name of the preset interface.
[0056] When the media processing client detects that the user has triggered the operation of interface 403, it can switch the first interface 401 to the second interface 404 used in this embodiment for generating video. The second interface 404 includes a first area 405, which may include an interface 406. The user can upload or update the original image by triggering the operation of interface 406. The second interface 404 may also include a second area 407, which may display cover thumbnails of multiple candidate videos. The user can select a reference video from the multiple candidate videos. For example, the user can select the candidate video as the reference video by clicking on the cover thumbnail of any candidate video.
[0057] In another implementation, as shown in Figure 4b, the first interface 411 can be an image display interface in a media processing client. For example, a user can preview an image p to be published through the first interface 411. The first interface 411 includes multiple functional interfaces for the currently displayed image p. The user can access the corresponding functional interface by triggering these interfaces. For example, the user can access the image p editing interface by triggering interface 412, or access the interface for generating a video based on the image p and the video selected by the user by triggering interface 413. Interface 413 can be a preset interface in the first interface 411 of this embodiment.
[0058] When the media processing client detects that the user has triggered the operation of interface 413, it can upload the currently displayed image P and switch the first interface 411 to the second interface 414 used for video generation in this embodiment. The second interface 414 includes a first area 415, which may include an interface 416. After image p is uploaded, a thumbnail of image p can be displayed in the first area 415. The user can also delete image p and select to upload other images by triggering the operation of interface 416. The second interface 414 may also include a second area 417, which may display thumbnails of cover images of multiple alternative videos for the user to select a reference video from.
[0059] In another implementation, as shown in Figure 4c, the first interface 421 can be the interface in the media processing client where a user browses videos generated by other users. For example, if user A generates video E using the video generation scheme provided in this embodiment and publishes video E, then the interface for user B to view video E published by user A can be the first interface. The first interface 421 includes a preset interface 422, which allows the user to access the video generation interface 423.
[0060] In step 303, the original image and reference video are determined through the second interface.
[0061] In this embodiment, the media processing client can determine the original image and reference video through the second interface. Specifically, the uploaded image displayed in the first area can be determined as the original image, and the candidate video selected through the second area can be determined as the reference video. As shown in Figure 4d, the user can upload the original image through the first area 432 in the second interface 431, and the image thumbnail of the uploaded image is displayed in the first area 432. The user can also delete the uploaded image or re-upload the image by operating on the first area 432. The user can also select a reference video through the second area 433. Specifically, the second area 433 displays the cover thumbnails of multiple candidate videos. The user can trigger an operation on the cover thumbnail corresponding to the candidate video to select a reference video from the candidate videos. For example, the user can click on the cover thumbnail of any candidate video, and the cover thumbnail can be highlighted in a preset manner. The cover thumbnail displays a preview button 434, and the user can preview the selected reference video by triggering the preview button 434. When previewing the reference video, you can choose to preview it silently by default, or you can turn on the sound during the preview to play the audio from the reference video.
[0062] In this embodiment, the original image includes a first object of a preset category, and the reference video includes a second object of a preset category that is moving. The object of the preset category can be a person, a cartoon character, an anthropomorphic animal, or an anthropomorphic object, etc. Therefore, the original image can include a person, a cartoon character, an anthropomorphic animal, or an anthropomorphic object, etc. The reference video can include moving people, cartoon characters, anthropomorphic animals, or anthropomorphic objects, etc.
[0063] In some embodiments, if the first interface is an image display interface, image recognition can be performed on the image currently displayed on the first interface to determine whether the image includes objects of a preset category. If the image includes objects of a preset category, a preset interface can be provided in the first interface. Furthermore, if the user triggers the preset interface, the media processing client can directly upload the image, use the image as the original image to be used, and display a thumbnail of the image in the first area of the switched second interface.
[0064] In other embodiments, after a user performs a trigger operation on the first area of the second interface, multiple candidate images can be output. These candidate images can be, for example, images from an image library. The user can select an image to upload from the multiple candidate images. The media processing client can identify the image to be uploaded to determine whether the image includes an object of a preset category. If the image includes an object of a preset category, the image can be uploaded directly, and a thumbnail of the image will be displayed in the first area of the second interface.
[0065] Since this embodiment can identify images before uploading them to determine whether they contain objects of a preset category, it can filter images in advance to avoid situations where the target video cannot be generated because the image does not contain objects of the preset category, thus improving the success rate of video generation.
[0066] In step 304, in response to a second trigger operation on the second interface, a target video corresponding to the original image is generated based on the reference video.
[0067] In this embodiment, the user can perform a second trigger operation on the second interface. After the media processing client detects the second trigger operation, it can generate a target video based on the reference video and the original image, causing the first object in the target video to move in the same way as the second object. For example, referring to Figure 4d, after uploading the original image through the first area 432 and selecting the reference video through the second area 433, the trigger button 435 in the second interface 431 changes from an unavailable state to an available state. The user can click the trigger button 435 to complete the second trigger operation on the second interface 431 and start generating the target video.
[0068] Specifically, since the reference video comprises multiple video frames, the reference keypoint feature information corresponding to the second object in each video frame of the reference video can be obtained first. This reference keypoint feature information can be the location information corresponding to the identified keypoints of the second object, such as coordinate information. The keypoints of the second object can be body keypoints, skeletal keypoints, facial keypoints, pose keypoints, etc. This embodiment does not limit the specific type of keypoints. As shown in Figure 4e, image 441 is a video frame included in the reference video, the person in image 441 is the second object, and image 442 can be the reference keypoint feature information corresponding to the second object.
[0069] In one implementation, computer vision technology and deep learning models can be directly used to detect the positional information of multiple key points corresponding to the second object in each video frame, thereby obtaining the reference key point feature information corresponding to the second object in each video frame. In another implementation, the reference key point feature information corresponding to the second object in each video frame of the candidate videos can be pre-acquired, and this information can be stored in association with the candidate videos. After the user selects any candidate video as a reference video, the reference key point feature information associated with that reference video can be retrieved from the pre-stored data.
[0070] Next, the first feature information corresponding to the first object and the second feature information corresponding to the background can be extracted from the original image. The first feature information may include, but is not limited to, the texture feature information, contour feature information, color feature information, and style feature information of the first object. The second feature information may include, but is not limited to, the texture feature information, style feature information, and color feature information of the background of the original image.
[0071] Finally, the target video can be generated based on the reference keypoint feature information, the first feature information, and the second feature information. Specifically, a noisy image can be acquired, and the noisy image, the reference keypoint feature information, the first feature information, and the second feature information can be input into the target model for denoising to obtain multiple target images. For example, for each video frame, using the reference keypoint feature information corresponding to that video frame, a multi-step denoising operation is performed iteratively based on the target model to generate the target image corresponding to that video frame, such that the first object in the target image has the same pose as the second object in the video frame. The multiple target images are then combined to obtain the target video.
[0072] It should be noted that the reference video may include not only image components (i.e., video frames) but also audio components. In one implementation, if the reference video includes audio, the audio portion can be directly obtained after generating multiple target images and combined with the target images to form the target video. In another implementation, if the reference video includes audio, the audio portion can be obtained after generating multiple target images and further processed, such as converting the audio portion to a specified style or replacing the audio portion, before combining the processed audio portion with the target images to form the target video.
[0073] Additionally, a video addition interface is displayed in the second area of the second interface. Users can add candidate videos through this interface. Specifically, users can trigger the video addition interface. After the media processing client detects this trigger, it outputs multiple candidate videos, which can be, for example, videos from a video library. Users can select the video to add from these candidate videos. After the media processing client determines the selected video, it further performs recognition processing on the video to determine whether it contains objects of a preset category. If it does not contain objects of the preset category, the video is rejected. If it does contain objects of the preset category, the client can obtain the key point feature information corresponding to the objects of the preset category in the video frames of the video to be added. For example, computer vision technology and deep learning models can be used to obtain the key point feature information corresponding to the objects of the preset category in the video frames of the video to be added. Then, the video to be added and the key point feature information are stored together, and a cover thumbnail of the video to be added is added and displayed in the second area.
[0074] As shown in Figure 4f, a video addition interface 453 is displayed in the second area 452 of the second interface 451. Users can add candidate videos through this interface. After the user triggers the video addition interface 453, the media processing client provides a video selection interface, allowing the user to select the video to be added. Once the user selects a video, a new candidate video option area 454 is created in the second area 452, displaying a thumbnail of the added video's cover image.
[0075] This disclosure provides a video generation scheme. By providing a preset interface in a first interface, upon detecting a first trigger operation on the preset interface, the first interface is switched to a second interface. The second interface determines an original image and a reference video. The original image includes a first object of a preset category, and the reference video includes a second object of a preset category undergoing motion. In response to a second trigger operation on the second interface, a target video corresponding to the original image is generated based on the reference video, causing the first object in the target video to move according to the motion pattern of the second object. This achieves the transfer of the motion pattern of the second object in the reference video to the first object in the original image, resulting in a target video corresponding to the original image. This efficiently completes the motion transfer of preset category objects, improves video generation quality, and enhances the user experience.
[0076] It should be noted that although the operations of the methods of this disclosure embodiment are described in a specific order in the above embodiments, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0077] Corresponding to the aforementioned video generation method embodiments, this disclosure also provides embodiments of a video generation apparatus.
[0078] As shown in FIG5, FIG5 is a block diagram of a video generation apparatus according to an exemplary embodiment of the present disclosure. The apparatus may include: a providing unit 501, a switching unit 502, a first determining unit 503, and a generation unit 504.
[0079] The providing unit 501 is configured to provide a preset interface in the first interface.
[0080] The switching unit 502 is configured to switch the first interface to the second interface after detecting the first trigger operation for the preset interface.
[0081] The first determining unit 503 is configured to determine an original image and a reference video through a second interface. The original image includes a first object of a preset category, and the reference video includes a second object of a preset category.
[0082] The generation unit 504 is configured to respond to a second trigger operation on the second interface, generate a target video corresponding to the original image based on the reference video, and make the first object in the target video move in the same way as the second object.
[0083] In some embodiments, the first interface is an image display interface, and the providing unit 501 is configured to: perform image recognition on the first image currently displayed on the first interface, and if the recognition result indicates that the first image includes an object of a preset category, display a preset interface in a preset area of the first interface.
[0084] In other embodiments, after detecting a first trigger operation for the preset interface, the device further includes: a first uploading unit configured to upload the first image as an original image to be used.
[0085] In other embodiments, the second interface includes a first area for uploading images and a second area for selecting videos. After uploading an image, the first area displays an image thumbnail of the uploaded image, and the second area displays cover thumbnails of multiple alternative videos.
[0086] In other embodiments, the first determining unit 503 is configured to: determine the uploaded image displayed in the first area as the original image, and determine the candidate video selected through the second area as the reference video.
[0087] In other embodiments, the device may further include: a first output unit, a selection unit, and an identification unit (not shown in the figure).
[0088] The first output unit is configured to output multiple candidate images in response to a trigger operation targeting the first region.
[0089] The selection unit is configured to determine the second image selected from a plurality of candidate images.
[0090] The recognition unit is configured to recognize the second image.
[0091] The second uploading unit is configured to upload the second image if the second image contains objects of a preset category.
[0092] In other embodiments, the generation unit 504 is configured to: acquire reference key point feature information corresponding to the second object in multiple video frames of the reference video, extract the first feature information corresponding to the first object and the second feature information corresponding to the background in the original image, and generate a target video based on the reference key point feature information, the first feature information and the second feature information.
[0093] In other embodiments, the generation unit generates a target video based on the reference key point feature information, the first feature information, and the second feature information in the following manner: acquiring a noisy image, inputting the reference key point feature information, the first feature information, and the second feature information into the target model for denoising operation to obtain multiple target images, and generating a target video based on the multiple target images.
[0094] In other embodiments, the second interface includes a second area for selecting videos, the second area displaying thumbnail images of multiple candidate video covers, and the second interface also includes a video adding interface for uploading candidate videos. The device further includes:
[0095] The second output unit is configured to output multiple candidate videos in response to a trigger operation that adds an interface to the video.
[0096] The second determining unit is configured to determine the video to be added from a plurality of candidate videos.
[0097] The acquisition unit is configured to acquire the key point feature information of objects of a preset category in the video frames of the video to be added.
[0098] The storage unit is configured to store the video to be added and the key point feature information to be added in association.
[0099] The display unit is configured to display a thumbnail of the cover of the video to be added in the second area.
[0100] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the embodiments of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0101] Referring now to FIG6, FIG6 is a schematic block diagram of an electronic device provided in some embodiments of the present disclosure. This electronic device 920 is, for example, suitable for implementing the video generation method provided in the embodiments of the present disclosure. The electronic device 920 can be a terminal device, etc., and can be used to implement a client or server. The electronic device 920 can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. It should be noted that the electronic device 920 shown in FIG6 is merely an example and does not impose any limitations on the functionality and scope of use of the embodiments of the present disclosure.
[0102] As shown in Figure 6, the electronic device 920 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 921, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 922 or a program loaded from a storage device 928 into a random access memory (RAM) 923. The RAM 923 also stores various programs and data required for the operation of the electronic device 920. The processing unit 921, ROM 922, and RAM 923 are interconnected via a bus 924. An input / output (I / O) interface 925 is also connected to the bus 924.
[0103] Typically, the following devices can be connected to I / O interface 925: input devices 926 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 927 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 928 including, for example, magnetic tapes, hard disks, etc.; and communication devices 929. Communication device 929 allows electronic device 920 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 6 shows electronic device 920 with various devices, it should be understood that it is not required to implement or possess all the devices shown, and electronic device 920 may alternatively implement or possess more or fewer devices. Each box shown in Figure 6 may represent one device, or multiple devices may be represented as needed.
[0104] According to embodiments of this disclosure, the video generation method described above can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the video generation method described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 929, or installed from a storage device 928, or installed from a ROM 922. When the computer program is executed by a processing device 921, the functions defined in the video generation method provided by embodiments of this disclosure can be implemented.
[0105] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the methods provided in this disclosure.
[0106] It should be noted that the computer-readable medium described in the embodiments of this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the embodiments of this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.
[0107] Computer program code for performing the operations of embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0108] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for storage media and computing devices are basically similar to the method embodiments, so they are described more simply; relevant parts can be referred to the descriptions of the method embodiments.
[0109] Those skilled in the art will recognize that the functions described in the embodiments of this disclosure in one or more of the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.
[0110] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the embodiments of this disclosure. It should be understood that the above descriptions are merely specific implementations of the embodiments of this disclosure and are not intended to limit the scope of protection of this disclosure. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating a video, the method comprising: A preset interface is provided in the first interface; After detecting a first trigger operation for the preset interface, the first interface is switched to the second interface; The second interface determines the original image and the reference video; the original image includes a first object of a preset category; the reference video includes a second moving object of the preset category; In response to a second trigger operation on the second interface, a target video corresponding to the original image is generated based on the reference video, so that the first object in the target video moves in the manner of the second object.
2. The method according to claim 1, wherein, The first interface is an image display interface; The first interface provides a preset interface, including: Perform image recognition on the first image currently displayed on the first interface; If the recognition result indicates that the first image includes an object of the preset category, the preset interface is displayed in a preset area of the first interface.
3. The method according to claim 2, wherein, After detecting a first trigger operation for the preset interface, the method further includes: uploading the first image as the original image to be used.
4. The method according to any one of claims 1-3, wherein, The second interface includes a first area for uploading images and a second area for selecting videos; after uploading an image, the first area displays an image thumbnail of the uploaded image; the second area displays cover thumbnails of multiple candidate videos.
5. The method according to claim 4, wherein, The step of determining the original image and reference video through the second interface includes: The uploaded image displayed in the first area is identified as the original image; and The candidate videos selected through the second area will be determined as reference videos.
6. The method according to claim 4, wherein, The method further includes: In response to a trigger operation targeting the first region, multiple candidate images are output; Determine the second image selected from the plurality of candidate images; The second image is then identified; If the second image includes an object of the preset category, then upload the second image.
7. The method according to any one of claims 1-3, wherein, The step of generating a target video corresponding to the original image based on the reference video includes: Obtain the reference key point feature information corresponding to the second object in multiple video frames of the reference video; Extract the first feature information corresponding to the first object and the second feature information corresponding to the background from the original image; The target video is generated based on the reference key point feature information, the first feature information, and the second feature information.
8. The method according to claim 7, wherein, The generation of the target video based on the reference key point feature information, the first feature information, and the second feature information includes: Acquire noisy images; The noisy image, the reference key point feature information, the first feature information, and the second feature information are respectively input into the target model for denoising to obtain multiple frames of target images. The target video is generated based on the multiple target images.
9. The method according to claim 7, wherein, The second interface includes a second area for selecting videos, which displays thumbnails of the cover images of multiple candidate videos; The second interface also includes a video adding interface for uploading alternative videos; wherein, the method further includes: In response to a trigger operation on the video adding interface, multiple candidate videos are output; Determine the video to be added from the plurality of candidate videos; Obtain the key point feature information of the object of the preset category in the video frame of the video to be added; The video to be added and the key point feature information to be added are stored together in association; The second area displays a thumbnail of the cover of the video to be added.
10. A video generation apparatus, the apparatus comprising: The unit is configured to provide a preset interface in the first interface; The switching unit is configured to switch the first interface to the second interface after detecting a first trigger operation for the preset interface; The first determining unit is configured to determine an original image and a reference video through the second interface; the original image includes a first object of a preset category; the reference video includes a second moving object of the preset category; The generation unit is configured to, in response to a second trigger operation on the second interface, generate a target video corresponding to the original image based on the reference video, such that the first object in the target video moves in the manner of the second object.
11. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-9.
12. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-9.
13. A computer program product, wherein, The computer program product includes a computer program that, when executed by a computer device, causes the computer device to perform the method according to any one of claims 1-9.
14. A computer program, wherein, When the computer program is executed by a computer device, the computer device performs the method according to any one of claims 1-9.