Image animation method, electronic device, and computer program product
By transferring the motion patterns of a reference video to the target object of the input image based on semantic mapping in the image animation method, the problem of monotonous output video and difficulty for users to intuitively understand motion patterns in the existing technology is solved, and high-quality, vivid video is generated.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MICROSOFT TECHNOLOGY LICENSING LLC
- Filing Date
- 2021-04-16
- Publication Date
- 2026-05-05
AI Technical Summary
Existing image animation methods typically apply pre-obtained motion patterns directly to the entire image without considering the semantic differences between different objects in the image. This results in output videos that are rather monotonous and rigid, and users find it difficult to intuitively understand the motion patterns.
By acquiring an input image and a reference video, the motion pattern of the reference object in the reference video is determined, and then transferred to the target object in the input image based on semantic mapping to generate an output video, so that the motion of the target object in the output video has the motion pattern of the reference object.
It achieves high-quality video generation, allowing users to intuitively understand motion patterns. Every pixel in the output video comes from the original input image, resulting in high realism and vividness.
Smart Images

Figure CN115222859B_ABST
Abstract
Description
Background Technology
[0001] Image animation refers to the automatic generation of dynamic videos from static images. Compared to static images, dynamic videos are more vivid and expressive, thus enhancing the user experience. Currently, image animation is widely used to generate dynamic backgrounds, live wallpapers, and more. However, the quality of the generated videos still needs improvement. Therefore, there is a need for image animation methods capable of generating high-quality videos. Summary of the Invention
[0002] According to the implementation of this disclosure, a scheme for generating video from images is proposed. In this scheme, an input image and a reference video are acquired. Based on the reference video, the motion pattern of a reference object in the reference video is determined. An output video is generated with the input image as the starting frame, and the motion of the target object in the input image in the output video follows the motion pattern of the reference object. In this way, the scheme can intuitively generate an output video based on an input image and a reference video, and the motion of the target object in the output video follows the motion pattern of the reference object in the reference video.
[0003] The summary section is provided to present the chosen concepts in a simplified form, which will be further described in the detailed description below. The summary section is not intended to identify key or principal features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0004] Figure 1 A block diagram of a computing device capable of implementing multiple implementations of the present disclosure is shown;
[0005] Figure 2 An architecture diagram of a system for image animation according to an implementation of this disclosure is shown;
[0006] Figure 3 An architecture diagram of a system for training an image animation model according to the implementation of this disclosure is shown;
[0007] Figure 4 A flowchart illustrating a method for animate images according to the present disclosure is shown;
[0008] Figure 5 A flowchart illustrating a method for training an image animation model according to an implementation of this disclosure is shown; and
[0009] In these accompanying figures, the same or similar reference symbols are used to indicate the same or similar elements. Detailed Implementation
[0010] This disclosure will now be discussed with reference to several example implementations. It should be understood that these implementations are discussed only to enable those skilled in the art to better understand and thus implement this disclosure, and not to imply any limitation on the scope of this disclosure.
[0011] As used herein, the term "comprising" and its variations are to be interpreted as open-ended terms meaning "including but not limited to". The term "based on" is to be interpreted as "at least partially based on". The terms "an implementation" and "an implementation" are to be interpreted as "at least one implementation". The term "another implementation" is to be interpreted as "at least one other implementation". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0012] As used in this paper, a "neural network" is capable of processing input and providing corresponding output. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus extending the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer, where the input layer receives the input to the neural network, and the output layer's output serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the layer above. A CNN is a type of neural network that includes one or more convolutional layers for performing convolution operations on their respective inputs. CNNs can be used in a variety of scenarios and are particularly well-suited for processing image or video data. In this paper, the terms "neural network," "network," and "neural network model" are used interchangeably.
[0013] As mentioned above, dynamic videos are more vivid and engaging than static images, thus enhancing the user experience. However, conventional image animation methods typically apply pre-obtained motion patterns directly to the entire image without considering the semantic differences between various objects within the image, resulting in monotonous and rigid output videos. Furthermore, pre-obtained motion patterns are often stored and invoked in the form of code, making it difficult for users to intuitively understand them. Consequently, users struggle to visualize the output video generated by applying motion patterns. For example, users might need to experiment with many different motion patterns to find the optimal one that produces the desired output video. Therefore, there is a need for image animation methods that can generate high-quality videos in an intuitive way.
[0014] The above discussion addresses some problems existing in conventional schemes for generating video from images. Based on an implementation of this disclosure, a scheme for generating video from images is proposed, aiming to solve one or more of the aforementioned problems and other potential issues. In this scheme, an input image and a reference video are acquired. Based on the reference video, the motion pattern of a reference object in the reference video is determined. An output video is generated with the input image as the starting frame, wherein the motion of the target object in the input image in the output video has the motion pattern of the reference object. Various example implementations of this scheme are further described in detail below with reference to the accompanying drawings.
[0015] Figure 1 A block diagram of a computing device 100 capable of implementing multiple implementations of the present disclosure is shown. It should be understood that... Figure 1 The computing device 100 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the implementation described in this disclosure. Figure 1 As shown, computing device 100 includes computing device 100 in the form of general computing device. Components of computing device 100 may include, but are not limited to, one or more processors or processing units 110, memory 120, storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.
[0016] In some implementations, computing device 100 can be implemented as various user terminals or service terminals with computing capabilities. Service terminals can be servers, large computing devices, etc., provided by various service providers. User terminals can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, sites, units, devices, multimedia computers, multimedia tablets, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is also foreseeable that computing device 100 can support any type of user-facing interface (such as "wearable" circuitry).
[0017] Processing unit 110 can be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 120. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 100. Processing unit 110 may also be referred to as a central processing unit (CPU), graphics processing unit (GPU), microprocessor, controller, or microcontroller.
[0018] Computing device 100 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 100, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 120 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Memory 120 may include graphics animation modules 122, which are configured to perform the functions of the various implementations described herein. Graphics animation modules 122 can be accessed and run by processing unit 110 to implement the corresponding functions.
[0019] Storage device 130 may be a removable or non-removable medium and may include machine-readable media capable of storing information and / or data and accessible within computing device 100. Computing device 100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 1 As shown, disk drives for reading from or writing to removable, non-volatile disks and optical disc drives for reading from or writing to removable, non-volatile optical discs can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces.
[0020] The communication unit 140 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 100 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 100 can operate in a networked environment using logical connections to one or more other servers, personal computers (PCs), or another general network node.
[0021] Input device 150 can be one or more various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 160 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 100 can also communicate as needed with one or more external devices (not shown) via communication unit 140. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 100, or with any device that enables computing device 100 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interfaces (not shown).
[0022] In some implementations, in addition to being integrated into a single device, some or all of the components of computing device 100 may be configured in the form of a cloud computing architecture. In a cloud computing architecture, these components can be remotely deployed and can work together to achieve the functionality described herein. In some implementations, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various implementations, cloud computing provides services over a wide area network (WAN), such as the Internet, using appropriate protocols. For example, cloud computing providers offer applications over a WAN, and these applications can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated at remote data center locations or they may be distributed. Cloud computing infrastructure can provide services through shared data centers, even if they appear as a single access point for users. Therefore, the components and functionality described herein can be provided from service providers at remote locations using a cloud computing architecture. Alternatively, they may also be provided from conventional servers, or they may be installed directly or otherwise on client devices.
[0023] The computing device 100 can generate video from images according to various implementations of this disclosure. For example... Figure 1 As shown, computing device 100 can receive input image 170 and reference video 171 via input device 150. Input device 150 can transmit input image 170 and reference video 171 to image animation module 122. Image animation module 122 generates output video 190 based on input image 170 and reference video 171. Output video 190 includes input image 170 as the starting frame, and at least one subsequent output frame, such as... Figure 1 The output frames shown are 190-1 and 190-2.
[0024] For example, input image 170 can be a landscape image to be processed. As shown, input image 170 can be a landscape image containing objects such as grass, blue sky, and white clouds. In other implementations, input image 170 can be a landscape image containing objects such as waterfalls, starry skies, and lakes. It should be understood that video can also be generated based on input image 170, which is not a landscape image. Reference video 171 can be an input video used to extract reference motion patterns. Reference video 171 consists of at least two consecutive frames. If input image 170 is a landscape image containing grass, blue sky, and white clouds, reference video 171 can also be a video containing grass, blue sky, and white clouds. In this case, motion patterns of grass, blue sky, and white clouds can be extracted from reference video 171 and applied to grass, blue sky, and white clouds in input image 170 respectively to generate output video 190. Figure 1 As shown, the clouds in the output video 190 move in position (as indicated by the circles) in consecutive frames 170, 190-1, and 190-2, and the motion of the clouds in the output video 190 follows the motion pattern of the clouds in the reference video 171. In some implementations, the objects in the reference video 171 may not correspond to the objects in the input image 170. In this case, the motion pattern of any object in the reference video 171 can be applied to the target object in the input image 170 according to the user's instructions to generate the output video 190 as desired by the user.
[0025] Figure 2 An architecture diagram of a system 200 for image animation according to the present disclosure is shown. System 300 can be implemented in... Figure 1 The computing device 100. The system 300 can be an end-to-end neural network model.
[0026] like Figure 2 As shown, computing device 100 can determine a first reference motion feature 210 based on the first frame 171-1 and the subsequent second frame 171-2 of reference video 171. The first reference motion feature 210 characterizes the motion from the first frame 171-1 to the second frame 171-2. The first reference motion feature 210 can be represented using optical flow. Optical flow is the instantaneous velocity of pixel motion of a moving object in space. When the time interval is very small (e.g., between two consecutive frames of a video), optical flow is also equivalent to the displacement of the target point. Optical flow can be represented using a two-channel image of the same size as the first and second frames (representing displacement in the x and y directions, respectively). Optical flow estimation can be used to determine the first reference motion feature 210 represented by optical flow. Optical flow estimation finds the correspondence between the first and second frames by utilizing the temporal changes of pixels in the image sequence and the correlation between adjacent frames, thereby calculating the motion information of the object between adjacent frames. In other words, optical flow estimation can determine a first reference motion feature 210 used to characterize the motion from the first frame 171-1 to the second frame 171-2. In some implementations, FlowNet 2.0 can be used to extract the optical flow F based on the t-th and t+1-th frames. tWhen using optical flow estimation to determine the first reference motion feature, the displacement of the target point between adjacent frames should be relatively small due to the inherent limitations of the optical flow method. Therefore, the object in the image should not undergo significant positional changes between adjacent frames. In this case, the reference video 171 depicting landscape changes can extract a more accurate first reference motion feature 210 compared to a reference video depicting ball sports (where the object's position changes more significantly). It should be understood that the scope of this disclosure does not limit the method for determining the first reference motion feature 210. When using other methods to determine the reference motion feature, the magnitude of change of the object in the reference video 171 between adjacent frames is not limited. In other words, the type of reference video 171 is not limited to a landscape image.
[0027] In some implementations, computing device 100 can determine a first set of reference objects in the first frame 171-1 and at least one object in the input image 170. The at least one object may include a target object. Reference Figure 2 The first set of reference objects in the first frame 171-1 may include blue sky, white clouds, trees, and a lake. At least one object in the input image 170 may include blue sky, white clouds, trees, a lake, and a bridge. In some implementations, the target object in the input image 170 may be white clouds. In other implementations, the target object in the input image 170 may include white clouds and trees.
[0028] Various methods can be used to determine the first set of reference objects in the first frame 171-1 and at least one object in the input image 170. For example, semantic segmentation methods can be used to determine objects with corresponding semantics. In this case, objects with the same semantics correspond to the same semantic segmentation mask. In some implementations, a first set of semantic segmentation masks 220 can be generated by performing semantic segmentation on the first frame 171-1. The first set of semantic segmentation masks 220 can indicate the corresponding positions of the first set of reference objects in the first frame 171-1. For example, reference... Figure 2 The first set of semantic segmentation masks 220 may include semantic segmentation masks corresponding to the blue sky, trees, and lakes, respectively. Semantic segmentation masks can be represented using binary masks. For example, assuming the first frame 171-1 has 300×300 pixels, the semantic segmentation mask for each semantic meaning can be represented by a 300×300 matrix. The binary value of each element in the matrix indicates whether an object with that semantic meaning exists at the corresponding location. (Reference) Figure 2In the first frame 171-1, the lake exists only at the bottom of the image; therefore, the semantic segmentation mask for the lake has a value of 1 only at the bottom of the matrix and zero values elsewhere. Similarly, at least one semantic segmentation mask can be generated by performing semantic segmentation on the input image 170. This at least one semantic segmentation mask can indicate the corresponding location of at least one object in the input image 170. (Reference) Figure 2 The system can generate semantic segmentation masks corresponding to blue sky, white clouds, trees, lakes, and bridges, respectively, based on the input image 170. Each semantic segmentation mask can indicate the corresponding position of an object with the corresponding semantic meaning in the input image 170. In some implementations, object determination methods other than semantic segmentation can be used to determine a first set of reference objects in the first frame 171-1 and at least one object in the input image 170. The scope of this disclosure is not limited herein.
[0029] In some implementations, computing device 100 can generate a first predicted motion feature 240 for input image 170 based on first reference motion feature 210 by determining a semantic mapping from at least one reference object in a first set of reference objects to at least one object in input image 170. As described above, the first reference motion feature 210 is used to characterize the motion from first frame 171-1 to second frame 171-2 and can be represented by optical flow. Similarly, the first predicted motion feature 240 can also be represented by optical flow. The first predicted motion feature 240 can characterize the motion that will occur in input image 170. It should be noted that, for example, the first reference motion feature 210 obtained by optical flow estimation can only describe the overall motion of the image and cannot describe the motion of each reference object in the image individually. In this case, if the first reference motion feature 210 is directly applied to input image 170, it is difficult to make the target objects in input image 170 move according to the desired motion pattern. For example, the tree in input image 170 may also move according to the motion pattern of the lake in reference video 171. Therefore, if the first reference motion feature 210 is directly used as the first predicted motion feature 240 for the input image 170, the quality of the generated output video 190 may not be high.
[0030] Therefore, in some implementations, computing device 100 can determine a motion pattern 230 for each reference object in reference video 171 based on first reference motion features 210, and use the motion patterns 230 to generate a first predicted motion feature 240. In some implementations, computing device 100 can transfer the motion pattern 230 of the reference object to the target object based on a semantic mapping from the reference object to the target object. In this case, the motion pattern 230 for the reference object can be transferred to the target object in input image 170 by determining a semantic mapping from at least one reference object in the first set of reference objects to at least one object in input image 170. For example, by mapping a lake in reference video 171 to a lake in input image 170, the motion pattern of the lake in reference video 171 can be transferred to the lake in input image 170. And by mapping a tree in reference video 171 to a tree in input image 170, the motion pattern of the tree in reference video 171 can be transferred to the tree in input image 170.
[0031] In some implementations, computing device 100 may also transfer the motion pattern 230 of the reference object to the target object based on predetermined rules that indicate the mapping from the reference object to the target object in the input image 170. The predetermined rules can be customized by the user. Examples of predetermined rules may include mapping based on object color, mapping based on object shape, and so on. Additionally, the predetermined rules may also indicate that certain objects in the input image 170 remain stationary in the generated output video 190 if they are not mapped.
[0032] In some implementations, the computing device 100 may also transfer the motion patterns of the additional reference objects to the additional target objects based on the mapping of additional reference objects in the additional reference video to additional target objects in the input image 170. The additional reference video may include one or more reference videos. In this case, the corresponding motion patterns 230 of the multiple reference objects can be transferred to the multiple target objects in the input video 170 by determining the mappings of multiple reference objects in the multiple reference videos to multiple target objects in the input image 170, respectively. For example, a lake in the first reference video can be mapped to a lake in the input image 170, thereby transferring the motion pattern 230 of the lake in the first reference video to the lake in the input image 170. Similarly, a tree in the second reference video can be mapped to a tree in the input image 170, thereby transferring the motion pattern of the tree in the second reference video to the tree in the input image 170.
[0033] In some implementations, computing device 100 can combine motion patterns 230 of reference objects in reference video 171 into combined motion patterns for input image 170 based on various mapping rules. Figure 2 (Not shown in the image). For example, see reference. Figure 2 A combined motion pattern for the input image 170 can be generated by mapping the reference objects blue sky and trees in the first frame 171-1 to the objects blue sky and trees in the input image 170, respectively. Additionally, when some objects in the input image 170 are not mapped, the values of the corresponding elements in the combined motion pattern corresponding to these objects can be set to zero. Based on the combined motion pattern, a first predicted motion feature 240 for the input image 170 can be determined. For example, the first predicted motion feature 240 can be extracted using a motion prediction convolutional neural network. The following will refer to... Figure 3 The details of generating the combined motion pattern and generating the first predicted motion feature 240 are described in detail.
[0034] In some implementations, computing device 100 may generate a first output frame 190-1 in output video 190, located after the starting frame, based on a first predicted motion feature 240 and input image 170. As described above, the first predicted motion feature 240 can be used to characterize the motion that will occur in input image 170, and this motion corresponds to the motion in reference video 171 from first frame 171-1 to second frame 171-2. In some implementations, the input image 170 may be warped based on the first predicted motion feature 240 to generate the first output frame 190-1. Image warping can refer to some kind of deformation or transformation between images. For example, image warping may include re-sampling the image and interpolating the sampled points. In some implementations, image warping may refer to optical flow mapping. In other words, the input image 170 may be optically flow-mapped based on optical flow to generate the first output frame 190-1. The scope of this disclosure does not limit the type of image warping used to generate the first output frame 190-1.
[0035] In some implementations, the computing device 100 may also generate a second output frame 190-2 in the output video 190, located after the first output frame 190-1, based on the generated first output frame 190-1. Similarly, a second reference motion feature may be determined based on the second frame 171-2 and the subsequent third frame of the reference video 171. The second reference motion feature is used to characterize the motion from the second frame 171-2 to the third frame. A second set of reference objects in the second frame 171-2 and at least one object in the first output frame 190-1 may be determined. It should be noted that since the first output frame 190-1 is generated by optical flow mapping of the input image 170, the objects included in the first output frame 190-1 should be consistent with the objects included in the input image 170. However, due to the motion of the objects, the corresponding semantic segmentation mask for each object may change. By determining the semantic mapping from at least one reference object in the second set of reference objects to at least one object in the first output frame 190-1, a second predicted motion feature for the first output frame 190-1 may be generated based on the second reference motion feature. It should be noted that it is generally desirable for the target object in the output video 190 to move according to the motion pattern of the same reference object in consecutive output frames. In this case, the semantic mapping relationship from the reference object to the target object is usually maintained. Based on the generated first predicted motion feature 240, the second predicted motion feature, and the input image, a second output frame 190-2 in the output video 190, located after the first output frame 190-1, can be generated. It should be noted that, unlike the generation of the first output frame 190-1, the input image 170 can be transformed based on the sum of the first predicted motion feature 240 and the second predicted motion feature to generate the second output frame 190-2. By directly transforming the input image 170 to generate the second output frame 190-2, the error accumulation of the generated output frames can be reduced, thereby making the output video 190 more realistic.
[0036] In some implementations, reference video 171 may also include multiple reference videos. In this case, for each of the multiple reference videos 171, corresponding reference motion features 210 corresponding to multiple sets of reference objects can be determined based on adjacent frames, and corresponding semantic segmentation masks 220 can be generated. Based on this, multiple sets of motion patterns 230 for the multiple reference videos 171 can be generated. Similarly, the multiple sets of motion patterns 230 for reference objects can be combined into a combined motion pattern for the input video 170 by determining the semantic mapping from at least one reference object in the multiple sets of reference objects to at least one object in the input image 170. For example, a lake in the first reference video can be mapped to a lake in the input image 170 and a tree in the second reference video can be mapped to a tree in the input image 170, thereby using the motion pattern of the lake in the first reference video and the motion pattern of the tree in the second reference video to form a combined motion pattern to generate a first predicted motion feature 240 for the input image 170.
[0037] In this manner, the implementation of this disclosure can generate an output video 190 from an input image 170 based on the motion of a reference object in reference video 171, and the motion of the target object in the generated output video 190 has the motion pattern of the reference object. Users can transfer the motion patterns of multiple reference objects in the reference video to their corresponding target objects. Furthermore, due to the introduction of reference video 171, users can intuitively understand the motion patterns to be applied to the input image 170. Therefore, users can easily visualize the motion effect of the target object in the generated output video 190 without having to go through multiple trials to select their preferred motion effect. Moreover, since the output video 190 is generated by directly performing optical flow mapping on the input image 170, each pixel in the output video 190 comes from the original input image 170, thus the output video 190 has high realism. Furthermore, embodiments of this disclosure can generate an output video 190 by directly inputting reference video 171 and input image 170, and the motion of the target object in the input image 170 in the output video 190 has the motion pattern of the reference object. In other words, the user does not need to predict the motion pattern of the target object based on the motion pattern 230 of the reference object, but directly transfers the motion pattern 230 of the reference object to the target object, so that the motion of the target object in the output video 190 has the motion pattern 230 of the reference object.
[0038] The above is for reference only. Figure 2 A system 200 for image animation is described. References will be made below. Figure 3 This describes a system 300 used for training image animation models. Figure 3 An architecture diagram of a system 300 for training an image animation model according to an implementation of this disclosure is shown. System 300 can be implemented in... Figure 1 In computing device 100. For example... Figure 3 As shown, system 300 acquires a training video, which includes a first training frame 301-1 and a subsequent second training frame 301-2. System 300 uses a machine learning model to generate a prediction video for the training video. The motion of objects in the prediction video follows the motion pattern of those objects in the training video, and the prediction video includes the first training frame 301-1 and a subsequent prediction frame 380 for the second training frame 301-2. System 300 trains the machine learning model based at least on the prediction frame 380 and the second training frame 301-2.
[0039] It should be noted that in system 300, the training video not only corresponds to the reference video 171 used by the image animation model during inference, but the first training frame 301 in the training video also corresponds to the input image 170 used during inference. Therefore, the predicted video generated by the machine learning model based on the first training frame 301 (corresponding to the input image 170) and the training video (corresponding to the reference video 171) is expected to be consistent with the training video itself. In other words, the second training frames 301-2 in the training video can be used as labels for the predicted frames 380 to train the image animation model. In this way, the training video itself can be used as a label to supervise the training of the machine learning model for generating videos from images. To enable the trained model to be widely applied to other unseen input images and reference videos, additional spatial transformation operations can be performed in the implementation of this disclosure to prevent the model from overlearning spatial relationships between training videos. Details of the spatial transformation operations will be referred to below. Figure 3 Detailed description.
[0040] like Figure 3 As shown, system 300 can generate training motion features 302 based on the first training frame 301-1 and the second training frame 301-2. Training motion features 302 are used to characterize the motion from the first training frame 301-1 to the second training frame 301-2. In some implementations, system 300 can also determine a set of training objects in the first training frame 301-1. (See reference...) Figure 2 As described, semantic segmentation can be performed on the first training frame 301-1 to generate a corresponding set of semantic segmentation masks 303. The set of semantic segmentation masks 303 can indicate the corresponding positions of a set of training objects in the first training frame 301-1.
[0041] System 300 can perform the same spatial transformation on both the training motion features 302 and a set of training semantic segmentation masks 303 to generate transformed training motion features 310 and a transformed set of training semantic segmentation masks 320. The spatial transformation can disrupt the spatial association between the first training frame 301-1 in the training video corresponding to the input image 170 and the training video corresponding to the reference video, thereby giving the trained image animation model stronger generalization ability. Furthermore, performing the same spatial transformation on both the training motion features 302 and the set of training semantic segmentation masks 303 can maintain the spatial consistency between the position of objects in the image and the motion features. Examples of spatial transformations can include horizontal flipping, vertical flipping, rotation, etc. For example, Figure 3 An example of a horizontally flipped spatial transformation is shown in the figure.
[0042] In some implementations, system 300 can generate motion patterns 340 for each training object based on transformed training motion features 310 and a transformed set of training semantic segmentation masks 320. For example... Figure 3 As shown, the motion pattern 340 of the training object in the first training frame 301-1 can be determined using the partial convolution module 330. Partial convolution has been widely used in image completion, image restoration and other fields. In the implementation of this disclosure, the training motion features 310 can be partially convolved using the transformed semantic segmentation mask 320 of each training object to determine the motion pattern of each training object. Specifically, partial convolution can be performed with reference to formula (1).
[0043] (1)
[0044] Where X represents the feature part corresponding to the convolution kernel in the transformed training motion feature 310, M represents the mask part corresponding to the convolution kernel in the semantic segmentation mask for a specific training object, W and b are the learnable weight parameters and biases, respectively, sum(1) represents the sum of the elements of the all-1 matrix of the convolution kernel size, sum(M) represents the sum of the elements of the mask part, and sum(1) / sum(M) can be used to compensate for the difference in area of different objects in the convolution sliding window.
[0045] As can be seen from formula (1), by using a mask specific to the training object to convolve the transformed training motion features 310, only the transformed training motion features 310 at the positions corresponding to the specific training object are passed to subsequent calculations. In other words, only the motion features specific to the reference object are preserved.
[0046] Furthermore, after each partial convolution, the semantic segmentation mask 320 can be updated using formula (2).
[0047] (2)
[0048] Equation (2) shows that as long as one element of the mask region corresponding to the convolution kernel corresponds to a reference object, the mask corresponding to the center position of the convolution kernel will be updated to 1. In other words, the transformed training motion feature 310 corresponding to the center position of the convolution kernel is considered to be related to the specific training object and is therefore retained for subsequent calculation. In this way, the motion pattern 340 for each training object can be determined based on the transformed training motion feature 310 describing the overall motion of the image.
[0049] In some implementations, predicted motion features 370 for the first training frames 301-1 can be generated based on the motion pattern 340 for each training object. It should be noted that here, the first training frame 301 corresponds to the input image 170 used during inference. Figure 3 As shown, a combined motion pattern 350 for the first training frame 301-1 can be generated based on the semantic segmentation mask 303 of the first training frame 301-1 and the motion pattern 340 for the training object. Based on the combined motion pattern 350 and the first training frame 301-1, predicted motion features 370 for the first training frame 301-1 can be determined. (See reference...) Figure 2 As described, based on the semantic segmentation mask 303 of at least one object in the first training frame 301-1, the motion patterns of the mapped training objects can be combined into a combined motion pattern 350 for the first training frame 301-1 using the semantic mapping from at least one object in the first training frame 301-1 to the training object. For example, refer to Figure 3 The lake in the first training frame 301-1 (corresponding to the input image 170) can be mapped to the lake in the first training frame 301-1 (corresponding to the first frame 171-1 in the reference video 171). It should be understood that in system 300, since the first training frame 301-1 corresponds to the first frame 171-1 in the input image 170 and the reference video 171, the semantic mapping operation can be omitted and the motion pattern 340 can be directly used to combine them into a combined motion pattern 350 for the first training frame 301-1. The combined motion pattern 350 can be determined using formula (3).
[0050] L ( i , j ) = z k ,when M k ( i , j = 1 (3)
[0051] in, L( i , j ) indicates position ( i , j The combined motion pattern 350 values (represented as a vector). M k This refers to the object in the first training frame 301-1 (corresponding to input image 170). k semantic segmentation mask, z k This refers to the object in the first training frame 301-1 (corresponding to the first frame 171-1 in reference video 171). k The motion pattern. Formula (3) shows that it can conform to the motion pattern. M k ( i , j = 1 condition position ( i , j The corresponding vector of ) is determined as z. k Therefore, based on z k To combine them into a combined motion pattern 350 for the first training frame 301-1.
[0052] In some implementations, the predicted motion feature 370 for the first training frame 301-1 can be determined based on the combined motion pattern 350 and the first training frame 301-1. As mentioned above, the predicted motion feature 370 can also be represented using optical flow. A convolutional neural network can be used to predict the motion feature. For example, the combined motion pattern 350 can be concatenated with the first training frame 301-1 and input into a U-Net network for predicting motion features to generate the predicted motion feature 370.
[0053] In some implementations, the predicted frame 380 can be generated based on the predicted motion features 370 and the first training frame 301-1. As described above, the first training frame 301-1 can be image transformed based on the predicted motion features 370 to generate the predicted frame 380. In some implementations, the first training frame 301-1 can be optically mapped based on the predicted motion features 370 represented by optical flow to generate the predicted frame 380.
[0054] In some implementations, at least one loss function can be determined for training the image animation model. A target loss function can be determined by weighted summation of at least one loss function. The image animation model can be trained by minimizing the target loss function.
[0055] In some implementations, the frame loss function can be determined based on the difference between the predicted frame 380 and the second training frames 301-2. For example, the L1 norm between predicted frame 380 and the second training frame 301-2 can be calculated using formula (4):
[0056] (4)
[0057] in and These represent predicted frame 380 and second training frames 301-2, respectively. This represents the frame loss function.
[0058] In some implementations, the motion feature loss function can be determined based on the difference between the predicted motion feature 370 and the trained motion feature 302. For example, the L1 norm between predicted motion feature 370 and trained motion feature 302 can be calculated using formula (5):
[0059] (5)
[0060] in and These represent the predicted motion feature 370 and the training motion feature 302 for the first training frame 301-1, respectively. This represents the motion feature loss function.
[0061] In some implementations, the smoothness loss function can be determined based on the spatial distribution of the predicted motion features 370 and the spatial distribution of the predicted frames 380. By introducing a smoothness loss function This can result in higher smoothness of the motion features predicted by the image animation model. For example, the smoothness loss function can be calculated by referring to formulas (6) and (7). :
[0062] (6)
[0063] (7)
[0064] in and These represent the pixels at position p and the neighboring position q in the predicted frame 380, respectively. and Let p and q represent the feature values at position p and the neighboring q position in the predicted motion feature 370, respectively. This represents the smoothness loss function. Furthermore, formula (7) defines the smoothness loss function in formula (6). function.
[0065] As mentioned above, the target loss function can be determined by weighted summation of at least one loss function. For example, the target loss function can be calculated using formula (8). , where α, β and γ This represents the corresponding coefficient.
[0066] (8)
[0067] The image animation model can be trained by minimizing the objective loss function. The trained image animation model can generate an output video 190 based on the input image 170 and the reference video 171, such that the motion of the target object in the output video 190 follows the motion pattern of the reference object in the reference video 171.
[0068] Figure 4 A flowchart of a method 400 for image animation according to some implementations of the present disclosure is shown. Method 400 may be implemented by computing device 100, for example, it may be implemented at image animation module 122 in memory 120 of computing device 100.
[0069] like Figure 4 As shown, at box 410, computing device 100 acquires input image 170 and reference video 171. At box 420, computing device 100 determines the motion pattern 230 of a reference object in reference video 171 based on reference video 171. At box 430, computing device 100 generates output video 190 with input image 170 as the starting frame. The motion of the target object in input image 170 in output video 190 has the motion pattern of the reference object in reference video 171.
[0070] In some implementations, generating the output video involves transferring the motion patterns of the reference object to the target object based on a semantic mapping from the reference object to the target object.
[0071] In some implementations, generating an output video includes at least one of the following: transferring the motion pattern of a reference object to a target object based on a predetermined rule indicating the mapping from a reference object to a target object; and transferring the motion pattern of an additional reference object to an additional target object based on the mapping from an additional reference object in the additional reference video to an additional target object in the input image.
[0072] In some implementations, determining the motion pattern of the reference objects includes: determining a first reference motion feature based on a first frame and a subsequent second frame of the reference video, the first reference motion feature being used to characterize the motion from the first frame to the second frame; determining a first set of reference objects in the first frame by generating a first set of semantic segmentation masks, the first set of semantic segmentation masks indicating the corresponding positions of the first set of reference objects in the first frame; and performing a partial convolution between the first reference motion feature and the first set of semantic segmentation masks to determine the motion pattern of the first set of reference objects.
[0073] In some implementations, generating the output video includes: determining at least one object in the input image by generating at least one semantic segmentation mask of the input image, the at least one object including a target object, the at least one semantic segmentation mask indicating the corresponding position of the at least one object in the input image; determining a combined motion pattern for the input image based on the corresponding motion pattern of the at least one reference object and the at least one semantic segmentation mask by determining a semantic mapping from at least one reference object in a first set of reference objects to the at least one object; generating a first predicted motion feature for the input image using a convolutional neural network based on the combined motion pattern and the input image; and performing an image transformation on the input image using the first predicted motion feature to generate a first output frame in the output video located after the starting frame.
[0074] In some implementations, generating the output video further includes: determining second reference motion features based on a second frame and a subsequent third frame of the reference video, the second reference motion features being used to characterize motion from the second frame to the third frame; generating a second set of semantic segmentation masks for the second frame, the second set of semantic segmentation masks indicating the corresponding positions of the first set of reference objects in the second frame; performing partial convolution on the second reference motion features and the second set of semantic segmentation masks to determine a second motion pattern of the first set of reference objects; and generating a second output frame in the output video located after the first output frame, the motion of the target object from the first output frame to the second output frame having the corresponding second motion pattern of the reference objects in the first set of reference objects.
[0075] As can be seen from the above description, the image animation scheme implemented according to this disclosure can intuitively apply the motion pattern of the reference object in the reference video to the input image to generate the output video, and the motion of the target object in the output video has the motion pattern of the reference object.
[0076] Figure 5 A flowchart of a method 500 for training an image animation model according to some implementations of the present disclosure is shown. Method 500 can be implemented by computing device 100, for example, it can be implemented at the image animation module 122 in the memory 120 of computing device 100.
[0077] like Figure 5 As shown, at box 510, computing device 100 acquires a training video. The training video includes a first training frame and a subsequent second training frame. At box 520, computing device 100 uses a machine learning model to generate a prediction video for the training video. The motion of objects in the prediction video follows the motion pattern of the objects in the training video, and the prediction video includes the first training frame and subsequent prediction frames for the second training frame. At box 530, computing device 100 trains the machine learning model based at least on the prediction frames and the second training frame.
[0078] In some implementations, generating a prediction video includes: determining the corresponding motion patterns of a set of training objects in the first training frame based on a first training frame and a second training frame; and generating prediction frames, wherein the motion of the set of training objects from the first training frame to the prediction frame has the corresponding motion patterns.
[0079] In some implementations, determining the corresponding motion patterns of a set of training objects includes: determining training motion features based on a first training frame and a second training frame, the training motion features being used to characterize the motion from the first training frame to the second training frame; determining a set of training objects in the first training frame by generating a set of training semantic segmentation masks, the set of training semantic segmentation masks indicating the corresponding positions of the set of training objects in the first training frame; performing the same spatial transformation on the training motion features and the set of training semantic segmentation masks respectively; and performing partial convolution on the transformed training motion features and the transformed set of training semantic segmentation masks to determine the corresponding motion patterns of the set of training objects.
[0080] In some implementations, generating a prediction frame includes: determining a combined motion pattern for a first training frame based on the corresponding motion patterns of a set of training objects and a set of training semantic segmentation masks; generating predicted motion features for the first training frame using a convolutional neural network based on the combined motion pattern and the first training frame; and performing image transformation on the first training frame using the predicted motion features to generate the prediction frame.
[0081] In some implementations, training a machine learning model includes: determining at least one loss function for training the machine learning model; determining a target loss function by weighted summation of the at least one loss function; and training the machine learning model by minimizing the target loss function.
[0082] In some implementations, determining at least one loss function includes at least one of the following: determining a frame loss function based on the difference between the predicted frame and the second training frame; determining a motion feature loss function based on the difference between the predicted motion features and the training motion features; and determining a smoothness loss function based on the spatial distribution of the predicted motion features and the spatial distribution of the predicted frame.
[0083] As can be seen from the above description, the scheme for training image animation models according to the present disclosure can use training videos to supervise the training of machine learning models without additional annotation work.
[0084] The following are some example implementations of this disclosure.
[0085] In a first aspect, this disclosure provides a computer-implemented method. The method includes: acquiring an input image and a reference video; determining a motion pattern of a reference object in the reference video based on the reference video; and generating an output video with the input image as a starting frame, wherein the motion of a target object in the output video exhibits the motion pattern of the reference object in the reference video.
[0086] In some implementations, generating the output video involves transferring the motion patterns of the reference object to the target object based on a semantic mapping from the reference object to the target object.
[0087] In some implementations, generating an output video includes at least one of the following: transferring the motion pattern of a reference object to a target object based on a predetermined rule indicating the mapping from a reference object to a target object; and transferring the motion pattern of an additional reference object to an additional target object based on the mapping from an additional reference object in the additional reference video to an additional target object in the input image.
[0088] In some implementations, determining the motion pattern of the reference objects includes: determining a first reference motion feature based on a first frame and a subsequent second frame of the reference video, the first reference motion feature being used to characterize the motion from the first frame to the second frame; determining a first set of reference objects in the first frame by generating a first set of semantic segmentation masks, the first set of semantic segmentation masks indicating the corresponding positions of the first set of reference objects in the first frame; and performing a partial convolution between the first reference motion feature and the first set of semantic segmentation masks to determine the motion pattern of the first set of reference objects.
[0089] In some implementations, generating the output video includes: determining at least one object in the input image by generating at least one semantic segmentation mask of the input image, the at least one object including a target object, the at least one semantic segmentation mask indicating the corresponding position of the at least one object in the input image; determining a combined motion pattern for the input image based on the corresponding motion pattern of the at least one reference object and the at least one semantic segmentation mask by determining a semantic mapping from at least one reference object in a first set of reference objects to the at least one object; generating a first predicted motion feature for the input image using a convolutional neural network based on the combined motion pattern and the input image; and performing an image transformation on the input image using the first predicted motion feature to generate a first output frame in the output video located after the starting frame.
[0090] In some implementations, generating the output video further includes: determining second reference motion features based on a second frame and a subsequent third frame of the reference video, the second reference motion features being used to characterize motion from the second frame to the third frame; generating a second set of semantic segmentation masks for the second frame, the second set of semantic segmentation masks indicating the corresponding positions of the first set of reference objects in the second frame; performing partial convolution on the second reference motion features and the second set of semantic segmentation masks to determine a second motion pattern of the first set of reference objects; and generating a second output frame in the output video located after the first output frame, the motion of the target object from the first output frame to the second output frame having the corresponding second motion pattern of the reference objects in the first set of reference objects.
[0091] In a second aspect, this disclosure also provides a computer-implemented method. The method includes: acquiring a training video. The training video includes a first training frame and subsequent second training frames. Generating a prediction video for the training video using a machine learning model. The motion of an object in the prediction video has a motion pattern of the object in the training video, and the prediction video includes the first training frame and subsequent prediction frames for the second training frame. Training the machine learning model based at least on the prediction frames and the second training frame.
[0092] In some implementations, generating a prediction video includes: determining the corresponding motion patterns of a set of training objects in the first training frame based on a first training frame and a second training frame; and generating prediction frames, wherein the motion of the set of training objects from the first training frame to the prediction frame has the corresponding motion patterns.
[0093] In some implementations, determining the corresponding motion patterns of a set of training objects includes: determining training motion features based on a first training frame and a second training frame, the training motion features being used to characterize the motion from the first training frame to the second training frame; determining a set of training objects in the first training frame by generating a set of training semantic segmentation masks, the set of training semantic segmentation masks indicating the corresponding positions of the set of training objects in the first training frame; performing the same spatial transformation on the training motion features and the set of training semantic segmentation masks respectively; and performing partial convolution on the transformed training motion features and the transformed set of training semantic segmentation masks to determine the corresponding motion patterns of the set of training objects.
[0094] In some implementations, generating a prediction frame includes: determining a combined motion pattern for a first training frame based on the corresponding motion patterns of a set of training objects and a set of training semantic segmentation masks; generating predicted motion features for the first training frame using a convolutional neural network based on the combined motion pattern and the first training frame; and performing image transformation on the first training frame using the predicted motion features to generate the prediction frame.
[0095] In some implementations, training the machine learning model includes: determining at least one loss function for training the machine learning model; determining a target loss function by weighted summation of the at least one loss function; and training the machine learning model by minimizing the target loss function.
[0096] In some implementations, determining the at least one loss function includes at least one of the following: determining a frame loss function based on the difference between the predicted frame and the second training frame; determining a motion feature loss function based on the difference between the predicted motion features and the training motion features; and determining a smoothness loss function based on the spatial distribution of the predicted motion features and the spatial distribution of the predicted frame.
[0097] In a third aspect, this disclosure provides an electronic device. The electronic device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform actions, the actions including: acquiring an input image and a reference video; determining a motion pattern of a reference object in the reference video based on the reference video; and generating an output video with the input image as a starting frame, the motion of a target object in the input image in the output video having the motion pattern of the reference object in the reference video.
[0098] In some implementations, generating the output video involves transferring the motion patterns of the reference object to the target object based on a semantic mapping from the reference object to the target object.
[0099] In some implementations, generating an output video includes at least one of the following: transferring the motion pattern of a reference object to a target object based on a predetermined rule indicating the mapping from a reference object to a target object; and transferring the motion pattern of an additional reference object to an additional target object based on the mapping from an additional reference object in the additional reference video to an additional target object in the input image.
[0100] In some implementations, determining the motion pattern of the reference objects includes: determining a first reference motion feature based on a first frame and a subsequent second frame of the reference video, the first reference motion feature being used to characterize the motion from the first frame to the second frame; determining a first set of reference objects in the first frame by generating a first set of semantic segmentation masks, the first set of semantic segmentation masks indicating the corresponding positions of the first set of reference objects in the first frame; and performing a partial convolution between the first reference motion feature and the first set of semantic segmentation masks to determine the motion pattern of the first set of reference objects.
[0101] In some implementations, generating the output video includes: determining at least one object in the input image by generating at least one semantic segmentation mask of the input image, the at least one object including a target object, the at least one semantic segmentation mask indicating the corresponding position of the at least one object in the input image; determining a combined motion pattern for the input image based on the corresponding motion pattern of the at least one reference object and the at least one semantic segmentation mask by determining a semantic mapping from at least one reference object in a first set of reference objects to the at least one object; generating a first predicted motion feature for the input image using a convolutional neural network based on the combined motion pattern and the input image; and performing an image transformation on the input image using the first predicted motion feature to generate a first output frame in the output video located after the starting frame.
[0102] In some implementations, generating the output video further includes: determining second reference motion features based on a second frame and a subsequent third frame of the reference video, the second reference motion features being used to characterize motion from the second frame to the third frame; generating a second set of semantic segmentation masks for the second frame, the second set of semantic segmentation masks indicating the corresponding positions of the first set of reference objects in the second frame; performing partial convolution on the second reference motion features and the second set of semantic segmentation masks to determine a second motion pattern of the first set of reference objects; and generating a second output frame in the output video located after the first output frame, the motion of the target object from the first output frame to the second output frame having the corresponding second motion pattern of the reference objects in the first set of reference objects.
[0103] In a fourth aspect, this disclosure provides an electronic device. The electronic device includes: a processing unit; and a memory coupled to the processing unit and containing instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform actions, the actions including: acquiring a training video, the training video including a first training frame and a subsequent second training frame; generating a prediction video for the training video using a machine learning model, the prediction video having motion patterns of objects in the prediction video having motion patterns of the objects in the training video, and the prediction video including the first training frame and subsequent prediction frames for the second training frame; and training the machine learning model based at least on the prediction frame and the second training frame.
[0104] In a fifth aspect, this disclosure provides a computer program product tangibly stored in a non-transient computer storage medium and including machine-executable instructions that, when executed by a device, cause the device to perform the methods described in the first or second aspect above.
[0105] In a sixth aspect, this disclosure provides a computer program product including machine-executable instructions that, when executed by a device, cause the device to perform the methods described in the first or second aspect above.
[0106] In a seventh aspect, this disclosure provides a computer-readable medium having stored thereon machine-executable instructions that, when executed by a device, cause the device to perform the methods described in the first or second aspect above.
[0107] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload programmable logic devices (CPLDs), and so on.
[0108] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0109] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0110] Furthermore, although the operations are described in a specific order, this should be understood as requiring that such operations be performed in the specific order shown or in sequential order, or requiring that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of a single implementation may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.
[0111] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A computer-implemented method, comprising: Acquire an input image and a reference video, wherein the input image includes multiple target objects and the reference video includes multiple reference objects corresponding to the multiple target objects; Based on the reference video, determine the motion patterns of the plurality of reference objects in the reference video; as well as An output video is generated with the input image as the starting frame, and the motion of the target object among the plurality of target objects in the output video has the motion pattern of the reference object corresponding to the target object; Determining the motion pattern of the reference object includes: Based on the first frame and the subsequent second frame of the reference video, a first reference motion feature is determined, which is used to characterize the motion from the first frame to the second frame. By generating a first set of semantic segmentation masks for the first frame, a first set of reference objects in the first frame is determined, and the first set of semantic segmentation masks indicates the corresponding positions of the first set of reference objects in the first frame; Partial convolution is performed on the first reference motion features and the first set of semantic segmentation masks to determine the motion pattern of the first set of reference objects.
2. The method according to claim 1, wherein generating the output video comprises: Based on the semantic mapping from the reference object to the target object, the motion pattern of the reference object is transferred to the target object.
3. The method of claim 1, wherein generating the output video comprises at least one of the following: Based on predetermined rules indicating the mapping from the reference object to the target object, the motion pattern of the reference object is transferred to the target object; and The motion pattern of the additional reference object is transferred to the additional target object in the input image based on the mapping of the additional reference object to the additional reference object based on the additional reference video.
4. The method according to claim 1, wherein generating the output video comprises: By generating at least one semantic segmentation mask of the input image, at least one object in the input image is determined, the at least one object including the target object, and the at least one semantic segmentation mask indicates the corresponding position of the at least one object in the input image; By determining the semantic mapping from at least one reference object in the first group of reference objects to the at least one object, a combined motion pattern for the input image is determined based on the corresponding motion pattern of the at least one reference object and the at least one semantic segmentation mask; Based on the combined motion pattern and the input image, a first predicted motion feature for the input image is generated using a convolutional neural network; as well as The input image is transformed using the first predicted motion feature to generate the first output frame in the output video, which is located after the starting frame.
5. The method according to claim 4, wherein generating the output video further comprises: Based on the second frame and the subsequent third frame of the reference video, a second reference motion feature is determined, which is used to characterize the motion from the second frame to the third frame; Generate a second set of semantic segmentation masks for the second frame, the second set of semantic segmentation masks indicating the corresponding positions of the first set of reference objects in the second frame; The second reference motion features are partially convolved with the second set of semantic segmentation masks to determine the second motion pattern of the first set of reference objects. as well as A second output frame is generated in the output video following the first output frame, and the motion of the target object from the first output frame to the second output frame has a corresponding second motion pattern of the reference object in the first group of reference objects.
6. A computer-implemented method, comprising: Acquire a training video, which includes a first training frame and a subsequent second training frame; Based on the first training frame and the second training frame, determine the corresponding motion pattern of each training object in a set of training objects in the first training frame. Based on the corresponding motion pattern, a predicted video is generated for the training video using a machine learning model, wherein the motion of an object in the predicted video has the motion pattern of the object in the training video, and the predicted video includes the first training frame and subsequent predicted frames for the second training frame. as well as The machine learning model is trained based at least on the predicted frame and the second training frame; Determining the corresponding motion patterns of the set of training subjects includes: Based on the first training frame and the second training frame, training motion features are determined, which are used to characterize the motion from the first training frame to the second training frame. By generating a set of training semantic segmentation masks for the first training frame, the set of training objects in the first training frame is determined, and the set of training semantic segmentation masks indicates the corresponding position of the set of training objects in the first training frame. Perform the same spatial transformation on the training motion features and the set of training semantic segmentation masks respectively; Partial convolution is performed on the transformed training motion features and the transformed set of training semantic segmentation masks to determine the corresponding motion patterns of the set of training objects.
7. The method of claim 6, wherein generating the predicted video comprises: The prediction frame is generated, and the motion of the group of training objects from the first training frame to the prediction frame has the corresponding motion pattern.
8. The method of claim 6, wherein generating the prediction frame comprises: Based on the corresponding motion patterns of the set of training objects and the set of training semantic segmentation masks, a combined motion pattern for the first training frame is determined. Based on the combined motion pattern and the first training frame, a convolutional neural network is used to generate predicted motion features for the first training frame. as well as The predicted motion features are used to perform image transformation on the first training frame to generate the predicted frame.
9. The method of claim 8, wherein training the machine learning model comprises: Determine at least one loss function for training the machine learning model; The target loss function is determined by weighted summation of the at least one loss function; as well as The machine learning model is trained by minimizing the target loss function.
10. The method of claim 8, wherein determining the at least one loss function comprises at least one of the following: Based on the difference between the predicted frame and the second training frame, a frame loss function is determined; Based on the difference between the predicted motion features and the trained motion features, a motion feature loss function is determined; and Based on the spatial distribution of the predicted motion features and the spatial distribution of the predicted frames, a smoothness loss function is determined.
11. An electronic device, comprising: Processing unit; as well as A memory, coupled to the processing unit and containing instructions stored thereon, which, when executed by the processing unit, cause the electronic device to perform actions, including: Acquire an input image and a reference video, wherein the input image includes multiple target objects and the reference video includes multiple reference objects corresponding to the multiple target objects; Based on the reference video, determine the motion patterns of the plurality of reference objects in the reference video; and An output video is generated with the input image as the starting frame, and the motion of the target object among the plurality of target objects in the output video has the motion pattern of the reference object corresponding to the target object; Determining the motion pattern of the reference object includes: Based on the first frame and the subsequent second frame of the reference video, a first reference motion feature is determined, which is used to characterize the motion from the first frame to the second frame. By generating a first set of semantic segmentation masks for the first frame, a first set of reference objects in the first frame is determined, and the first set of semantic segmentation masks indicates the corresponding positions of the first set of reference objects in the first frame; Partial convolution is performed on the first reference motion features and the first set of semantic segmentation masks to determine the motion pattern of the first set of reference objects.
12. The electronic device of claim 11, wherein generating the output video comprises: Based on the semantic mapping from the reference object to the target object, the motion pattern of the reference object is transferred to the target object.
13. The electronic device of claim 12, wherein generating the output video comprises at least one of the following: Based on predetermined rules indicating the mapping from the reference object to the target object, the motion pattern of the reference object is transferred to the target object; and The motion pattern of the additional reference object is transferred to the additional target object in the input image based on the mapping of the additional reference object to the additional reference object based on the additional reference video.
14. The electronic device of claim 11, wherein generating the output video comprises: By generating at least one semantic segmentation mask of the input image, at least one object in the input image is determined, the at least one object including the target object, and the at least one semantic segmentation mask indicates the corresponding position of the at least one object in the input image; By determining the semantic mapping from at least one reference object in the first group of reference objects to the at least one object, a combined motion pattern for the input image is determined based on the corresponding motion pattern of the at least one reference object and the at least one semantic segmentation mask; Based on the combined motion pattern and the input image, a first predicted motion feature for the input image is generated using a convolutional neural network; as well as The input image is transformed using the first predicted motion feature to generate the first output frame in the output video, which is located after the starting frame.
15. An electronic device comprising: Processing unit; as well as A memory, coupled to the processing unit and containing instructions stored thereon, which, when executed by the processing unit, cause the electronic device to perform actions, including: Acquire a training video, which includes a first training frame and a subsequent second training frame; Based on the first training frame and the second training frame, determine the corresponding motion pattern of each training object in a set of training objects in the first training frame. Based on the corresponding motion patterns, a predicted video is generated for the training video using a machine learning model, wherein the motion of objects in the predicted video has the motion patterns of the objects in the training video, and the predicted video includes the first training frame and subsequent predicted frames for the second training frame; and The machine learning model is trained based at least on the predicted frame and the second training frame; Determining the corresponding motion patterns of the set of training subjects includes: Based on the first training frame and the second training frame, training motion features are determined, which are used to characterize the motion from the first training frame to the second training frame. By generating a set of training semantic segmentation masks for the first training frame, the set of training objects in the first training frame is determined, and the set of training semantic segmentation masks indicates the corresponding position of the set of training objects in the first training frame. Perform the same spatial transformation on the training motion features and the set of training semantic segmentation masks respectively; Partial convolution is performed on the transformed training motion features and the transformed set of training semantic segmentation masks to determine the corresponding motion patterns of the set of training objects.
16. The electronic device of claim 15, wherein generating the predicted video comprises: The prediction frame is generated, and the motion of the group of training objects from the first training frame to the prediction frame has the corresponding motion pattern.
17. The electronic device of claim 15, wherein generating the prediction frame comprises: Based on the corresponding motion patterns of the set of training objects and the set of training semantic segmentation masks, a combined motion pattern for the first training frame is determined. Based on the combined motion pattern and the first training frame, a convolutional neural network is used to generate predicted motion features for the first training frame. as well as The predicted motion features are used to perform image transformation on the first training frame to generate the predicted frame.
18. The electronic device of claim 17, wherein training the machine learning model comprises: Determine at least one loss function for training the machine learning model; The target loss function is determined by weighted summation of the at least one loss function; as well as The machine learning model is trained by minimizing the target loss function.
19. A computer program product comprising machine-executable instructions that, when executed by a device, cause the device to perform an action, the action comprising: Acquire an input image and a reference video, wherein the input image includes multiple target objects and the reference video includes multiple reference objects corresponding to the multiple target objects; Based on the reference video, determine the motion patterns of the plurality of reference objects in the reference video; as well as An output video is generated with the input image as the starting frame, and the motion of the target object among the plurality of target objects in the output video has the motion pattern of the reference object corresponding to the target object; Determining the motion pattern of the reference object includes: Based on the first frame and the subsequent second frame of the reference video, a first reference motion feature is determined, which is used to characterize the motion from the first frame to the second frame. By generating a first set of semantic segmentation masks for the first frame, a first set of reference objects in the first frame is determined, and the first set of semantic segmentation masks indicates the corresponding positions of the first set of reference objects in the first frame; Partial convolution is performed on the first reference motion features and the first set of semantic segmentation masks to determine the motion pattern of the first set of reference objects.
20. A computer program product comprising machine-executable instructions that, when executed by a device, cause the device to perform an action, the action comprising: Acquire a training video, which includes a first training frame and a subsequent second training frame; Based on the first training frame and the second training frame, determine the corresponding motion pattern of each training object in a set of training objects in the first training frame; based on the corresponding motion pattern, use a machine learning model to generate a predicted video for the training video, wherein the motion of the objects in the predicted video has the motion pattern of the objects in the training video, and the predicted video includes the first training frame and subsequent predicted frames for the second training frame. as well as The machine learning model is trained based at least on the predicted frame and the second training frame; Determining the corresponding motion patterns of the set of training subjects includes: Based on the first training frame and the second training frame, training motion features are determined, which are used to characterize the motion from the first training frame to the second training frame. By generating a set of training semantic segmentation masks for the first training frame, the set of training objects in the first training frame is determined, and the set of training semantic segmentation masks indicates the corresponding position of the set of training objects in the first training frame. Perform the same spatial transformation on the training motion features and the set of training semantic segmentation masks respectively; Partial convolution is performed on the transformed training motion features and the transformed set of training semantic segmentation masks to determine the corresponding motion patterns of the set of training objects.
Citation Information
Patent Citations
Video generation method for action migration and neural network training method and device
CN110210386A
Moving target pose tracking method and device
CN111192293A