Method, apparatus and computer program product for processing images

By generating and fusing three-dimensional models from source object images and animation models from target object videos, the method addresses the challenge of low-quality video synthesis, achieving accurate and efficient replacement of target objects with source objects in videos.

CN120318376APending Publication Date: 2025-07-15DELL PROD LP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410051817.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the video synthesis, the direct construction of an animation model requires high visual information of the source object, resulting in the inability to generate high-quality synthetic videos and high computing costs.

Method used

By generating the three-dimensional model of the source object and the animation model of the target object and fusing it, a second video that replaces the target object is generated, the static characteristics of the three-dimensional model are used to reduce the visual information needs, and combining the action information of the animation model to achieve high-quality video replacement.

Benefits of technology

It reduces the requirements for visual information, improves the quality and computing efficiency of video replacement, and reduces computing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318376A_ABST
    Figure CN120318376A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a device and a computer program product for processing an image. The method comprises the steps that multiple images and a first video are obtained, the multiple images indicate visual information of a source object under multiple visual angles, and the first video indicates animation of a target object. The method further includes generating a three-dimensional model for the source object based on the plurality of images, and generating a plurality of animation models for the target object based on the first video. The method further includes generating a second video for the source object by fusing the three-dimensional model for the source object and the plurality of animation models for the target object, where the second video replaces the target object in the first video with the source object. Through the mode, the source object can be replaced to the position of the target object in the video only by aiming at the visual information of the source object under the multiple visual angles, and the source object accurately simulates the action of the target object in the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and more particularly, to methods, devices, and computer program products for processing images. Background Art

[0002] With the rapid development of computer technology, video data has grown explosively, and a large number of different types of videos fill the cyberspace. People can freely download and share this video data and watch the content played therein.

[0003] Video processing technology allows users to edit video data, such as adjusting the clarity or playing speed of a video, or splicing different videos into one video, etc. Video synthesis technology is a common video processing technology, and people can synthesize elements other than video data into video data and display such synthesized elements when the video is played, so as to edit the desired video data. Summary of the Invention

[0004] Embodiments of the present disclosure propose a method, a device, and a computer program product for processing images. In a first aspect of the embodiments of the present disclosure, a method for processing images is provided. The method includes obtaining a plurality of images and a first video, where the plurality of images indicate visual information of a source object from multiple perspectives, and the first video indicates the animation of a target object; generating a three-dimensional model for the source object based on the plurality of images; generating a plurality of animation models for the target object based on the first video; and fusing the three-dimensional model for the source object and the plurality of animation models for the target object to generate a second video for the source object, where the second video uses the source object to replace the target object in the first video.

[0005] In a second aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement a method for processing images, the method including obtaining a plurality of images and a first video, where the plurality of images indicate visual information of a source object from multiple perspectives, and the first video indicates the animation of a target object; generating a three-dimensional model for the source object based on the plurality of images; generating a plurality of animation models for the target object based on the first video; and fusing the three-dimensional model for the source object and the plurality of animation models for the target object to generate a second video for the source object, where the second video uses the source object to replace the target object in the first video.

[0006] In a third aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, a method for processing images is implemented. The method includes obtaining a plurality of images and a first video, where the plurality of images indicate visual information of a source object from multiple perspectives, and the first video indicates an animation of a target object; generating a three-dimensional model for the source object based on the plurality of images; generating a plurality of animation models for the target object based on the first video; and fusing the three-dimensional model for the source object and the plurality of animation models for the target object to generate a second video for the source object, where the second video uses the source object to replace the target object in the first video.

[0007] It should be understood that the content described in the summary of the invention section is not intended to limit the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] With reference to the accompanying drawings and the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 is a schematic diagram of an exemplary environment in which the embodiments of the present disclosure can be implemented;

[0010] Figure 2 is a flowchart of a method for processing images according to an embodiment of the present disclosure;

[0011] Figure 3 is a schematic diagram of processing images according to an embodiment of the present disclosure;

[0012] Figure 4 is a flowchart of generating a three-dimensional model according to an embodiment of the present disclosure;

[0013] Figure 5 is a flowchart of generating a plurality of animation models for a target object according to an embodiment of the present disclosure;

[0014] Figure 6 is a flowchart of fusing a three-dimensional model and a plurality of animation models according to an embodiment of the present disclosure;

[0015] Figure 7 is a logical schematic diagram of a method for processing images according to an embodiment of the present disclosure;

[0016] Figure 8 is a schematic diagram of a method for processing images according to an embodiment of the present disclosure

[0017] Figure 9It is a logical block diagram for generating a second video according to an embodiment of the present disclosure;

[0018] Figure 10 It is a schematic diagram of the effect according to an embodiment of the present disclosure; and

[0019] Figure 11 It shows a block diagram of a device that can implement multiple embodiments of the present disclosure. Detailed implementation manners

[0020] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0021] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0022] In the related art of video synthesis, it is usually to directly construct an animation model for the source object using the visual information of the source object, and then control the form of the animation model to generate a series of animations. However, if an animation model is directly constructed, the requirements for the visual information of the source object will be very high, and the existing materials (such as the number of images for the source object, the resolution, or the video stream for the source object) are often insufficient to construct a relatively good animation model for the source object, thus a high-quality synthesized video cannot be obtained.

[0023] To this end, the present disclosure proposes a method for processing images. In an embodiment of the present disclosure, a plurality of images and a first video are obtained, where the plurality of images indicate visual information of a source object from multiple perspectives, and the first video indicates an animation of a target object; a three-dimensional model for the source object is generated based on the plurality of images; a plurality of animation models for the target object are generated based on the first video; and the three-dimensional model for the source object and the plurality of animation models for the target object are fused to generate a second video for the source object, where the second video replaces the target object in the first video with the source object. Since the static three-dimensional model can express the three-dimensional visual information of the source object more clearly and accurately than the dynamic model, and the plurality of animation models also accurately express the respective action information of the target object in the first video from a three-dimensional perspective, fusing these two models can achieve the purpose of replacing the target object in the first video with the source object, and enable the replaced second video to clearly and accurately reproduce the scenario where the source object simulates the respective actions of the target object. In addition, since the three-dimensional model for the source object is a static model, its requirement for the visual information of the source object is significantly lower than that of the animation model. Therefore, by fusing the three-dimensional model and the animation model to indirectly establish an animation model for the source object, the method of the present disclosure can use lower-quality visual information to construct a higher-quality animation model and the second video with the target object replaced, while also reducing the computational cost.

[0024] Figure 1 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As Figure 1 shown, the environment 100 may include a client 101, a network 102, and a service unit 103, and the service unit 103 is communicatively coupled to the client 101 via the network 102. The network 102 may be, for example, a wide area network (WAN), a local area network (LAN), a wireless network, a public telephone network, an intranet, or any other type of network well known to those skilled in the art.

[0025] In this embodiment, the method for processing images is executed by the service unit 103. The method executed by the service unit 103 includes the following steps. The service unit 103 obtains a plurality of images and a first video from the client 101, where the plurality of images indicate visual information of a source object from multiple perspectives, and the first video indicates an animation of a target object. It should be noted that the plurality of images may be a plurality of images captured by, for example, a camera, or may be a plurality of video frames constituting video data. The first video may come from the client 101 or may be video data stored locally in the service unit 103.

[0026] The service unit 103 generates a three-dimensional model for the source object based on multiple images. The service unit 103 can use a local model or call an external model (such as a neural network for image processing) to generate a three-dimensional model for the source object. The three-dimensional model for the source object generated by the service unit 103 is a static model, which can richly display the visual information of the static source object. For example, when the service unit 103 is externally connected to a monitor, the source object can be observed from various perspectives of the three-dimensional view. Since there is no need to reflect the motion information of the model, the service unit 103 can construct a realistic three-dimensional model with a relatively low computational cost, reducing the requirements for the hardware configuration of the service unit 103.

[0027] The service unit 103 generates multiple animation models for the target object based on the first video. Each of the multiple animation models reflects the respective actions made by the target object in the first video. In this animation model, there is no need for the service unit 103 to render the visual information of the model (such as color information, clothing information, decoration information), so the computational complexity is greatly reduced on the premise of ensuring the accurate expression of the actions of the target object.

[0028] The service unit 103 fuses the three-dimensional model for the source object and the multiple animation models for the target object to generate a second video for the source object, where the second video uses the source object to replace the target object in the first video. The service unit 103 can externally connect to a monitor (not shown) to play this second video, or can transmit the second video to the client 101 through the network 102 for playing.

[0029] As Figure 1 shown, in the environment 100, the network 102 can be used to transmit data between the client 101 and the service unit 103. For example, the network 102 can be used to transmit multiple images from the client 101 to the service unit 103, so that the service unit 103 can process the multiple images according to the embodiments of the present disclosure. The network 102 has a theoretical bandwidth, which refers to the maximum transmission speed supported by the network 102, indicating the maximum amount of data that the network 102 can transmit under ideal conditions, usually measured in bits per second (bps). For example, if the theoretical bandwidth of the network 102 is 100 Mbps, it means that under ideal conditions, it can transmit one hundred million bits of data per second. However, in practice, due to possible other factors in the network (such as signal interference, bandwidth sharing, transmission delay, etc.), the actual transmission speed may not reach 100 Mbps.

[0030] As understood by those of ordinary skill in the art, examples of the service unit 103 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The servers can be directly or indirectly connected through wired or wireless communication methods, which are not limited in this application.

[0031] The client 101 can be any type of mobile computing device, including mobile computers (e.g., personal digital assistants (PDAs), laptop computers, notebook computers, tablet computers, netbooks, etc.), mobile phones (e.g., cellular phones, smartphones, etc.), wearable computing devices (e.g., smartwatches, head-mounted devices, including smart glasses, etc.) or other types of mobile devices. In some embodiments, the client 101 can also be a stationary computing device, such as a desktop computer, a gaming console, a smart TV, etc.

[0032] Figure 2 is a flowchart of a method for processing images according to some embodiments of the present disclosure. As Figure 2 shown, the flowchart 200 includes blocks 210 - 270. At block 210, a plurality of images and a first video are obtained, where the plurality of images indicate the visual information of the source object from multiple perspectives, and the first video indicates the animation of the target object.

[0033] The plurality of images can be images captured using, for example, a camera device, or several video frames in a sequence of video frames constituting video data. The content of the plurality of images can include the visual information of the source object from multiple perspectives. For example, the plurality of images can include the visual information of the front image, side image, and back image of the source object. In some embodiments, in each of the plurality of images, the visual information of the source object occupies the main part of the entire image, and the posture, shape, etc. are consistent in each image. For example, a camera can be used to capture the source object in the same static state from multiple perspectives. The first video can be a video in any format that reflects a series of actions made by the target object, and these actions include static actions. The source object is the replacement object, and the target object is the object to be replaced. Since one purpose of the present disclosure is to replace the target object in the first video with the source object, it is necessary to obtain images reflecting the source object and the first video reflecting the target object.

[0034] At block 230, a three-dimensional model of the source object is generated based on multiple images. The three-dimensional model reflects the model of the source object constructed in a three-dimensional space, in which visual information of the source object can be obtained from any perspective. The visual information reflected in this three-dimensional space includes both the visual information included in the multiple images and the visual information not included in the multiple images, and the visual information not included can be predicted based on the included visual information. Therefore, the three-dimensional model can display all the visual information of the source object in all directions, including but not limited to color, clothing, decorations, etc. The three-dimensional model is a static model and no specific action information needs to be determined for it.

[0035] At block 250, multiple animation models of the target object are generated based on the first video. The multiple animation models at least include a series of actions of the target object in the first video. For example, if the first video captures the content of the target object playing football, then the multiple animation models indicate the respective actions of the target object when playing football in the first video, and each animation model corresponds to one action. It should be noted that the animation model may not reflect the visual information of the target object, that is, there is no need to depict the visual information such as the color and clothing of the target object in the animation model, but it is necessary to include pose and shape information because these information are closely related to the actions.

[0036] At block 270, the three-dimensional model of the source object and the multiple animation models of the target object are fused to generate a second video of the source object, where the second video uses the source object to replace the target object in the first video. In order to replace the target object in the first video with the source object to generate the second video, since the three-dimensional model does not contain action information while the animation model indicates action information, it is necessary to use the three-dimensional model of the source object to simulate the animation model of the target object, that is, imitate a series of actions of the target object in the first video, so as to obtain the second video in which the source object moves.

[0037] Because the static three-dimensional model can express the three-dimensional visual information of the source object more clearly and accurately compared to the dynamic model, and the multiple animation models also accurately express the respective action information of the target object in the first video from a three-dimensional perspective, so fusing these two models can achieve the purpose of replacing the target object in the first video with the source object, and make the replaced second video clearly and accurately reproduce the scenario of the source object simulating the respective actions of the target object. In addition, since the three-dimensional model of the source object is a static model, its requirement for the visual information of the source object is significantly lower than that of the animation model. Therefore, by fusing the three-dimensional model and the animation model to indirectly establish an animation model of the source object, the method of the present disclosure can use lower-quality visual information to construct a higher-quality animation model and the second video with the target object replaced, while also reducing the computational cost.

[0038] Figure 3 is a schematic diagram of processing an image according to an embodiment of the present disclosure. In Figure 3 the illustrated embodiment, multiple images are from video data. As Figure 3 shown, a shooting device 301 is used to shoot a source object 302, thereby obtaining multiple images 303 including the source object; a shooting device 304 is used to shoot a target object 305, thereby obtaining a first video 306 including the target object. The multiple images 303 and the first video 306 are input into a computing device 307 capable of implementing the method of the present disclosure in real time, and a three-dimensional model 308 of the source object and multiple animation models (not shown) of the target object can be obtained. From Figure 3 it can be seen that the three-dimensional model 308 contains rich visual information of the source object. The computing device 307 can fuse the three-dimensional model with the multiple animation models, thereby obtaining a second video 309 of the source object.

[0039] Regarding block 230, the present disclosure provides some more specific embodiments for generating a three-dimensional model. Figure 4 is a flowchart of generating a three-dimensional model according to an embodiment of the present disclosure. The flowchart 400 includes blocks 410-440. In block 410, multiple images are sampled to obtain multiple image key frames. In order to reduce the amount of calculation and ensure the quality of the three-dimensional model, it is not necessary to utilize the visual information of each image, but image key frames can be selected for processing through multiple image samplings. This can also improve the processing efficiency.

[0040] In block 420, according to the multiple image key frames, a sparse point cloud for the target object is determined. A point cloud is a data structure used to represent the shape of an object in three-dimensional space. It consists of a series of points, and each point has a corresponding coordinate value indicating the position of the point in three-dimensional space. Each point can also have other information values, such as color information, so that the point cloud can represent a three-dimensional model with color. A sparse point cloud is a type of point cloud. Compared with a dense point cloud, a sparse point cloud uses a smaller number of points to display various visual information of the three-dimensional model. Generating a sparse point cloud requires less calculation, has lower requirements for hardware configuration, and has a faster generation speed.

[0041] In some embodiments, for any two image key frames among the multiple image key frames, first, feature points within each image key frame are determined, and then a fundamental matrix of the image key frame is determined. The fundamental matrix indicates the relative pose and depth information of the image key frame; then, according to the fundamental matrix, the feature points are matched. The matched feature points indicate the common information of the points belonging to each image key frame, and thus can be used as a point in the sparse point cloud. A sparse point cloud for the source object is generated according to the matched feature points.

[0042] At block 430, based on the sparse point cloud of the target object, the first model is trained according to preset training conditions. The sparse point cloud can serve as the basic information and reference for constructing a three-dimensional model. Through this sparse point cloud, a neural network for constructing a three-dimensional model, such as a neural radiance field network, can be trained. By iteratively training the first model, it can predict visual information not included in the multiple images based on the multiple images, thereby adjusting and enriching the structure and information of the point cloud to make it more clearly and accurately express the visual information of the source object.

[0043] In an embodiment, the preset training conditions include: iteratively training the first model a first preset number of times based on the sparse point cloud of the target object; and reducing the learning rate of the first model in response to having performed a second preset number of iterations, where the second preset number is less than the first preset number. For example, the first preset number can be set to 50,000 times, and the second preset number can be set to 5,000 times, 10,000 times, 15,000 times, 20,000 times, etc. Additionally, the learning rate of the first model can be set to 1 / (e^4). By adopting the training conditions of the foregoing embodiment, a first model with good accuracy can be obtained relatively quickly with limited computing resources.

[0044] At block 440, according to multiple image key frames, the trained first model is used to determine a three-dimensional model. After the iterative training is completed, the multiple image key frames are input into the trained first model, and the trained first model can construct a three-dimensional model for the source object based on the existing visual information and predicted visual information of the multiple image key frames.

[0045] Since a static three-dimensional model can express the three-dimensional visual information of the source object more clearly and accurately than a dynamic model, this provides a necessary basis for the second video to clearly and accurately reproduce the various actions of the source object simulating the target object.

[0046] For block 250, the present disclosure provides some specific embodiments. Figure 5 It is a flowchart for generating multiple animation models for a target object according to an embodiment of the present disclosure. The flowchart 500 includes blocks 510 - 540. At block 510, each video frame of the first video is sampled to obtain multiple video key frames. Similar to block 410, in order to reduce the computational amount and ensure the quality of the animation model, it is not necessary to utilize the visual information of each image, but rather video key frames can be selected for processing through multiple image samplings. This can also improve the processing efficiency.

[0047] At block 520, based on one of the multiple video key frames, a second model is used to generate an animation model for the target object. The animation model is a three-dimensional model that mainly reflects the action information of the target object in the one video key frame, and there is no need to render other information such as visual information like clothing and color on the animation model. At block 520, one video key frame is selected to generate the animation model, which can be the model basis for other animation models. In an embodiment, a Skinned Multi-Person Linear (SMPL) model can be used to generate an animation model for the target object.

[0048] At block 530, the deformation information of the target object is extracted from each of the multiple video key frames. In order to adjust the animation model obtained at block 520 to obtain the actions of the target object reflected in other video key frames, it is necessary to extract the deformation information of the target object from each video key as the main basis for adjusting the animation model.

[0049] At block 540, according to the deformation information of the target object in each video key frame, an animation model for the target object is adjusted to obtain the animation model of the target object in each video key frame. For example, if the animation model is generated based on the first video key frame, then when determining the animation model corresponding to the second video key frame, the deformation information of the target object can be extracted from this video key frame, the deformation information is converted into action parameters related to the animation model, and after applying the action parameters to the animation model, the animation model can present the actions of the target object in the second video key frame, thus obtaining the animation model corresponding to the second video key frame.

[0050] In this embodiment, since each of the video key frames is for the target object and the generation of the animation model mainly considers the action information rather than the visual information, one video key frame can be selected from the multiple video key frames to generate the animation model. The animation model can use parameters related to the action information (i.e., deformation information) to obtain animation models with other actions. Therefore, there is no need to generate an animation model for each video key frame, which can greatly save computing resources and improve computing efficiency.

[0051] In an embodiment, deformation information of a target object is extracted from each of a plurality of video key frames, including: extracting pose information and shape information of the target object from each of the plurality of video key frames; correspondingly, an animation model for the target object is adjusted according to the deformation information of the target object in each video key frame, including for each video key frame, using the pose information and shape information of the target object extracted from the video key frame to adjust an animation model for the target object to obtain an animation model of the target object in each video key frame. As an example, the pose information may indicate postures of the target object such as standing, semi-squatting, squatting, crawling, jumping, etc., and the shape information may indicate shapes of the target object such as height, fatness, body proportion, etc.

[0052] For box 270, the present disclosure also provides some more specific embodiments. Figure 6 is a flowchart of fusing a 3D model and a plurality of animation models according to an embodiment of the present disclosure. In Figure 6 the illustrated embodiment, flowchart 600 includes box 610 - box 640. In box 610, a typical animation model is determined from a plurality of animation models for the target object, where the typical animation model for the target object has a consistent pose and shape with the 3D model of the source object. The consistent pose of the typical animation model and the 3D model is beneficial to the alignment of the two. For example, the typical animation model and the 3D model may be models of the target object and the source object standing respectively.

[0053] In box 620, the 3D model is aligned with the typical animation model to obtain an aligned 3D model. Generally, both the 3D model and the animation model are represented by point clouds. At this time, this alignment does not require the complete coincidence of corresponding points in the point clouds, only that the overall point clouds are very close to each other.

[0054] In box 630, a plurality of skinning weights of the animation model for the target object are transmitted to the aligned 3D model. Since the 3D model is a static model without relevant parameters for adjusting actions, and the animation model has skinning weights as relevant parameters for adjusting actions. Also, since both can be implemented in the form of point clouds at the same time, the adjustment of the actions of the 3D model is achieved by borrowing the plurality of skinning weights of the plurality of animation models to make it a dynamic model. The skinning weight indicates the correlation degree between the feature points in the point cloud and the bones of the model. By setting the skinning weight, the corresponding action can be adjusted on the model.

[0055] At block 640, a second video for the source object is generated using the aligned three-dimensional model according to multiple skin weights. After obtaining the skin weights, the aligned three-dimensional model can be used as the object to be adjusted, and a series of actions can be made on the aligned three-dimensional model. These actions are consistent with the actions made by the above-mentioned multiple animation models. This series of actions are spliced into a video, which is the second video. In this embodiment, by passing the skin weights of the animation model to the three-dimensional model, the dynamic adjustment of the static three-dimensional model is realized, so that the source object accurately simulates a series of actions of the target object.

[0056] In an embodiment, aligning the three-dimensional model with a typical animation model includes: rotating the three-dimensional model by a first angle according to a rotation vector to obtain a first three-dimensional model, where the first angle is the angle indicated by the rotation vector; displacing the first three-dimensional model along a first direction by a first distance according to a displacement vector to obtain a second three-dimensional model, where the first direction and the first distance are the direction and distance indicated by the displacement vector respectively; determining the distance between the second three-dimensional model and the typical animation model; and in response to the distance being less than a preset threshold, determining the second three-dimensional model as the aligned three-dimensional model for the source object.

[0057] In this embodiment, the aligned three-dimensional model is represented by a point cloud, and the rotation and displacement of the three-dimensional model are completed by adjusting the coordinates of the points in the point cloud. Through rotation and displacement operations, the alignment of the aligned three-dimensional model and the typical animation model can be achieved, and this alignment method is simple and accurate.

[0058] For example, the alignment operation can be performed according to the following formula (1).

[0059] P’ = R * P + T Formula (1)

[0060] Where P’ represents the aligned three-dimensional model, P represents the three-dimensional model before alignment, R represents the rotation vector, and T represents the displacement vector.

[0061] In an embodiment, determining the distance between the second three-dimensional model and the typical animation model includes, for each point in the second three-dimensional model, determining the distance between the point and the corresponding point in the typical animation model as the first distance of the point; and determining the sum of the first distances of all points in the second three-dimensional model as the distance between the second three-dimensional model and the typical animation model. As shown above, when determining whether to align, the sum of the distances between corresponding points in two point clouds can be calculated. A larger sum indicates misalignment, and a smaller sum indicates that alignment can be recognized.

[0062] For example, the distance between the second three-dimensional model and the typical animation model can be determined according to the following formula (2).

[0063]

[0064] where S represents the distance between the second 3D model and the typical animation model, p i represents the coordinates of the i-th point in the second 3D model, N represents the total number of points in the second 3D model, i is a positive integer, R represents the rotation vector, T represents the displacement vector, m i represents the coordinates of the point corresponding to the i-th point in the typical animation model. This method helps to simply and accurately determine whether two point clouds are aligned.

[0065] In an embodiment, according to multiple skin weights, using the aligned 3D model to generate a second video for the source object, including controlling the aligned 3D model according to multiple skin weights to obtain a 3D animation for the source object; and determining a two-dimensional projection of the 3D animation for the source object according to a preset perspective as the second video. Since the obtained 3D scene can provide visual information from various perspectives, when performing the two-dimensional projection, it can be projected according to any preset perspective, so that second videos at different angles can be obtained.

[0066] As an example, control the aligned 3D model according to formula (3) to indirectly obtain multiple animation models for the source object.

[0067] t p (θ, β) = C + B s (β) + B p (θ) Formula (3)

[0068] where θ is the parameter representing the pose in the skin weight, β is the parameter representing the shape in the skin weight, T p (θ, β) is the aligned 3D model with pose θ and shape β, C is the initial mixed weight of the aligned 3D model, which indicates the initial parameters of the aligned 3D model, B s (β) is the action information generated by the pose θ, B p (θ) is the action information generated by the shape β.

[0069] Through the control method of formula (3), dynamic control can be performed on the aligned 3D model, so that the source object can simulate a series of actions of the target object. This avoids the generation work of a large number of models, can save computing resources and improve the speed of generating the second video.

[0070] In an embodiment, the method for processing an image further includes determining background information according to the first video, where the background information does not include the target object; and embedding the background information of the first video into the two-dimensional projection as an optimized second video for the source object. Embedding the background information into the two-dimensional projection can more realistically reproduce the scene in the first video in the second video and can improve the user experience.

[0071] Figure 7 is a logical schematic diagram of a method for processing images according to an embodiment of the present disclosure. In Figure 7 the illustrated embodiment, multiple image keyframes 701 are sampled from multiple images (not shown), and the multiple image keyframes 701 are input into a neural radiance field network 702 to generate a three-dimensional model 703 for the source object.

[0072] In this process, multiple video keyframes 704 can be sampled from a first video (not shown) in parallel, and the multiple video keyframes 704 are input into a skinned multi-person linear model 705 to generate an animation model 706 for the target object. Then, the three-dimensional model 703 and the animation model 706 are aligned to obtain an aligned three-dimensional model 707. Skinning weights 708 corresponding to each video keyframe 704 are extracted from the animation model 706, and the skinning weights 708 are transmitted to the aligned three-dimensional model 707 and two-dimensionally projected thereon to obtain a two-dimensional projection 709. The background 710 of each video keyframe 704 is extracted from the multiple video keyframes 704 and inserted into the two-dimensional projection 709 respectively to obtain a second video 711.

[0073] Figure 8 is a schematic diagram of a method for processing images according to an embodiment of the present disclosure. In this embodiment, multiple images 801 are sampled to extract image keyframes 802, where the source object stands in a T-shaped pose in the multiple image keyframes 802. Then, a sparse point cloud 803 for the source object is generated, and the sparse point cloud 803 includes a number of points, each point having coordinate information and color information. Using these points as training data, the neural radiance field network is iteratively trained to enable it to generate a three-dimensional model 804 based on the multiple image keyframes.

[0074] On the other hand, a first video 805 is sampled to obtain multiple video keyframes 806, and an animation model for the target object is generated via a skinned multi-person linear model. Here, only the first video keyframe can be used to generate the animation model 807. A model standing in a T-shaped pose is preset in the skinned multi-person linear model, and this model is used as the basis for multiple animation models. Furthermore, the pose information and shape information of each video keyframe can be extracted from the multiple video keyframes 806 to generate skinning weights, and the skinning weights are used to adjust the animation model 807 to obtain a series of actions of the target object.

[0075] The three-dimensional model 804 and the three-dimensional model 804 are fused through rotation and displacement vectors to obtain an aligned three-dimensional model 808. The background 809 obtained from the multiple video keyframes 806 is embedded into the aligned three-dimensional model 808, and combined with the skinning weights, a second video 810 can be obtained.

[0076] Figure 9 It is a logical block diagram for generating the second video 909 according to an embodiment of the present disclosure. At 901, the skin weights of multiple animation models are extracted. At 902, the three-dimensional model for the source object is loaded. At 903, the vertex information of the typical animation model is obtained. According to these vertex information, at 904, the point cloud of the typical animation model is determined. The point clouds of the three-dimensional model 902 and the typical animation model 904 are normalized, and then the two point clouds are aligned at 906. After alignment, the aligned three-dimensional model can be denormalized at 907, and the skin weights obtained at 901 are transmitted to the aligned three-dimensional model. By controlling the aligned three-dimensional model to perform a series of actions through the skin weights, the second video 909 can be generated.

[0077] Figure 10 It is a schematic diagram of the effect according to an embodiment of the present disclosure. As Figure 10 shown, the first column on the left is the first video 1001, which reflects a series of actions of the left object and the right object playing football. The second video generated according to an embodiment of the present disclosure can be the aligned three-dimensional model without background shown at 1002, where the left object in the first video 1001 is replaced with the source object as the target object. The second video with background is shown at 1003, and the difference from 1002 is that the background is embedded. From Figure 10 the effect shown, the source object can clearly and accurately simulate the actions of the target object in the background of the first video.

[0078] Figure 11 It shows a schematic block diagram of an example device 1100 that can be used to implement the embodiments of the present disclosure. As shown in the figure, the device 1100 includes a computing unit 1101, which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) 1102 or the computer program instructions loaded from the storage unit 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.

[0079] Multiple components in device 1100 are connected to I / O interface 1105, including: input unit 1106, such as a keyboard, mouse, etc.; output unit 1107, such as various types of displays, speakers, etc.; storage unit 1108, such as a disk, optical disc, etc.; and communication unit 1109, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1109 allows device 1100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0080] Computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 1101 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1101 executes the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by computing unit 1101, one or more steps of method 200 described above can be executed. Alternatively, in other embodiments, computing unit 1101 can be configured to execute method 200 in any other suitable manner (e.g., by means of firmware).

[0081] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0082] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0083] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. Additionally, although the operations are depicted in a particular order, this should be understood to require that the operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, the various features that are described in the context of a single implementation can also be implemented separately or in any suitable sub-combination in multiple implementations.

[0084] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for processing an image, comprising: Obtaining a plurality of images and a first video, wherein the plurality of images indicate visual information of a source object from multiple perspectives, and the first video indicates an animation of a target object; Generating a three-dimensional model for the source object based on the plurality of images; Generating a plurality of animation models for the target object based on the first video; And Fusing the three-dimensional model for the source object and the plurality of animation models for the target object to generate a second video for the source object, wherein the second video uses the source object to replace the target object in the first video.

2. The method according to claim 1, wherein generating a three-dimensional model for the source object based on the plurality of images comprises: Sampling the plurality of images to obtain a plurality of image key frames; Determining a sparse point cloud for the target object according to the plurality of image key frames; Training a first model based on the sparse point cloud for the target object according to preset training conditions; And Determining the three-dimensional model using the trained first model according to the plurality of image key frames.

3. The method according to claim 2, wherein the preset training conditions include: Performing iterative training on the first model a first preset number of times based on the sparse point cloud of the target object; And Reducing the learning rate of the first model in response to having performed a second preset number of times of iterative training, wherein the second preset number is less than the first preset number.

4. The method according to claim 1, wherein generating a plurality of animation models for the target object based on the first video comprises: Sampling each video frame of the first video to obtain a plurality of video key frames; Generating an animation model for the target object using a second model according to one video key frame among the plurality of video key frames; Extracting deformation information of the target object from each video key frame among the plurality of video key frames; And Adjusting the animation model for the target object according to the deformation information of the target object in each video key frame to obtain the animation model of the target object in each video key frame.

5. The method according to claim 4, wherein extracting deformation information of the target object from each video key frame among the plurality of video key frames comprises: Extracting pose information and shape information of the target object from each video key frame among the plurality of video key frames; And adjusting the animation model for the target object according to the deformation information of the target object in each video key frame comprises: For each video key frame, using the pose information and shape information of the target object extracted from the video key frame to adjust the animation model for the target object to obtain the animation model of the target object in each video key frame.

6. The method according to claim 4, wherein the fusing the three-dimensional model of the source object and multiple animation models of the target object to generate a second video for the source object includes: determining a typical animation model from the multiple animation models of the target object, wherein the typical animation model of the target object has a consistent pose and shape with the three-dimensional model of the source object; aligning the three-dimensional model with the typical animation model to obtain an aligned three-dimensional model; transmitting multiple skinning weights of the animation model of the target object to the aligned three-dimensional model; and generating a second video for the source object using the aligned three-dimensional model according to the multiple skinning weights.

7. The method according to claim 6, wherein the aligning the three-dimensional model with the typical animation model includes: rotating the three-dimensional model by a first angle according to a rotation vector to obtain a first three-dimensional model, wherein the first angle is the angle indicated by the rotation vector; displacing the first three-dimensional model along a first direction by a first distance according to a displacement vector to obtain a second three-dimensional model, wherein the first direction and the first distance are respectively the direction and distance indicated by the displacement vector; determining the distance between the second three-dimensional model and the typical animation model; and in response to the distance being less than a preset threshold, determining the second three-dimensional model as the aligned three-dimensional model for the source object.

8. The method according to claim 7, wherein the determining the distance between the second three-dimensional model and the typical animation model includes: for each point in the second three-dimensional model, determining the distance between the point and the corresponding point in the typical animation model as the first distance of the point; and determining the sum of the first distances of the respective points in the second three-dimensional model as the distance between the second three-dimensional model and the typical animation model.

9. The method according to claim 6, wherein the generating a second video for the source object using the aligned three-dimensional model according to the multiple skinning weights includes: controlling the aligned three-dimensional model according to the multiple skinning weights to obtain a three-dimensional animation for the source object; and determining a two-dimensional projection of the three-dimensional animation for the source object according to a preset viewing angle as the second video.

10. The method according to claim 9, further comprising: determining background information according to the first video, wherein the background information does not include the target object; and embedding the background information of the first video into the two-dimensional projection as an optimized second video for the source object.

11. An electronic device for processing images, comprising: at least one processor; and coupled to the at least one processor and having instructions stored thereon, which when executed by the at least one processor cause the electronic device to perform actions, the actions including: acquiring a plurality of images and a first video, wherein the plurality of images indicate visual information of a source object from multiple perspectives, and the first video indicates an animation of a target object; Generate a three-dimensional model for the source object based on the multiple images; Generate multiple animation models for the target object based on the first video; and Fuse the three-dimensional model for the source object and the multiple animation models for the target object to generate a second video for the source object, wherein the second video uses the source object to replace the target object in the first video.

12. The electronic device according to claim 11, wherein generating the three-dimensional model for the source object based on the multiple images comprises: Sampling the multiple images to obtain multiple image key frames; Determining a sparse point cloud for the target object according to the multiple image key frames; Training a first model based on the sparse point cloud for the target object according to preset training conditions; and Determining the three-dimensional model according to the multiple image key frames using the trained first model.

13. The electronic device according to claim 12, wherein the preset training conditions comprise: Performing iterative training on the first model a first preset number of times based on the sparse point cloud of the target object; and Reducing the learning rate of the first model in response to having performed a second preset number of times of iterative training, wherein the second preset number is less than the first preset number.

14. The electronic device according to claim 11, wherein generating the multiple animation models for the target object based on the first video comprises: Sampling each video frame of the first video to obtain multiple video key frames; Generating one animation model for the target object using a second model according to one video key frame among the multiple video key frames; Extracting the deformation information of the target object from each video key frame among the multiple video key frames; and Adjusting the one animation model for the target object according to the deformation information of the target object in each video key frame to obtain the animation model of the target object in each video key frame.

15. The electronic device according to claim 14, wherein extracting the deformation information of the target object from each video key frame among the multiple video key frames comprises: Extracting the pose information and shape information of the target object from each video key frame among the multiple video key frames; And adjusting the one animation model for the target object according to the deformation information of the target object in each video key frame comprises: For each video key frame, using the pose information and shape information of the target object extracted from the video key frame to adjust the one animation model for the target object to obtain the animation model of the target object in each video key frame.

16. The electronic device according to claim 14, wherein fusing the three-dimensional model for the source object and the multiple animation models for the target object to generate a second video for the source object comprises: Determine a typical animation model from multiple animation models for the target object, where the typical animation model for the target object has a consistent pose and shape with the three-dimensional model for the source object; Align the three-dimensional model with the typical animation model to obtain an aligned three-dimensional model; Transmit multiple skinning weights of the animation model for the target object to the aligned three-dimensional model; and Generate a second video for the source object using the aligned three-dimensional model according to the multiple skinning weights.

17. The electronic device according to claim 16, wherein the aligning the three-dimensional model with the typical animation model includes: Rotate the three-dimensional model by a first angle according to a rotation vector to obtain a first three-dimensional model, where the first angle is the angle indicated by the rotation vector; Displace the first three-dimensional model along a first direction by a first distance according to a displacement vector to obtain a second three-dimensional model, where the first direction and the first distance are respectively the direction and distance indicated by the displacement vector; Determine the distance between the second three-dimensional model and the typical animation model; and In response to the distance being less than a preset threshold, determine the second three-dimensional model as the aligned three-dimensional model for the source object.

18. The electronic device according to claim 17, wherein the determining the distance between the second three-dimensional model and the typical animation model includes: For each point in the second three-dimensional model, determine the distance between the point and the corresponding point in the typical animation model as the first distance of the point; and Determine the sum of the first distances of the respective points in the second three-dimensional model as the distance between the second three-dimensional model and the typical animation model.

19. The electronic device according to claim 16, wherein the generating a second video for the source object using the aligned three-dimensional model according to the multiple skinning weights includes: Control the aligned three-dimensional model according to the multiple skinning weights to obtain a three-dimensional animation for the source object; and Determine a two-dimensional projection of the three-dimensional animation for the source object according to a preset perspective as the second video.

20. A computer program product, the computer program product being tangibly stored on a non-volatile computer-readable medium and including machine-executable instructions that, when executed, cause the machine to perform actions, the actions including: Obtain a plurality of images and a first video, where the plurality of images indicate visual information of a source object from multiple perspectives, and the first video indicates an animation of a target object; Generate a three-dimensional model for the source object based on the plurality of images; Generate multiple animation models for the target object based on the first video; and Fuse the three-dimensional model for the source object and the multiple animation models for the target object to generate a second video for the source object, where the second video uses the source object to replace the target object in the first video.