Video processing method and related device

By acquiring facial image features of real-world objects, adjusting the facial features of digital objects, and performing consistency processing, the problem of poor visual effects in digital object videos is solved, and more realistic and coherent digital object videos are generated.

CN122002091APending Publication Date: 2026-05-08HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
Filing Date
2024-11-04
Publication Date
2026-05-08

Smart Images

  • Figure CN122002091A_ABST
    Figure CN122002091A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video processing method and a related device. The video processing method comprises the following steps: acquiring a first picture, wherein the first picture comprises a face image of a target object; a first video is acquired, the first video comprises a plurality of first video frames, each first video frame in the plurality of first video frames comprises a first digital object, and the first digital object is a digital object of the target object; extracting the facial features of the facial image of the target object in the first picture, and adjusting the facial features of the first digital object contained in each of the plurality of first video frames based on the facial features of the facial image of the target object; a plurality of second video frames in the second video respectively correspond to a plurality of first video frames, each second video frame comprises a second digital object, and the second digital object is the first digital object of which the facial feature is processed. By adopting the mode, the reality sense of the digital object in the video or the visual effect of the video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a video processing method and related apparatus. Background Technology

[0002] With the development of virtual reality and metaverse technologies, digital objects (such as digital humans), as virtual characters, are attracting increasing attention due to their ability to achieve various applications in virtual digital space. A digital object is a model constructed using computer graphics technology that imitates an object in the real world (such as a real person). While digital objects themselves are static graphics, digital object videos can be generated to exhibit dynamic characteristics. A digital object video is a video that uses multiple consecutive video frames to showcase the dynamic movement of a digital object.

[0003] To generate a digital object video, an initial 3D model can be constructed by simulating a real-world object using 3D modeling software. Further details, textures, and skeletal structures can be added to the initial 3D model, followed by rendering to obtain the 3D digital object. After obtaining the 3D digital object, animations related to it can be set, and rendering the animations will then produce the digital object video.

[0004] The digital object videos generated using the above method have poor visual effects due to the coarseness of the digital objects obtained from 3D modeling. Summary of the Invention

[0005] This application provides a video processing method for improving the visual effects of digital objects in a video. This application also provides a video processing apparatus, a computing device cluster, a computer-readable storage medium, and a computer program product.

[0006] A first aspect of this application provides a video processing method, comprising: acquiring a first image, the first image containing a facial image of a target object; acquiring a first video, the first video including multiple first video frames, each of the multiple first video frames containing a first digital object, the first digital object being a digital object of the target object; extracting facial features of the facial image of the target object in the first image, and adjusting the facial features of the first digital objects contained in each of the multiple first video frames based on the facial features of the facial image of the target object; and obtaining a second video, the multiple second video frames in the second video each corresponding to multiple first video frames, each second video frame containing a second digital object, the second digital object being a first digital object whose facial features have been processed.

[0007] Computing devices in a computing device cluster can acquire a first image and / or a first video from local or other computing devices. The first image contains one or more target objects, each containing facial features. The target objects can be real-world objects. A first image containing the facial image of the target object can be obtained by taking a picture of the target object. The target object is, for example, a person in the real world, and the facial image of the target object is, for example, a human face image. The digital object of the target object is the first digital object, such as a digital person obtained by virtualizing a real person in the real world. A video containing the first digital object can be used as the first video. The first video includes multiple first video frames, each containing the first digital object. When the video is played continuously, dynamic images containing the first digital object can be presented.

[0008] Both the target object and the first digital object possess facial features, which can be extracted from the first image or first video frame. Facial features can be local features such as the nose, mouth, eyes, and skin, or global features encompassing the entire face. After extracting the facial features of the target object's facial image, the facial features of the first digital object contained in multiple first video frames are adjusted based on these features. This makes the facial features of the first digital object more closely match the target object, resulting in a second digital object. After adjusting multiple first video frames, multiple second video frames are obtained, and these second video frames are ordered chronologically to form the second video.

[0009] In the first aspect mentioned above, the first digital object is an object obtained by virtualizing the target object. The target object can be regarded as a kind of original information, reflecting more realistic information that the digital object possessed before virtualization. This information is more realistic than the virtualized digital object. Therefore, by processing the face of the first digital object of the target object in the first video using the facial image of the target object, the facial features of the second digital object in the resulting second video can reflect the facial features of the target object, thereby increasing the realism of the facial image of the digital object in the video. When the target object is an actual object in reality, the facial features of the digital object in the video can reflect the real facial details of the object in the real world, thereby enhancing the realism of the digital object in the video and improving the visual effect of the video. In addition, by adjusting the first video using the first image to obtain the second video, it is also possible to easily adjust the faces of digital objects in existing videos, and flexibly change the visual effect of the faces of digital objects in the video.

[0010] In one possible implementation, before acquiring the first video, the method further includes: acquiring multiple second images, which are images of the target object from different perspectives; performing 3D reconstruction on the multiple second images to obtain a 3D model of the first digital object; and rendering the 3D model of the first digital object to generate a first video, which contains dynamic images of the first digital object.

[0011] Computing devices in a computing cluster can acquire multiple second images from local or other computing devices. These second images can be photos of the target object taken from different perspectives: a forward-facing view, a side-facing view, and a view with its back to the forward-facing view. Feature points are extracted from these second images, and their correspondence in 3D space is determined. Based on this correspondence, the feature points are converted into point cloud data in 3D space. A triangular mesh is generated from the 3D point cloud data to obtain a geometric model. The textures from the multiple second images are mapped onto this geometric model to obtain a 3D model of the first digital object. This 3D model is then imported into a rendering engine, where textures are adjusted, lighting effects are added, and dynamic effects of the first digital object are set, such as its actions or movement paths. After rendering, a first video is obtained. The resulting video reflects the characteristics of the target object from different perspectives, and the first video has a high degree of matching with the target object. Furthermore, a video containing the first digital object can be obtained simply by processing multiple second images; users only need to provide multiple second images, making it highly convenient.

[0012] In one possible implementation, before obtaining the second video, the method further includes: performing consistency processing and fusion processing on the first digital objects contained in the multiple second video frames and the first video frames corresponding to the multiple second video frames; wherein, the consistency processing includes at least one of the following: pose consistency processing, style consistency processing, or background consistency processing; the fusion processing includes stitching the facial image in the second video frame with the non-facial image of the first digital object contained in the first video frame.

[0013] For the sake of brevity, the following text will refer to any video frame obtained by processing the first video frame as a "second video frame." That is, in addition to the existence of second video frames within the second video, any intermediate video frame obtained by processing the first video frame before obtaining the second video can be considered a second video frame. This intermediate video frame may differ from or be identical to the second video frames contained in the second video. When processing (consistency processing and fusion processing) multiple second video frames and the first digital objects contained in the first video frames of each of the multiple second video frames, the multiple second video frames referred to may be intermediate video frames generated after processing the first video frame using the first image. Since further processing is required, these intermediate video frames will differ from the second video frames in the second video.

[0014] Consistency processing refers to processing operations that ensure consistency across different frames of a second video. It's important to note that "consistency" does not mean complete identicalness, but rather that the visual effects presented across multiple frames of the second video exhibit continuity or similarity. For example, to maintain the continuity of the movements of a second digital object in the second video during playback, pose consistency processing can be used. This ensures that the pose of the second digital object in a second video frame matches the pose of the first digital object in its corresponding first video frame. Based on this, the original continuous movements of the first digital object in the first video can be preserved. Specifically, the movements of the first digital object in the first video already possess a certain continuity. However, after adjusting the facial features of the first digital object based on the target object's facial features, the facial features change. For instance, in the first image, the target object faces forward, but multiple first video frames contain first digital objects facing to the side. After adjusting the facial features based on the target object, a digital object facing forward is created. Pose consistency processing then restores this to a second digital object facing to the side, thus preserving the original continuity of the first video.

[0015] After adjusting the facial features of the first digital object in the first video frame using the facial features of the target object in the first image, the style of the first digital object's facial features, such as skin tone and skin smoothness, changes due to the influence of the target object. However, other body features of the first digital object, such as the neck and limbs, are not changed during facial feature adjustment. This results in disjointed skin tone and skin smoothness among different body features, and inconsistent styles between facial features and other features. Therefore, style consistency processing ensures that the style of the second digital object remains consistent across different video frames in the second video, improving the video's visual effect. Furthermore, the first digital object in the first video frame does not cover all image areas; the areas not covered by the first digital object are background areas. After adjusting the facial features of the first digital object in the first video frame, the background areas are not processed synchronously, resulting in inconsistent visual effects between the facial features and the surrounding background areas, such as deviations in hue and brightness. Background consistency processing ensures that the background in different second video frames matches the foreground where the second digital object is located, thereby improving the video's visual effect.

[0016] After adjusting the facial features of the first digital object, since the facial feature process focuses on the facial image, while there are non-facial images in the first video frame, the connection between the facial image and the original non-facial image changes after the facial features are adjusted. For example, the connection between the face and the neck changes. At this time, by stitching the facial image of the second digital object in the processed second video frame with the non-facial image of the first digital object, the connection between the facial image and the non-facial image becomes natural and smooth, thereby improving the visual effect.

[0017] In one possible implementation, before performing consistency processing and fusion processing on the second video frame corresponding to each first video frame and the first digital object contained in each first video frame, the method further includes: responding to adjustment information sent by the user, the adjustment information being used to adjust at least one of skin color and skin texture; adjusting the facial features of the second digital objects contained in the multiple second video frames according to the adjustment information to update the multiple second video frames; and using the updated multiple second video frames for consistency processing and fusion processing.

[0018] To allow users to flexibly adjust the visual effects of facial features in digital objects within a video, users can send adjustment information, and the computing device can respond to this information to further adjust the facial features. Adjustment information includes skin tone and texture. Skin tone can be categorized as light (white, ivory) or dark (brown, black), and the degree of lightness or darkness can be adjusted using this information. Skin texture can be categorized as spots, wrinkles, pores, acne, or skin laxity. Spot types can be categorized as freckles, sunspots, or age spots, and their colors as brown, black, or light-colored, with shapes that can be round or irregular. Wrinkles can be categorized as horizontal or diagonal, and their depth can be adjusted. Pores can be categorized by diameter, such as large or fine pores. Acne types can be categorized as blackheads or papules, and their colors as white, red, or skin-toned, with adjustable size. Skin laxity can be categorized as loose or firm. This adjustment information can be used to eliminate elements in the skin texture that the user does not want to have, such as eliminating spots, pimples, wrinkles or pores, and can also change the attributes of the skin texture, setting the corresponding shape and size of elements such as spots and pimples.

[0019] After performing facial feature processing on the first digital object in the first video frame to obtain the second video frame, the facial features of the second digital object in the second video frame are adjusted based on the adjustment information, thereby obtaining a second video frame containing the second digital object after adjusting the facial features. The second video frame containing the second digital object after adjusting the facial features can be further used for consistency processing and splicing processing to obtain a second video frame after consistency processing and splicing processing. The second video frames after consistency processing and splicing processing can be combined into a second video.

[0020] In one possible implementation, pose consistency processing includes: establishing a mapping relationship between a first keypoint and a second keypoint; wherein the first keypoint is a facial keypoint of a second digital object contained in each second video frame, and the second keypoint is a facial keypoint of a first digital object contained in each first video frame, and the mapping relationship is the correspondence between the position of the first keypoint in each second video frame and the position of the second keypoint in each corresponding first video frame; according to the mapping relationship, aligning the positions of the first keypoints in their respective second video frames according to the positions of the second keypoints.

[0021] After processing the first video frame to obtain the second video frame, facial key points of the second digital object are extracted from the second video frame, and facial key points of the first digital object are extracted from the first video frame. The facial key points extracted from the second and first digital objects correspond to each other, and are key points of specific locations on the face, such as key points of glasses, nose, and mouth. Multiple key points of facial features can be extracted based on a classifier, and the coordinates of each key point in its respective video frame can be determined, and the coordinates of the same feature points can be mapped. For example, based on classifier A, keypoints such as the top left corner 1, top center left eye 1, and top right corner 1 of the second digital object in the second video frame are detected. Based on the same classifier A, keypoints such as the top left corner 2, top center left eye 2, and top right corner 2 of the first digital object in the first video frame are also detected. Then, the coordinates (x1, y1) of the top left corner 1 and the coordinates (x2, y2) of the top left corner 2 are mapped to obtain the positional correspondence between the top left corner 1 and the top left corner 2 of the left eye. Based on a similar principle, the correspondence between each keypoint can be obtained, thus establishing the mapping relationship between the first keypoint and the second keypoint. Since the pose of the first digital object in multiple first video frames is consistent, after the pose changes due to adjustments to the facial features of the first digital object, the first keypoints of the facial features of the second digital object obtained after facial feature adjustments are aligned using the second keypoints of the original facial features of the first digital object as a reference. The purpose of alignment is to ensure that the positions of the first keypoints and the second keypoints are consistent. Specifically, the deviation between the positions of the first key point and the second key point can be calculated, and the first key point can be corrected based on the deviation value, so that the first key point and the second key point are consistent, thereby achieving pose consistency processing, which can improve the coherence of the second digital object between different second video frames in the second video.

[0022] In one possible implementation, style consistency processing includes: adjusting the style of the first digital object in the first video frame based on the style of the facial features of the second digital object; wherein the style includes lighting or skin color.

[0023] After processing the facial features of the first digital object in the first video frame to obtain the second video frame, the style of the facial features of the second digital object in the second video frame changes compared to the style of the facial features of the first digital object. However, the style of the non-facial features of the second digital object remains unchanged compared to the style of the non-facial features of the first digital object. Therefore, the styles of the facial features and non-facial features of the second digital object in the second video frame do not change synchronously, resulting in inconsistency. Therefore, by adjusting the style of the full-body features of the first digital object in the first video frame based on the style of the facial features of the second digital object, a digital object with a consistent full-body style is obtained and updated as the second digital object, thus ensuring that the full-body style of the second digital object in the updated second video frame remains consistent. Styles include elements like lighting or skin tone. Lighting refers to the contrast between light and shadow on a digital object, and mainly involves two parts: illumination and shadow. Illumination is the distribution of light intensity on a digital object; for example, highlights appear very bright, transition areas are gradual areas from bright to dark, and reflected light adds some brightness to dark areas, making them appear less dark. Shadows are created by uneven lighting, with lower light intensity in shadowed areas. Lighting and shadow can affect the realism of digital objects; adjusting lighting and shadow can enhance the realism of a second digital object in a second video frame.

[0024] In one possible implementation, background consistency processing includes: eliminating color difference between the foreground region and the background region; wherein the foreground region is the region where the first digital object is located in each first video frame, and the background region is the region in each first video frame excluding the foreground region.

[0025] The first digital object is the focus area in the first video frame, thus acting as the foreground. The background area contrasts with the area containing the first digital object, serving to highlight it. A color mismatch or lack of contrast between the foreground and background areas creates a perceived color difference, akin to a color discrepancy. This color difference may be inherent in the first video or arise during the processing of the digital object's facial features. Such a color difference can create a sense of disharmony for the viewer and interfere with the realism of the digital object in the video, thus requiring its elimination. Eliminating this color difference can be achieved by bringing the background color closer to the foreground. For example, a unified white balance adjustment can be applied to both the background and foreground areas, or the color difference between them can be smoothed to create a gradient effect, eliminating abrupt boundaries and increasing the harmony between the background and foreground areas, thereby improving the visual effect of the second video.

[0026] In one possible implementation, facial features of the target object's facial image in the first image are extracted, and facial features of the first digital object contained in each of the multiple first video frames are adjusted based on the facial features of the target object's facial image. This includes: extracting the facial mask of the target object in the first image, and the facial mask of the first digital object in each first video frame; replacing the facial mask of the first digital object in each first video frame with the facial mask of the target object, and eliminating mask gaps; wherein, the mask gap is the gap between the replaced mask and the adjacent regions of the mask.

[0027] To maximize the inclusion of the target object's facial features in the first digital object of the first video, and to reflect the target object's facial details on the digital object, facial features in the first video frame of the first video can be adjusted by facial mask replacement. A facial mask represents a mask or outline of the facial region of the target object or the first digital object. It isolates the entire facial region from the target object or the first digital object. The facial mask includes specific areas of the face such as the contour lines, cheeks, facial features, and eyebrows. Facial mask replacement involves scaling, rotating, or translating the target object's facial mask proportionally, and then replacing the pixels of the first digital object's facial mask in the first video frame of the first video with the pixels of the target object's facial mask in the first image using a pixel-by-pixel replacement method, thereby obtaining a second video frame. Multiple second video frames obtained after facial mask replacement can be combined into a second video. Alternatively, the second video frames after facial mask replacement can be further adjusted according to adjustment information, or combined with the first video frames in the first video to achieve consistency processing and splicing processing, thereby obtaining the processed second video frames. Multiple processed second video frames can be combined into a second video. Thus, the second digital object in the obtained second video fully reflects the characteristics of the target object's facial features, enhancing the visual effect of the second digital object in the second video.

[0028] In one possible implementation, obtaining the first video includes: receiving a description of the first video and generating the first video based on the description. The description of the first video can be provided by a computing device in a computing device cluster via a network, or it can be obtained locally on the computing device executing the method. A user can input the description of the first video locally on the computing device via an input device. The description of the first video includes information about the target object, such as its gender, age, height, body type, and other physical characteristics, as well as its action characteristics, such as walking, dancing, sitting posture, or facial expressions. Furthermore, the description can include scene or background information, such as indoors, outdoors, a park, or a street. The computing device performs text analysis and semantic understanding on the description of the first video, extracts key information, performs 3D modeling based on the key information, generates corresponding actions or animations, generates a background or scene, and renders the first video. For example, a description of the first video could be "Please generate a video of a digital human broadcasting content X." In this case, the description can be combined to generate a corresponding digital human video in which the digital human broadcasts content X. This method can flexibly generate the first video based on various descriptions.

[0029] In one possible implementation, adjusting the facial features of a first digital object contained in multiple first video frames based on the facial features of the target object's facial image includes: extracting a first video frame whose facial orientation is the same as the target object's facial orientation as a reference video frame; replacing the facial image of the first digital object contained in the reference video frame with the target object's facial image; displaying the replaced reference video frame to the user for reference, allowing the user to confirm the approximate effect; and upon receiving confirmation from the user, further adjusting the facial features of the first digital objects in the multiple first video frames based on the replaced reference video frame, and performing operations such as consistency processing and stitching. If the user is not satisfied with the effect of the displayed reference video frame, they may request a new first image or require specific adjustments, such as adjusting skin tone or texture. In this case, the first video frames need to be adjusted according to the user's adjustment information to ensure that the visual effect of the digital object in the resulting video is confirmed by the user, thereby improving the visual experience.

[0030] In one possible implementation, adjusting the facial features of a first digital object contained in multiple first video frames based on the facial features of the target object's facial image includes: mapping the facial features of the target object's facial image and the facial features of the first digital objects contained in the multiple first video frames to a preset feature space, generating two sets of corresponding feature vectors respectively; performing linear interpolation on the two sets of feature vectors to obtain a new embedding vector; decoding the embedding vector to convert it into an image containing a more realistic second digital object; performing the above processing on each of the multiple first video frames in the first video to obtain the second video. This method can add facial details of the target object and retain some details of the digital object, improving the realism of the digital object in the second video.

[0031] In one possible implementation, the facial features of a first digital object contained in multiple first video frames are adjusted based on the facial features of the target object's facial image. This includes: separating the shape and texture of the face in the target object's facial image, and separating the shape and texture of the face in the facial images of the first digital objects contained in the multiple first video frames. The texture of the target object's facial shape is then fused with the shape of the facial features of the first digital object in the first video frames to obtain a second video frame. This method can preserve the facial texture of the target object in the second digital object in the second video frame, thereby increasing the facial details of the target object in the digital object, improving the realism of the digital object in the video, and also preserving some features of the digital object.

[0032] In one possible implementation, the second video is obtained by processing the first image and the first video using a generative model. This generative model extracts facial features from the target object's facial image in the first image, and adjusts the facial features of the first digital object contained in multiple first video frames based on these facial features; thus, the second video is obtained. This generative model can encode and decode the first image and the first video to generate the second video. By generating the second video using a generative model, the probability distribution of the first image and the first video can be learned, resulting in a realistic second video.

[0033] In one possible implementation, the generative model includes at least one of a first sub-generative model, a second sub-generative model, or a third sub-generative model. The first sub-generative model is used to adjust the facial features of a first digital object in each of multiple first video frames using the facial features of the target object in the first image. The second sub-generative model is used to adjust the facial features of second digital objects contained in multiple second video frames according to the adjustment information, thereby updating the multiple second video frames. The third sub-generative model is used to perform consistency processing and fusion processing on the multiple second video frames and the first digital objects contained in their respective corresponding first video frames to obtain a second video.

[0034] In this possible implementation, the first, second, and third sub-generative models can be independently trained generative models, thus employing independent sub-generative models to achieve facial feature adjustment, consistency processing, and stitching processing. Each model focuses on its own task, allowing for high flexibility in model training adjustments. Alternatively, an integrated generative model can be constructed, comprising three distinct stages, each corresponding to the functions of the first, second, and third sub-generative models. This integrated generative model improves task coherence.

[0035] In one possible implementation, the second sub-generative model includes a latent space and a decoder. The latent space is used to sample the conditional vector and the first video frame. The conditional vector is a vector obtained by encoding adjustment information. The decoder is used to generate the second video frame based on the data sampled from the latent space. By sampling in the latent space, the influence of the adjustment information on the first video frame is taken into account. This influence can be reflected in the second digital object in the generated second video frame, thereby accurately generating a second video frame that matches the adjustment information.

[0036] In one possible implementation, the color values ​​of the corresponding pixels in the image can be further changed for the second video frame generated by the second sub-generative model, allowing for more flexible adjustments.

[0037] In one possible implementation, the generative model is a generative adversarial network (GAN) model. The second video generated by this GAN model has a high degree of realism.

[0038] In one possible implementation, after obtaining the second video, the second video is sent to a computing device in the computing device cluster used for displaying the video, and the computing device used for displaying the video displays the second video.

[0039] A second aspect of this application provides a video processing method, comprising: a computing device (video display device) in a computing device cluster sending a first image to other computing devices (video processing devices) in the computing device cluster, the first image containing a facial image of a target object; the video display device receiving a second video sent by the video processing device; and the video display device displaying the second video. The second video is obtained by the video processing device in the following manner: acquiring a first video, the first video comprising multiple first video frames, each of the multiple first video frames containing a first digital object, the first digital object being a digital object of the target object; extracting facial features of the facial image of the target object in the first image; adjusting the facial features of the first digital objects contained in each of the multiple first video frames based on the facial features of the target object's facial image; and obtaining a second video, the multiple second video frames in the second video each corresponding to multiple first video frames, each second video frame containing a second digital object, the second digital object being a first digital object whose facial features have been processed. By employing the above method, by transferring the processing procedures dependent on high-performance computing devices to other computing devices, computing devices with lower processing capabilities in the computing device cluster can also present highly realistic digital object videos to users.

[0040] In one possible implementation, before receiving the second video, the method further includes: sending second images to other computing devices (video processing devices) in the computing device cluster. The multiple second images are images of the target object from different perspectives. The other devices (video processing devices) in the computing device cluster are used to perform 3D reconstruction on the multiple second images to obtain a 3D model of the first digital object; the 3D model of the first digital object is then rendered to generate a first video, which contains dynamic images of the first digital object.

[0041] In one possible implementation, after sending the second image, the method further includes receiving a retransmission instruction. This retransmission instruction is provided by other computing devices (video processing devices) when they determine that a first video cannot be generated based on the second image, and is used to instruct the retransmission of a new second image. If the quality of the second image is too poor—for example, lacking important body parts, having too low an image resolution, or an incorrect viewing angle—it may prevent other computing devices (video processing devices) from obtaining a first digital object that meets the minimum standards, and thus from generating the corresponding first video. After receiving the retransmission instruction, a new second image is re-uploaded according to the instruction. Alternatively, a prompt message can be sent to the user, instructing them to re-upload the second image. This method can improve the accuracy of generating a first video containing digital objects.

[0042] In one possible implementation, before receiving the second video, the method further includes: receiving a user's description of the first video and sending the description of the first video to the video processing device. The description of the first video is used to enable the video processing device to generate a first video corresponding to the description. This approach allows for flexible adjustment of the generated first video.

[0043] In one possible implementation, before receiving the second video, the method further includes sending the first video to another computing device (video processing device). This method allows the video display device to specify the first video.

[0044] In one possible implementation, before receiving the second video, the method further includes: receiving adjustment material sent by a video processing device, displaying the adjustment material, receiving user selection information on the adjustment material to generate adjustment information, and sending the adjustment information to the video processing device. The adjustment information is used to adjust at least one of skin color and skin texture. The video processing device is used to adjust the facial features of a second digital object contained in multiple second video frames according to the adjustment information to update the multiple second video frames. The updated multiple second video frames are used for consistency processing and fusion processing.

[0045] A third aspect of this application provides a video processing apparatus having the function of performing video processing as described in the first aspect or any possible implementation thereof. This function can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above function, such as a first acquisition module, a second acquisition module, and a processing module.

[0046] In one possible implementation, the video processing apparatus includes:

[0047] The first acquisition module is used to acquire a first image, which contains a facial image of the target object;

[0048] The second acquisition module is used to acquire a first video, which includes multiple first video frames. Each of the multiple first video frames contains a first digital object, which is a digital object of the target object.

[0049] The processing module is used to extract facial features of the target object's facial image in the first image, adjust the facial features of the first digital object contained in each of the multiple first video frames based on the facial features of the target object's facial image, and obtain a second video. The multiple second video frames in the second video each correspond to multiple first video frames, and each second video frame contains a second digital object, which is a first digital object whose facial features have been processed.

[0050] The video processing apparatus described in the third aspect is capable of generating digital object videos with realistic facial features.

[0051] In one possible implementation, the processing module is also used for:

[0052] Acquire multiple second images, which are images of the target object from different perspectives;

[0053] A 3D model of the first digital object is obtained by reconstructing multiple second images in 3D.

[0054] The first video is generated by rendering the three-dimensional model of the first digital object. The first video contains dynamic images of the first digital object.

[0055] In one possible implementation, the processing module is also used for:

[0056] The first digital object contained in the first video frame and the first video frame corresponding to the first video frame of the second video frame are subjected to consistency processing and fusion processing; wherein, the consistency processing includes at least one of the following: pose consistency processing, style consistency processing, or background consistency processing; the fusion processing includes stitching the facial image in the second video frame with the non-facial image of the first digital object contained in the first video frame.

[0057] In one possible implementation, the processing module is also used for:

[0058] Respond to adjustment information sent by the user, the adjustment information being used to adjust at least one of skin tone and skin texture;

[0059] According to the adjustment information, the facial features of the second digital object contained in multiple second video frames are adjusted to update the multiple second video frames; the updated multiple second video frames are used for consistency processing and fusion processing.

[0060] In one possible implementation, the processing module is also used for:

[0061] Establish a mapping relationship between the first keypoint and the second keypoint; wherein, the first keypoint is the facial keypoint of the second digital object contained in each second video frame, and the second keypoint is the facial keypoint of the first digital object contained in each first video frame. The mapping relationship is the correspondence between the position of the first keypoint in each second video frame and the position of the second keypoint in each corresponding first video frame.

[0062] Based on the mapping relationship, align the positions of the first key points in the corresponding second video frames according to the positions of the second key points.

[0063] In one possible implementation, the processing module is also used for:

[0064] The style of the first digital object in the first video frame is adjusted based on the style of the facial features of the second digital object; wherein the style includes lighting or skin color.

[0065] In one possible implementation, the processing module is also used for:

[0066] Eliminate color difference between the foreground and background regions; wherein, the foreground region is the region where the first digital object is located in each first video frame, and the background region is the region other than the foreground region in each first video frame.

[0067] In one possible implementation, the processing module is also used for:

[0068] Extract the facial mask of the target object in the first image, and the facial mask of the first digital object in each first video frame;

[0069] The face mask of the first digital object in each first video frame is replaced with the face mask of the target object, and the mask gap is eliminated; wherein, the mask gap is the gap between the replaced mask and the adjacent area of ​​the mask.

[0070] A fourth aspect of this application provides a video display device, the video display device comprising:

[0071] The sending module is used to send the first image. The first image contains a facial image of the target object.

[0072] A receiving module is used to receive a second video. The second video is obtained by the video processing device and provided to the video display device in the following manner: acquiring a first video, the first video including multiple first video frames, each of the multiple first video frames containing a first digital object, the first digital object being a digital object of a target object; extracting facial features of the target object's facial image from a first image; adjusting the facial features of the first digital objects contained in each of the multiple first video frames based on the facial features of the target object's facial image; obtaining a second video, the multiple second video frames in the second video each corresponding to multiple first video frames, each second video frame containing a second digital object, the second digital object being a first digital object whose facial features have been processed.

[0073] The display module is used to display the second video.

[0074] In one possible implementation, the sending unit is further configured to: send multiple second images to the video processing device, wherein the multiple second images are images of the target object from different perspectives. The video processing device is configured to perform 3D reconstruction on the multiple second images to obtain a 3D model of the first digital object; and render the 3D model of the first digital object to generate a first video, wherein the first video contains dynamic images of the first digital object.

[0075] In one possible implementation, the video display device is configured as follows:

[0076] The receiving module is also used for retransmission indication information. This retransmission indication information is fed back by the video processing device when it determines that the first video cannot be generated based on the second image, and is used to instruct the retransmission of a new second image.

[0077] The sending module is also used to resend the second image according to the retransmission instruction information.

[0078] In one possible implementation, the video display device is configured as follows:

[0079] The receiving module is also configured to receive a user's description of the first video. The description of the first video is used to instruct the video processing device to generate a first video corresponding to the description.

[0080] The sending module is also used to send a description of the first video to the video processing device.

[0081] In one possible implementation, the sending module is further configured to: send the first video to the video processing device.

[0082] In one possible implementation, the video display device is configured as follows:

[0083] The receiving module is also used to receive adjustment materials sent by the video processing device, and to receive the user's selection information on the adjustment materials in order to generate adjustment information;

[0084] The display module is also used to display and adjust materials.

[0085] The sending module is also used to send the adjustment information to the video processing device.

[0086] The fifth aspect of this application provides a computing device cluster including at least one computing device, each computing device including a processor and a memory, wherein the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, causing the computing device cluster to perform a method as described in the first aspect or any possible implementation thereof.

[0087] A sixth aspect of this application provides a computing device cluster including at least one computing device, each computing device including a processor and a memory, wherein the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, causing the computing device cluster to perform a method as described in the second aspect above or any possible implementation thereof.

[0088] A seventh aspect of this application provides a chip system including one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the interface circuits are configured to receive signals from the memory of a computing device cluster and send signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the computing device cluster performs a method as described in the first aspect or any possible implementation thereof.

[0089] The eighth aspect of this application provides a chip system including one or more interface circuits and one or more processors; the interface circuits and processors are interconnected via lines; the interface circuits are configured to receive signals from the memory of a computing device cluster and send signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the computing device cluster performs a method as described in the second aspect above or any possible implementation thereof.

[0090] A ninth aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a cluster of computing devices, cause the cluster of computing devices to perform a method as described in the first aspect or any possible implementation thereof.

[0091] The tenth aspect of this application provides a computer-readable storage medium storing instructions that, when executed on a cluster of computing devices, cause the cluster of computing devices to perform a method as described in the second aspect or any possible implementation thereof.

[0092] The eleventh aspect of this application provides a computer program product comprising computer program code that, when run on a computer, causes the computer to perform a method as described in the first aspect or any possible implementation thereof.

[0093] The twelfth aspect of this application provides a computer program product comprising computer program code that, when run on a computer, causes the computer to perform a method as described in the second aspect or any possible implementation thereof.

[0094] The thirteenth aspect of this application provides a video processing system, including a cloud-side device and an end-side device. The cloud-side device is used to execute the method of the first aspect or any possible implementation thereof, and the end-side device is used to execute the method of the second aspect or any possible implementation thereof.

[0095] The fourteenth aspect of this application provides a video processing system, including a computing device cluster, the computing device cluster being used to perform the method of the first aspect or any possible implementation thereof.

[0096] The relevant features and effects of the third, fifth, seventh, ninth, eleventh, thirteenth and fourteenth aspects of this application can be understood by referring to the corresponding descriptions in the first aspect or any possible implementation of the first aspect.

[0097] The relevant features and effects of aspects four, six, eight, ten, and twelfth of this application can be understood by referring to the corresponding descriptions in aspect two or any possible implementation of aspect two. Attached Figure Description

[0098] Figure 1A This is a schematic diagram of the architecture of a professional graphics workstation provided in an embodiment of this application;

[0099] Figure 1B This is a schematic diagram of the architecture of a cloud server provided in an embodiment of this application;

[0100] Figure 1C A schematic diagram of the architecture of the end-to-cloud collaborative video processing system provided in this application embodiment is provided for the purpose of this application embodiment;

[0101] Figure 1D This is a schematic diagram of the architecture of the cloud system provided in an embodiment of this application;

[0102] Figure 2 This is a schematic diagram of an embodiment of the video processing method provided in this application;

[0103] Figure 3 This is an exemplary schematic diagram of facial feature processing provided in the embodiments of this application;

[0104] Figure 4 This is an exemplary schematic diagram of selecting a first video frame for facial feature processing provided in an embodiment of this application;

[0105] Figure 5 This is an exemplary schematic diagram of processing a first video frame in conjunction with a first image, provided in an embodiment of this application;

[0106] Figure 6 This is a schematic diagram of an embodiment of the video processing method in a cloud-edge integrated scenario provided in this application.

[0107] Figure 7 This is an exemplary schematic diagram of processing digital human videos provided in an embodiment of this application;

[0108] Figure 8 This is an exemplary schematic diagram of end-to-cloud interaction to achieve facial realism adjustment in the embodiments of this application;

[0109] Figure 9This is an exemplary schematic diagram of the structure of the realism adjustment model provided in the embodiments of this application;

[0110] Figure 10 This is a schematic diagram of an embodiment of the video processing method provided in this application;

[0111] Figure 11 This is an exemplary schematic diagram illustrating video processing of a local first video in an embodiment of this application;

[0112] Figure 12 This is a schematic diagram of an embodiment of the video processing method provided in this application;

[0113] Figure 13 This is a schematic diagram of the structure of an embodiment of the video processing apparatus provided in this application;

[0114] Figure 14 This is a schematic diagram of the structure of an embodiment of the video display device provided in this application;

[0115] Figure 15 This is a schematic diagram of the structure of one embodiment of the computing device provided in this application;

[0116] Figure 16 This is a schematic diagram of the structure of one embodiment of the computing device cluster provided in this application;

[0117] Figure 17 This is a schematic diagram of the structure of one embodiment of the computing device cluster provided in this application. Detailed Implementation

[0118] The embodiments of this application are described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. As those skilled in the art will understand, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0119] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0120] This application provides a video processing method to enhance the realism of digital objects in a video. This application also provides a video processing apparatus, a computing device cluster, a computer-readable storage medium, and a computer program product.

[0121] The video processing method provided in this application can be executed on at least one computing device in any type of computing device cluster, as long as the computing device has the corresponding video processing capability to perform video processing. Several types are given below.

[0122] like Figure 1A As shown, combined with Figure 1A This application describes the architecture of a computing device within a computing device cluster, namely a professional graphics workstation, as provided in its embodiments. The video processing method provided in this application can be executed locally on the professional graphics workstation. The dedicated graphics workstation includes a host computer, input / output devices, etc. The host computer may include a high-performance graphics processor, a high-performance central processing unit, and a large-capacity memory. Input / output devices include, for example, a keyboard and a monitor. The input / output devices can be used to allow users to upload a first image and to display a second video to the user. The host computer can process the first video based on the first image uploaded by the user to obtain a second video, which is then displayed on the monitor. Alternatively, the host computer can also generate the first video using a second image uploaded by the user or a description of the first video. Alternatively, the host computer can also receive the first video uploaded by the user. Furthermore, the host computer can respond to user selections via input devices to read pre-stored first images or first videos from memory, thereby generating the second video. The professional graphics workstation can be equipped with applications for video processing to interact with the user and implement video processing.

[0123] In addition to adopting Figure 1A Besides the structure shown in the professional graphics workstation diagram, it can also be implemented using other high-performance computing devices, and can also... Figure 1A The structure shown can be modified. For example, input devices can be wireless keyboards, wireless mice, touch screens, cameras (for capturing images and gesture recognition), audio input devices (for voice control), etc., and output devices can be projectors, virtual reality glasses, etc. The number and form of the host are not limited.

[0124] like Figure 1B As shown, Figure 1B This is a schematic diagram of the architecture of a cloud server provided in an embodiment of this application. The cloud server can be a computing device in a computing device cluster, and may include storage nodes, scheduling nodes, and worker nodes.

[0125] Storage nodes are used to store data, such as a first image, a second image, a first video, or a second video. Additionally, storage nodes can store training samples used to train a generative model. The generative model can then process the first image and the first video to generate the second video. Storage nodes can send their stored data to a scheduling node.

[0126] The scheduling node's functionality can be implemented through software or hardware. The scheduling node can retrieve the first image from the storage node, send the first image to the worker nodes, and assign tasks. Furthermore, the scheduling node can also retrieve a second image or a first video from the storage node and send it to the worker nodes.

[0127] A scheduling node, as an example of a software functional unit, can be responsible for distributing video processing tasks. A scheduling node may include code running on a computing instance. This computing instance can include at least one of a physical host (computing device), a virtual machine, or a container. The scheduling node manages and allocates video processing tasks to multiple worker nodes, such as graphics processing unit (GPU) accelerators or rendering nodes, enabling efficient execution of the video processing tasks.

[0128] Furthermore, the aforementioned computing instance can be one or more. For example, a scheduling node may include code running on multiple hosts / virtual machines / containers to handle large-scale video rendering or generative model inference tasks. The multiple hosts / virtual machines / containers used to run this code can be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run this code can be distributed within the same availability zone (AZ) or in different AZs, each AZ comprising one or more geographically proximate data centers. Typically, a region can include multiple AZs.

[0129] As an example of a hardware functional unit, a scheduling node can include at least one computing device, such as a server. Alternatively, a scheduling node can also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0130] Worker nodes can be used to process a first video by combining a first image to obtain a second video. Furthermore, worker nodes can be used to generate the first video by combining the description of the second image and the first video. Worker nodes can execute video processing tasks assigned by the scheduling node. After obtaining the second video, it can be fed back to the storage node. Worker nodes can be physical machines, virtual machines (VMs), or containers, etc. A worker node can include one or more central processing units (CPUs) and graphics processing units (GPUs), and can be either a CPU or a GPU.

[0131] like Figure 1C As shown, Figure 1C This is a schematic diagram of the architecture of the edge-cloud collaborative video processing system provided in the embodiments of this application. The edge-cloud collaborative video processing system includes multiple computing devices in a computing device cluster, and the multiple computing devices include cloud-side devices and edge-side devices.

[0132] The edge-cloud collaborative video processing system provided in this application includes a cloud-side device and an edge-side device, which communicate with each other via a network. The cloud-side device can acquire a first image provided by the edge-side device, and process a first video locally on the cloud-side device based on the first image to obtain a second video. Alternatively, the cloud-side device can acquire a first image and a second image provided by the edge-side device, generate a first video based on the second image, and process the first video in conjunction with the first image to obtain the second video. Alternatively, the cloud-side device can acquire a first image and a description of the first video provided by the edge-side device, generate a first video based on the description of the first video, and process the first video in conjunction with the first image to obtain the second video. Alternatively, the cloud-side device can acquire a first image and a first video provided by the edge-side device, and process the first video in conjunction with the first image to obtain the second video. After obtaining the second video, the cloud-side device can provide the second video to the edge-side device.

[0133] Various types of applications can be installed on the edge device, such as live streaming applications, news broadcasting applications, conferencing applications, and video creation applications. In live streaming applications, users can be represented by digital avatars in the live video stream. In news broadcasting applications, digital avatars replace real people in news broadcasting videos. In conferencing applications, participants can join the meeting using digital avatars in the video conference screen. In video creation, creators can create videos with digital avatars based on their creative needs. Users can upload a first image and a first video in any application. Furthermore, users can upload a second image (a full-body photo of the user) in any application instead of uploading the first video, so that the cloud-side device can generate the first video based on the full-body photo. Alternatively, users can input a description of the first video to be generated in any application, so that the cloud-side device can retrieve the first video based on the description.

[0134] The cloud-side device can acquire a first image and a first video from the edge device via the network. By combining the facial features of the target object in the first image and processing the facial features of the first digital object in the first video, a second video is obtained and provided to the edge device. Alternatively, the cloud-side device can acquire a second image from the edge device and generate the first video based on the second image, instead of directly acquiring the first video from the edge device.

[0135] The edge device acquires the second video provided by the cloud device, plays the second video, and displays a realistic digital object video, thereby realizing various functions such as digital news broadcasting and virtual live streaming.

[0136] The cloud-side device can be a worker node in the cloud system. The cloud system also includes a scheduling node. After receiving the first image and obtaining the first video from the edge device, the scheduling node can perform the corresponding video processing. The scheduling node can also assign the first image and the first video to one or more worker nodes in the cloud system, whereby one or more worker nodes will perform the corresponding video processing.

[0137] End-side devices can be terminal equipment. Terminal equipment, also known as user equipment (UE), mobile station (MS), mobile terminal (MT), etc., is a device that includes wireless communication functions (providing voice / data connectivity to users), such as a handheld device with wireless connectivity. Examples of current terminal equipment include: mobile phones, tablets, laptops, PDAs, wireless routers, mobile internet devices (MID), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving cars, wireless terminals in vehicle-to-everything (V2X) communication, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, and wireless terminals in smart homes, etc. For example, wireless terminals in the Internet of Vehicles (IoV) can be in-vehicle equipment, vehicle-mounted equipment, in-vehicle modules, or vehicles themselves. Wireless terminals in industrial control can be robots. For example, wireless terminals in autonomous driving can be drones. These terminal devices can run Android, iOS, Windows, or other operating systems. Applications that need to display videos containing digital objects can run on these terminal devices, such as live streaming applications, news broadcasting applications, conferencing applications, and video creation applications.

[0138] The above Figure 1C The cloud-side device can be located within the cloud system; the cloud system architecture can be found in [reference needed]. Figure 1D To understand. For example Figure 1D As shown, the cloud system includes a cloud platform and basic resources. The cloud platform includes a cloud platform manager, and the scheduling node described above can be... Figure 1DThe cloud platform manager in the system. Basic resources can include multiple servers, each of which can be a worker node, or each server can include multiple worker nodes.

[0139] exist Figure 1D The working node can be a computing device card or a virtual machine (VM). The computing device card can be at least one of a central processing unit (CPU), a graphics processing unit (GPU), and a network processing unit (NPU).

[0140] The cloud platform manager maintains or periodically collects information about each worker node in the basic resources, such as the resource usage status (resource utilization rate or resource idle rate) on each worker node. This information can serve as auxiliary decision-making information when allocating rendering tasks.

[0141] The cloud platform manager can obtain the first image provided by the client. After obtaining the first video, the cloud platform manager can execute the corresponding video processing procedures. The cloud platform manager can also assign the first image and first video to one or more worker nodes in the cloud system, whereby one or more worker nodes will execute the corresponding video processing procedures.

[0142] After the worker node completes the video processing task, the cloud platform manager can provide the user with the second video obtained after video processing.

[0143] The hardware architecture involved in the embodiments of this application has been described above. The implementation of the video processing method that can be executed by the above hardware structure will be described below.

[0144] like Figure 2 As shown, one embodiment of the video processing method provided in this application includes:

[0145] Step 201. Obtain the first image, which contains the facial image of the target object.

[0146] The target object can be a real, existing object in the real world, such as a real person. The first image can be obtained by taking a photograph of the target object, thus reflecting its true information. Alternatively, the first image can be generated in other ways, such as using a generative model.

[0147] Since the facial image of the first digital object in the first video frame of the first video needs to be adjusted based on the facial image in the first image to obtain the second video, the quality of the first image affects the effect of the obtained second video. In order to improve the effect of the second video, one or more of the following constraints can be set on the first image:

[0148] The first image has one or more attributes, such as sharpness, resolution, or completeness, that are greater than a first preset value. This first preset value includes a sharpness benchmark value, a resolution benchmark value, or a completeness benchmark value. Sharpness can be evaluated using the Laplacian operator; a higher Laplacian operator response value indicates greater sharpness. Sharpness can also be evaluated using edge detection algorithms; the detection of distinct edges indicates higher sharpness, while the detection of blurry edges indicates lower sharpness. Resolution can be measured in pixels, for example, a basic resolution of 8 megapixels is required. Completeness refers to the completeness of the content in the first image, such as detecting whether the first image includes complete facial key points, such as facial features.

[0149] The first image is a picture of the target object taken with a high-definition camera. By determining the image metadata of the first image, such as the exchangeable image file format (EXIF), the model of the camera that took the first image can be determined, and whether the camera is a high-definition camera of the corresponding model can be confirmed. If so, it can be determined that the first image was taken with a high-definition camera. Furthermore, it can be determined whether the target object exists in the first image, such as whether there is a person whose entire facial area is captured, thus confirming that the first image was taken with a high-definition camera.

[0150] The first image contains more information than the second preset value. Information entropy can be used to measure the information content of the first image; the higher the information entropy, the greater the information content, indicating more detail in the first image. Alternatively, a color histogram can be used; the more colors in the color histogram, the greater the information content. Other methods can also be used to measure information content. The second preset value is a value that conforms to the benchmark information content; this value can be set according to actual calibration.

[0151] The first image is a picture of a preset type, such as a portrait, close-up photo, or facial scan. For this type of first image, the facial image of the target object contains rich detail, exceeding the detail of the facial features of the first digital object. This rich detail includes fine wrinkles, pores, skin folds, fine hairs, facial expressions, realistic lighting, and facial rosiness or tone.

[0152] Furthermore, if there is flexibility in the desired effect of the second video, the above constraints can be omitted, and the first image obtained can be directly processed.

[0153] The first image can come from various sources, such as user uploads, downloads from servers, or data captured by image acquisition devices (such as cameras), and there are no restrictions here.

[0154] Step 202. Obtain the first video, which includes multiple first video frames. Each of the multiple first video frames contains a first digital object, which is the digital object of the target object.

[0155] The first video is a digital media file containing multiple video frames, including multiple first video frames. Embedded within these multiple first video frames are first digital objects. These first digital objects are virtual digital objects generated based on the characteristics of the target object. For example, the target object in the first image is person A, and the first digital object is constructed using person A's body data. Body data of person A is acquired through 3D scanning, motion capture, or depth cameras; this data includes the body's geometry, posture, and facial expressions. The acquired data is processed to extract key information, such as the body's contours, joint positions, and dynamic movement patterns. The first digital object is created using computer graphics techniques (such as mesh modeling). This process can utilize point cloud data or skeletal animation systems to transform the extracted body data into a renderable 3D model, and can also add textures and material properties to the digital object to enhance its realism. Textures can be created by capturing surface images of the actual object or using generative modeling methods. Alternatively, although the first digital object may not be directly constructed using the target object's data, it may be a digital object selected from a large library of digital object assets that is similar to the target object. This similarity can manifest in characteristics such as age, gender, height, body type, and the type of object, such as people or pets. In this case, the first digital object can also be considered as a digital object similar to the target object.

[0156] The first video can be obtained in various ways, such as from a pre-stored video library, by generating the first video based on multiple second images, or by generating the first video based on a description of the first video, etc. These will be introduced in detail below.

[0157] The first video is obtained from a pre-stored video. For example, it can be read directly from local storage (such as a solid-state drive or flash memory) or received from other devices (servers or terminal devices, etc.) as the first video. The criteria for the first video are that multiple consecutive first video frames contain a first digital object corresponding to the target object. This method is suitable for scenarios where other content of the generated second video is not required, and only realistic facial features are desired. Alternatively, this method can directly read an existing first video and quickly process it to obtain the second video.

[0158] Generating a first video based on multiple second images includes: acquiring multiple second images, each representing an image of a target object from different perspectives; performing 3D reconstruction on the multiple second images to obtain a 3D model of the first digital object; and rendering the 3D model of the first digital object to generate a first video, which contains dynamic footage of the first digital object. Each second image may include an object from a specific perspective, such as a person from a specific perspective, such as a frontal view, a side view, or a rear view. The second images can be uploaded by the user, read from local storage, or downloaded from a server.

[0159] An example of how to perform 3D reconstruction from multiple second images to obtain a 3D model of the first digital object is shown below:

[0160] Feature points are detected and matched in multiple second images from different perspectives. The 3D coordinates of the matched feature points are calculated to generate a preliminary point cloud model. The point cloud model is then converted into a 3D surface model. The textures from the multiple second images from different perspectives are mapped onto the 3D surface model to generate a 3D model of the first digital object. This method can generate a high-precision 3D model of the first digital object using second images from multiple perspectives.

[0161] A second image is selected, and features such as height, gender, age, or body type of the objects within it are identified. These features are then input into a pre-defined parametric digital object model to obtain a 3D digital object. By combining information such as clothing or decorations of the object from multiple second images, textures for clothing or decorations are added to the 3D digital object, and after rendering, a 3D model of the first digital object is obtained. This method adjusts the pre-defined parametric digital object model to quickly generate a 3D model of the first digital object.

[0162] Multiple second images from different perspectives are input into a trained deep learning model. The model learns the structure and lighting information in the second images to generate a polygonal mesh model, which is then rendered and textured to generate a 3D model of the first digital object. This method can generate a high-precision 3D model of the first digital object based on limited photo data.

[0163] Besides the examples above, other methods can also be used to construct the first digital object, which are not limited here.

[0164] The following is an example of how to render the 3D model of the first digital object to generate the first video:

[0165] Tools such as Blender and Maya are used to rig the 3D model of the first digital object, which is then used to control the model's movements. Combining keyframe animation or motion capture data, animation is added to the 3D model of the first digital object, increasing actions, expressions, and poses. Rendering is then performed, with parameters including lighting, shadows, and reflections. Multiple output images are treated as video frames and combined to create the first video. This method can generate a first video with high-quality rendering.

[0166] Real-time rendering is achieved using tools such as Unity and Unreal Engine. By importing the 3D model of the first digital object into these tools, and using built-in animation tools or importing external animation data, the model is applied to the 3D model of the first digital object. Real-time rendering lighting, materials, and environmental effects are then set, the animation is recorded, and exported as a video file to obtain the first video. This method allows for the rapid generation of the first video using real-time rendering tools.

[0167] Besides using the above method to render the first video, other methods can also be used to render the first video, and no limitation is made here.

[0168] Besides the methods described above for generating the first video, it is also possible to generate the first video based on a description of the first video, without relying on a second image. The process involves obtaining a description of the first video and generating the first video based on that description. This description can be uploaded by the user or a pre-defined description. The main process of generating the first video based on its description is described below:

[0169] The system receives a text description of the first video from the user. Based on natural language processing (NLP) technology, it extracts key information from the description (such as age, gender, height, body type, hairstyle, skin color, etc.). Named entity recognition is used to identify the features of the target object and establish its preliminary outline. For example, the description "a tall, blond young man wearing sportswear is running" includes multiple features: height, hair color, age, clothing, and actions.

[0170] Load a pre-defined 3D digital object library (material library). This library contains basic objects such as people, pets, and animals of various kinds. Based on the parsed text features, select the basic digital object that best matches the description. For example, based on information such as gender, age, and body type, select a matching 3D human model from the database. After selecting the basic model, further detail adjustments are made to the model based on specific features described (such as hairstyle, skin color, and clothing). After modifying the model's appearance, such as changing hair color, changing to the clothes described, and modifying skin color and texture, the 3D model of the first digital object is obtained. For example, if the description requires "fair skin and wearing glasses," facial and appearance details that meet the requirements will be added to the digital object based on these features.

[0171] If the text describes the target object's actions (such as "running," "jumping," "waving," etc.), these actions need to be parsed. These actions are then mapped onto the skeletal animation of the first digital object's 3D model. The actions are applied to the first digital object's 3D model using 3D animation software (such as Maya, Blender, etc.) or a skeletal animation engine to make it exhibit the dynamic effects described. For example, the "running" action in the description will cause the model's limbs to move in a coordinated skeletal motion, mimicking the movement of a runner.

[0172] Based on the description of the first video, a suitable background or scene can be generated or selected for the first digital object. If the description includes scene information (such as "running in the park"), a park scene can be selected from the background material library, or an environment matching the description can be generated. Using scene rendering technology, the 3D model of the first digital object is composited with the background, ensuring that elements such as lighting, shadows, and proportions are consistent between the digital object and the background, creating a highly realistic scene.

[0173] The system further renders the 3D model of the first digital object frame-by-frame, dynamically depicting its appearance within the scene. Based on the animation's settings, each frame of the video is generated, ensuring smooth motion. During each frame rendering, lighting, shadows, and other scene effects are processed to enhance the video's realism. After rendering, all frames are composited, and the video is encoded to generate a complete, playable video file, serving as the first video. Post-processing can be performed after generating the first video to optimize image quality, frame rate, etc. Fine-tuning can also be done according to the user's specific needs, such as adjusting video resolution and color saturation.

[0174] Besides using the above method to generate the first video, other methods can also be used to generate the first video, which are not limited here.

[0175] Step 203. Extract the facial features of the target object's face image in the first image, and adjust the facial features of the first digital object contained in each of the multiple first video frames based on the facial features of the target object's face image.

[0176] To facilitate subsequent processing, facial features of the target object's face in the first image can be extracted first. An example of the extraction method is as follows: To improve extraction accuracy, the first image can be preprocessed. Preprocessing includes: eliminating noise, adjusting the image to a uniform size and resolution, normalizing the image, and converting the RGB image to grayscale or other color spaces (such as YUV or HSV). After preprocessing, a pre-trained Haar feature cascade classifier can be used for face detection. By scanning different regions in the image, potential facial areas can be quickly found. Alternatively, detection methods based on deep neural networks (such as Faster R-CNN, SSD, YOLO) can accurately detect faces in complex environments. Further detection of multiple key points in the target object's face image in the first image is then performed. Based on the positions of these key points, geometric features are calculated, such as the distance between the eyes, the distance from the nose to the chin, and the width of the cheeks. These features can be used to determine the structure and proportions of the face. Finally, feature vectors are used to encode the facial features of the target object. Using the geometric features extracted earlier, the facial information is transformed into a fixed-length vector. The value in each vector represents a different facial feature, which typically includes information such as expression, posture, texture, and face shape, thus obtaining the facial features of the target object's facial image.

[0177] Adjusting the facial features of the first digital object contained in each of the multiple first video frames of the first video requires using the facial features of the target object in the first image as a reference object. This ensures that the facial features of the second digital object in each of the multiple second video frames of the resulting second video match the facial features of the target object. The main principles include: directly replacing the facial image of the first digital object in the first video frame with the facial image of the target object to obtain the second video frame; or extracting a certain attribute from the facial features of the target object and using this attribute to adjust the corresponding attribute in the facial features of the first digital object; or performing feature fusion between the facial features of the target object and the facial features of the first digital object. Several different implementation methods for facial feature adjustment are illustrated below:

[0178] A second video is obtained by replacing the facial images of a first digital object in multiple first video frames with the facial image of the target object. Specifically, for facial image replacement, the bounding boxes of the facial images of both the target object and the first digital object are detected, and the corresponding facial images are cropped based on these bounding boxes. After aligning the facial images, the scaling ratio of the target object's facial image is adjusted, and the adjusted target object's facial image is mapped onto the first digital object's facial image, thus achieving facial image replacement. This method directly replaces the facial image, preserving rich facial details of the target object and improving the realism of the digital object.

[0179] The generator extracts latent vectors representing attributes such as facial features, wrinkles, skin tone, blemishes, or expressions from the facial features of the target object in the first image. It then extracts latent vectors representing the corresponding attributes from the facial features of the first digital object in the first video frame. Based on the latent vectors corresponding to the target object, it modifies the latent vectors corresponding to the first digital object. The modified latent vectors are then input into a generator to produce a second video frame containing the attributes of the target object. The second digital object in the generated second video frame contains the corresponding attributes of the target object's facial features.

[0180] Facial features of the target object in the first image and facial features of the first digital object in the first video frame are extracted using a convolutional neural network. These two sets of facial features are mapped into a feature space to generate two feature vectors. Interpolation is then performed on these two feature vectors to obtain a new embedding vector. The interpolation coefficients are used to control the weights; by setting the weights corresponding to the facial features of the target object to be greater than those of the facial features of the first digital object, the generated image's face more closely resembles the target object's face. This embedding vector is then decoded and converted into an image using a generative adversarial network. This image contains a more realistic second digital object while incorporating the original facial features of the target object.

[0181] Besides using the above method to obtain the second video, other implementation methods can also be used, which are not limited here.

[0182] In order to flexibly adjust the facial features of the second digital object in the second video frame, the system can respond to adjustment information sent by the user, the adjustment information being used to adjust at least one of skin color and skin texture; according to the adjustment information, the facial features of the second digital object contained in multiple second video frames are adjusted to update multiple second video frames; the updated multiple second video frames are used for consistency processing and fusion processing.

[0183] Users send adjustment information through some form of interaction (such as a UI interface, input device, or command). This adjustment information can include adjustments to skin tone or skin texture. Users can adjust the skin tone of a target object, for example, by selecting a color palette, using a slider, or entering a color code. Users can adjust the skin's smoothness, pore size, wrinkle depth, number of blemishes, etc. This can be achieved through parameter selection in the user interface, such as selecting "increase skin smoothness" or "reduce wrinkles." Specific features can also be adjusted, such as the tone, luster, and moisture of specific areas (forehead, cheeks). After receiving the adjustment information, the system determines the specific color of the skin tone entered by the user. For example, if the user selects "dark tone," it will be interpreted as a darker color value within the RGB range. For skin texture adjustments, specific adjustment parameters can be generated based on the user's input. For example, if the user selects "smooth skin," it will be interpreted as reducing the roughness of the skin texture, or if the user selects "reduce pores," an image filter will be applied to reduce the detail of the skin's pores. For localized adjustments, color replacement can be applied to specific facial areas (such as the forehead and chin) without affecting other areas. By analyzing high-frequency components in the image (representing texture details), these high-frequency information can be selectively weakened, making the skin appear smoother. For adjustments to increase skin radiance, a reflection model is applied to the skin area to simulate the smooth reflection effect of natural light. The reflective properties of the skin can be enhanced by adjusting the brightness and color of highlight areas. After adjusting the facial features of the second digital object contained in multiple second video frames according to the adjustment information, updated second video frames are obtained. These multiple second video frames can then undergo further consistency processing and fusion. This allows the second video frames to flexibly match user requirements.

[0184] like Figure 3 As shown, Figure 3 This is a schematic diagram illustrating facial feature processing as provided in an embodiment of this application.

[0185] When processing the facial features of the first digital object in the first video frame based on the facial features of the target object, in order to preserve the facial features of the target object, the facial features of the target object can be used to replace the facial features of the first digital object in the first video frame. Furthermore, adjustments can be made to specific facial features. The adjustment method involves adjusting the facial features of the replaced first digital object according to the adjustment information to obtain the second digital object.

[0186] Figure 3In this process, the first video can be obtained based on the second image, a description of the first video, or a pre-stored first video. After obtaining the first video, a first video frame and consecutive video frames (multiple consecutive first video frames) are extracted from it. The facial features of the target object in the first image are replaced in a first video frame to obtain a second video frame. The user sends adjustment information based on this second video frame, and the facial features of the second video frame are adjusted according to the adjustment information to obtain an updated second video frame. Subsequently, the updated second video frame is merged with the multiple consecutive first video frames. The merging process can simply replace the facial features of the second digital object in the updated second video frame with the facial features of the first digital object in the multiple consecutive first video frames, or it can perform consistency processing and splicing processing to ensure strong consistency between different second video frames of the resulting second video, presenting a natural and coherent visual effect.

[0187] Examples of how facial feature replacement can be implemented are as follows:

[0188] The process involves extracting the facial mask of the target object from the first image and the facial mask of the first digital object in each first video frame. The facial mask of the target object is then used to replace the facial mask of the first digital object in each first video frame, eliminating mask gaps. Specifically, a facial mask image of the target object and a facial mask image of the first digital object in each first video frame are created. The mask covers the contour of the face and is a binary image where white areas represent mask regions. The facial mask of the target object is used to replace the facial mask in the first video frame. After mask replacement, discontinuous or missing regions exist between the mask region and other regions in the image; these regions are called mask gaps. Small gaps in the mask can be filled using morphological operations, such as dilation to expand the mask region and fill small gaps, erosion to shrink the mask region and remove small noise, opening (erosion followed by dilation) to remove small noise, and closing (dilation followed by erosion) to fill small gaps. Additionally, interpolation methods can be used to fill missing regions in the mask.

[0189] Facial feature points of the target object in the first image and facial feature points of the first digital object in the first video frame are extracted, and the deformation mapping from the facial feature points of the target object to the facial feature points of the first digital object is calculated. The faces of both the target digital object and the first digital object are divided into multiple meshes, such as triangular meshes. The mesh of the digital object is mapped onto the mesh of the first digital object, and interpolation is performed on specific feature points and their neighborhoods to gradually transform the face of the first digital object into the face of the target object. The resulting face after replacement using this method has high realism.

[0190] Face replacement can also be achieved using generative models. Here are a few examples of how:

[0191] A face replacement model is constructed based on a Generative Adversarial Network (GAN), which includes a generator and a discriminator. The generator combines a first image and a first video frame to generate a new synthetic image containing the facial features of the target object, as well as other parts of the first digital object. The discriminator determines whether the generated synthetic image is realistic; if so, it is used as the image with replaced facial features. During the training phase, the generator combines training samples corresponding to the first image and the video frame to generate a synthetic image. The discriminator determines whether the synthetic image is realistic and provides feedback. If it is not realistic, the generator adjusts its parameters based on the feedback and regenerates the synthetic image. The discriminator further judges until it determines that the regenerated synthetic image is realistic, at which point training is complete. Using this method, the resulting image with replaced facial features presents a natural and realistic visual effect.

[0192] The system is built upon an autoencoder, comprising an encoder and a decoder. The encoder encodes the first image and the first video frame into low-dimensional latent vectors, which are then fused. This fused latent vector contains features such as the shape and pose of facial features. The decoder generates an image based on this fused latent vector, incorporating portions of the first image and other parts of the first video frame. During training, training samples corresponding to both the first image and the video frame can be used for joint training. This allows the model to receive both sets of samples simultaneously and optimize a joint loss function. This function measures the difference between the generated image and the corresponding input training samples. By minimizing this difference, the generated image contains the facial features of the first image and other parts of the training samples from the video frame, resulting in a more realistic image.

[0193] In addition to the models mentioned above, other generative models can also be used to achieve face replacement, and no specific limitations are imposed here.

[0194] For images with replaced facial features, a realism adjustment model can be used to further refine the facial features, controlling the realism of the features based on the adjustment information. Realism adjustment models can be built based on generative models; several implementation methods are illustrated below:

[0195] The realism adjustment model can be built upon a diffusion model. The image with replaced facial features is used as input for processing. This model extracts features such as wrinkles, skin tone, or blemishes from the image according to the adjustment information, progressively converting the image into a noisy image. The noisy image is then denoised, and the features to be adjusted are used to guide the generation process during denoising, making the generated realistic image more realistic in terms of the adjusted features. During the training of this realism adjustment model, a pre-trained network can be used to extract the features to be adjusted from the training samples, such as skin tone, wrinkles, and blemishes. The training samples contain images with facial features, and noise is progressively added to the images until the images are nearly completely random noise. During training, the model learns how to recover images from different levels of noise, while simultaneously increasing the detail of skin tone, wrinkles, and blemishes. To enhance detail, the model can also denoise the image at different scales, such as global denoising at a larger scale to ensure the overall structure of the image, and local denoising at a smaller scale to enhance details such as wrinkles, blemishes, and skin.

[0196] A realism adjustment model can be built based on generative adversarial networks (GANs). This model includes a generator network and a discriminator network. The generator network can be a multi-scale network, including a global network and a local detail network. The global network processes the macroscopic structure of the entire face, making the generated realistic image appear lifelike overall. The local detail network, combined with adjustments to the material, processes local areas of the face, such as around the eyes, forehead, and corners of the mouth, adding details to these areas, such as blemishes and wrinkles. The discriminator network includes a global discriminator and a local discriminator. The global discriminator is used to judge the realism of the entire generated realistic image. The local discriminator is used to judge the realism of local areas, especially areas prone to blemishes and wrinkles. During training, the global generator and global discriminator can be trained first to ensure the generated image is realistic overall. Then, based on the local generator and local discriminator, the training focuses on local details, improving the realism of wrinkles, blemishes, and skin tone. Finally, joint training is performed, combining the global and local generators with the discriminator to further optimize the model.

[0197] Besides the methods mentioned above, other methods such as autoregressive models can also be used to construct realism adjustment models, which will not be elaborated here. Furthermore, the aforementioned realism adjustment models can further include an image post-processing component. Image post-processing can further adjust the color of certain facial feature pixels in the image for precise adjustments.

[0198] After facial feature adjustment, consistency processing and fusion processing are performed on the first digital objects contained in multiple second video frames and their corresponding first video frames. The consistency processing includes at least one of the following: pose consistency processing, style consistency processing, or background consistency processing. The fusion processing involves stitching together the facial image in the second video frame with the non-facial image of the first digital object contained in the first video frame. The following describes the implementation of consistency processing using a single video frame as an example.

[0199] The following are some examples of implementation methods for pose consistency processing.

[0200] A mapping relationship is established between the first and second keypoints. The first keypoint is the facial keypoint of the second digital object contained in each second video frame, and the second keypoint is the facial keypoint of the first digital object contained in each first video frame. The mapping relationship is the correspondence between the position of the first keypoint in each second video frame and the position of the second keypoint in each corresponding first video frame. Based on the mapping relationship, the positions of the first keypoints in their respective second video frames are aligned according to the positions of the second keypoints. Specifically, the first keypoints of the facial features of the second digital object contained in the second video frame, and the second keypoints of the facial features of the first digital object contained in the first video frame, are detected, for example, using the Dlib library for keypoint detection. The first or second keypoints, such as specific positions of facial features like the eyes, nose, and mouth, can be 68 or 81 keypoints. 81 keypoints provide additional, finer facial details compared to 68 keypoints. After detecting the first and second keypoints, they are matched to construct the mapping relationship. Affine transformations can be used for keypoint alignment. By combining this mapping relationship, the positions of the second keypoint are synchronized with those of the first keypoint, and facial features are replaced. In this way, the poses of facial features in the resulting video frames are consistent.

[0201] A 3D reconstruction is performed on the facial features of the second digital object in the second video frame, and also on the facial features of the first digital object in the first video frame. During the 3D reconstruction, the depth, contour, and structure of the face are captured. A trained deep neural network can be used to directly predict the 3D structure of the face from the image. Head pose estimation is performed on the two 3D models obtained from the 3D reconstruction to obtain parameters such as rotation and translation. Alignment is then performed using the pose estimation results to ensure that the two 3D models have the same rotation angle and position. Furthermore, the closest point pairs between the two 3D models can be selected, and the error between these points can be calculated. This error is then used to optimize the rotation and translation parameters, minimizing the error for further alignment. After alignment, the facial texture of the second digital object in the second video frame can be mapped onto the 3D model of the first digital object in the first video frame. After rendering, the 3D model is rendered onto a 2D image to obtain the second video frame with pose consistency processing. This method achieves high-precision pose consistency processing.

[0202] Besides using the above methods to achieve pose consistency processing, other methods can also be used, which are not limited here.

[0203] The following is an example of how to implement style consistency processing: Adjust the style of the first digital object in the first video frame based on the style of its facial features in the second digital object; where style includes lighting or skin tone. The overall style of the first digital object in the first video frame can be adjusted according to the style of its facial features in the second video frame. For example, if the skin of the second digital object's face in the second video frame is a cool tone, then the skin of the limbs, neck, and other parts of the first digital object in the first video frame can also be adjusted to a cool tone. Performing this processing on each frame of the first video ensures style consistency across multiple second video frames. Alternatively, Gaussian blurring can be applied to each frame of the first video frame to ensure style consistency in the resulting second video frames. Other methods can also be used to achieve style consistency processing, which are not limited here.

[0204] The following example illustrates how background consistency processing is implemented: eliminating color differences between the foreground and background regions. The foreground region is the area containing the first digital object in each first video frame, and the background region is the area excluding the foreground region in each first video frame. Color differences may exist between the digital object and the background region in the first video frame. These color differences may be caused by the aforementioned style consistency processing, may be pre-existing, or may arise during other processing. To maintain background consistency, these color differences can be eliminated. Methods for eliminating color differences include color matching between the background and foreground, applying a color gradient to the junction area between the foreground and background, or feathering the foreground edges. Other methods can also be used for background consistency processing, which are not limited here.

[0205] Step 204. Obtain the second video. The multiple second video frames in the second video each correspond to multiple first video frames. Each second video frame contains a second digital object, which is a first digital object whose facial features have been processed.

[0206] In the process of processing the facial features of the first digital object in the first video, all first video frames containing the first digital object can be processed separately, or only a portion of the first video frames containing the first digital object can be selected and processed separately. The processed first video frames and the unprocessed video frames in the first video are then combined to obtain the second video.

[0207] When processing multiple video frames, one frame can be selected, and the first digital object in that frame can be processed using the facial features of the target object to obtain a second video frame containing a second digital object. Then, the first digital object in the first video frame can be adjusted using the second digital object in the second video frame, and combined with the unprocessed video frames from the first video to obtain the second video. Alternatively, for each first video frame, the facial features of the first digital object in that first video frame can be adjusted according to the facial features of the target object to obtain a second video frame. This second video frame can then be combined with the unprocessed video frames from the first video to obtain the second video. If no unprocessed video frames exist in the first video, then the processed second video frames can be combined.

[0208] like Figure 4 As shown, Figure 4 This is an exemplary schematic diagram of selecting a first video frame for facial feature processing provided in an embodiment of this application.

[0209] After acquiring the first image and the first video, a first video frame is selected from multiple video frames of the first video. Based on the facial features of the target object, the first digital object in the first video frame is processed to obtain a second video frame containing a second digital object. Then, the second digital object is fused with the first digital object from the multiple video frames to obtain the second video. When selecting the first video frame from the multiple video frames of the first video, a frame can be selected randomly or based on specific conditions. For example, a video frame where the target object's face is facing forward can be selected. Processing the first digital object in the first video frame to obtain the second video frame containing the second digital object mainly involves processing the facial features of the first digital object with reference to the facial features of the target object. This process can involve transferring overall facial features, transferring facial feature attributes, or fusing facial features. Specific implementation methods can refer to the facial feature processing methods described above.

[0210] like Figure 5 As shown, Figure 5 This is an exemplary schematic diagram of processing a first video frame in conjunction with a first image, provided in an embodiment of this application.

[0211] When processing the first video frame, a generative model can be used. The generative model includes several sub-generative models: a face replacement model, a realism adjustment model, and a consistency processing model. The face replacement model replaces facial features from the first image into the first video frame of the first video, generating a video frame with replaced facial features. This video frame is then input into the realism adjustment model, which adjusts the realism of local facial features and generates an adjusted image. The adjusted image is then input into the consistency processing model, along with multiple consecutive video frames. The consistency processing model integrates the facial features from the adjusted image into the digital objects of the multiple consecutive video frames, performing consistency processing to generate the second video. The second video generated by the generative model reflects the details of the facial features in the first image, resulting in a highly realistic second digital object.

[0212] The main steps of the video processing method according to the embodiments of this application have been described above. The following will further elaborate on the above video processing method in combination with several different scenarios.

[0213] like Figure 6 As shown, Figure 6 This is a schematic diagram of an embodiment of a video processing method in an edge-cloud combined scenario provided in this application, the method including:

[0214] 601. The end-side device sends a first image and a second image to the cloud-side device. Correspondingly, the cloud-side device receives the first image and the second image.

[0215] In edge-cloud integrated video processing scenarios, users can customize the overall image and facial details of digital objects in the generated digital object video on their edge devices. Performance-critical processing steps are executed on cloud devices, which can be relatively low-performance electronic devices such as smartphones, virtual reality devices, or tablets. Users can upload images and process them in the cloud to obtain highly realistic digital object videos.

[0216] The edge device can receive a first image and a second image uploaded by the user through a graphical interface. The first image is, for example, a portrait of a real person, including features of the human body from the shoulders up. Each of the multiple second images contains a full-body photograph of the same person from one angle; for example, two photographs might be a full-body photograph of the person taken from a frontal angle and a full-body photograph taken from a side angle. Alternatively, the first image can also be an image of other objects with facial features, such as an image of a pet; this is not a limitation.

[0217] The end-side device can determine whether the facial features of the target object in the first image meet the requirements (or determine whether the first image meets the constraints) in the following way:

[0218] When a user uploads the first image, a prompt message can be displayed on the device's screen, prompting the user to upload a photo with rich details. After receiving the uploaded photo, the device assumes that the facial features of the target object in the photo meet the requirements. This method can quickly obtain the first image.

[0219] The edge device can also receive multiple second images uploaded by the user. When receiving these images, the edge device can display a second prompt, suggesting the user upload multiple full-body photos from different angles, such as front, side, and back views. In the full-body photos, the person can be in a T-pose or A-pose, with no obstructions to any part of the body, and the photo should have uniform brightness, thus enhancing the realism of the constructed digital human. After receiving the second images uploaded by the user, the edge device can send them to the cloud device to generate the digital human without needing to determine if they meet the requirements. If the cloud device reports that it cannot generate a digital human because the second image does not meet the requirements, it prompts the user to re-upload the second images. Alternatively, the edge device can also first evaluate whether the second images meet the requirements.

[0220] The edge device can call the application programming interface (API) provided by the cloud device to upload a first image and a second image. Correspondingly, the cloud device receives the first and second images uploaded by the edge device through the same API.

[0221] 602. The cloud-based device constructs a three-dimensional model of the first digital object based on the second image.

[0222] Each second image corresponds to a full-body photograph of the human body from a specific angle, and multiple second images include full-body photographs from front, side, and back angles. The cloud-based device constructs a digital representation using full-body photographs from front, side, and back angles.

[0223] 603. The cloud-based device renders the three-dimensional model of the first digital object to obtain the first video.

[0224] The cloud-based device renders the digital human to obtain a video of the digital human that requires further processing.

[0225] 604. The cloud-based device processes the facial features of the first digital object in the first video based on the facial features of the target object in the first image to obtain the second video.

[0226] like Figure 7 As shown, Figure 7 This is an exemplary schematic diagram illustrating the processing of a digital human video according to an embodiment of this application. In this processing, sections 701-703 describe the process of generating the digital human video, and sections 704-707 describe the process of processing the digital human video to obtain a second video.

[0227] 701. Obtain multi-angle full-body photos and portraits uploaded by users.

[0228] Figure 7 The multi-angle full-body photos shown include frontal and side views. Portrait photos cover the body from the shoulders up.

[0229] 702. Reconstruct a 3D digital human based on multi-angle full-body photographs.

[0230] 703. Generate a digital human video based on the 3D digital human.

[0231] 704. Extract a single frame of human image from a digital human video, as well as multiple consecutive frames of human image.

[0232] The digital human video contains multiple consecutive frames of human images. By identifying whether a face exists in each video frame, it is determined to be a human image; otherwise, it is not. After identifying the human images, those that are temporally consecutive in the video are extracted as multiple consecutive frame images. These multiple consecutive frame images can then be used in a consistency processing model to perform consistency processing with the processed facial features, thereby altering the facial features of the digital human in the video frames.

[0233] To help users identify the desired adjustment materials for further adjustments to the digital human's facial features in the video, a portrait frame can be extracted from the video frame. The digital human in this portrait frame is then replaced with the face in the portrait photo and displayed to the user. At this point, the user can intuitively see the initial effect of the digital human after the face replacement is more realistic. Based on this, the user can select skin tone, wrinkles, etc., based on the digital human with the replaced face to make the digital human's facial features more in line with their needs.

[0234] When extracting a single frame of a human face, multiple consecutive frames may contain images from different angles. To obtain complete facial features for subsequent processing, a single frame from a frontal view can be extracted from the digital human video. By selecting a frontal view frame, the overall facial features are fully captured, allowing users to easily preview the initial results of the face replacement.

[0235] Here's an example of how to extract a frontal view portrait: Read the digital human video frame by frame, acquire the image of each frame, detect the face in each frame, and determine the angle of the face by using key points in the face. For example, by using the relative positions of key points such as the eyes, nose, and mouth—that is, the left and right eyes and corners of the mouth are all on the horizontal line, and the nose is centered—we can determine that the face is frontal. When an image containing a frontal face is detected, the image containing the area above the shoulders can be cropped as the extracted portrait frame, omitting features from other parts of the face.

[0236] 705. Use a portrait replacement model to replace the face of a portrait photo into a single frame of portrait.

[0237] A single frame of portrait and a portrait photo are input into a face replacement model, which then performs face replacement. The model replaces the face in the single frame of portrait with the face from the real portrait photo, resulting in a replaced face image. This replaced face image includes the face and other parts, such as the neck, hair, and shoulders. The face is derived from the real portrait photo, while the other parts are extracted from a single frame of portrait extracted from a digital human video.

[0238] 706. Uses realistic models to adjust skin tone, wrinkles, and blemishes.

[0239] The realism adjustment model can pre-set an optimized set of parameters for features such as skin tone, wrinkles, and blemishes, adjusting the resulting facial image after face replacement to further enhance the realism of these features. Furthermore, the realism adjustment model can also receive input adjustment materials and adjust the realism of skin tone, wrinkles, and blemishes according to the user's wishes. In this case, facial realism adjustment can be achieved through edge-cloud interaction.

[0240] like Figure 8 As shown, one embodiment of this application's edge-cloud interaction for achieving realistic face adjustment includes:

[0241] 801. The cloud-side device sends the adjustment material to the edge device. Correspondingly, the edge device receives the adjustment material.

[0242] Adjustments can be made to features such as blemishes, wrinkles, and skin tone, including the number, size, and color variations of blemishes; the depth, density, and directionality of wrinkles; and the basic tone, brightness, and contrast of skin tone. The cloud-based device can also transmit facial images to the edge device.

[0243] 802. End-side device displays and adjusts materials.

[0244] The edge device can display adjustment materials, such as the number of blemishes, the depth of wrinkles, and the brightness of skin tone, for users to select. To enhance the display effect and make it easier for users to intuitively determine how to select adjustment materials, the edge device can display the facial image and adjustment materials together.

[0245] 803. The end-side device receives selection information for the adjustment material.

[0246] The selection information is used to indicate the selection status of the adjustment material, such as selecting an adjustment material with skin tone as skin tone 1, number of spots as x, and wrinkle depth as y.

[0247] 804. The end-side device sends the selection information of the adjusted material to the cloud-side device. Correspondingly, the cloud-side device receives the selection information.

[0248] 805. The cloud-based device uses a realistic adjustment model to process selection information and facial images to obtain realistic portraits.

[0249] like Figure 9 As shown, Figure 9 This is an exemplary schematic diagram of the structure of the realism adjustment model provided in the embodiments of this application.

[0250] After receiving selection information, the cloud-based device inputs this selection information (skin color, wrinkles, or blemishes) as conditional information into the realism adjustment model. The model encodes this conditional information into a conditional vector, which is then mapped to a latent space. This latent space is a low-dimensional representation space used by the realism adjustment model to encode complex features and relationships when generating images. Facial image features are also encoded into the same latent space, allowing the features to be combined with the conditional information. When generating a realistic portrait, the realism adjustment model decodes the encoded conditional information and facial image features together to generate a new image as the realistic portrait. Because the user's selection information is considered during image generation, the generated realistic portrait better reflects the user's needs for realism adjustments, maintaining the realism of the generated portrait while improving the flexibility of portrait adjustments.

[0251] Furthermore, the realism adjustment model can also perform post-processing, which further adjusts the generated image according to the adjusted material. For example, it maps the color space of the generated image to the HSV color space, and further adjusts the three color components of the HSV color space according to the adjusted material, thereby further improving the matching degree with the adjusted material.

[0252] After obtaining a realistic portrait, the cloud-side device can send the portrait to the edge device, which then displays it for the user to preview. If the user needs to further select new materials for adjustment, they can do so, thereby flexibly adjusting the realism of the portrait from the user's perspective.

[0253] 707. A consistency processing model is used to perform consistency processing to obtain the second video.

[0254] A second video is obtained by performing consistency processing on multiple consecutive video frames based on realistic human images. The realistic human images output from the realism adjustment model, along with multiple consecutive frame human images, are input into the consistency processing model for consistency processing of the video and human images, resulting in the second video.

[0255] The aforementioned realistic portrait is a single image, while multiple consecutive frames exist. During fusion, each realistic portrait can be used to perform consistency processing on each frame, thereby obtaining multiple consecutive frames after consistency processing. These multiple consecutive frames after consistency processing are then combined with the original video frames in the first video that have not been processed, according to the time sequence of the video frames, to obtain the second video.

[0256] During consistency processing, pose consistency, style consistency, and background consistency can be performed sequentially. The following describes how to perform consistency processing on a single frame of a portrait from multiple consecutive frames. Furthermore, other frames of a portrait from multiple consecutive frames can also be processed in the same way:

[0257] After performing the above processing on each frame in a series of frames, multiple consecutive portrait video frames with consistency processing are obtained. By combining multiple consecutive portrait video frames with other video frames, a complete second video can be obtained.

[0258] 605. The cloud-side device sends the second video to the end-side device. Correspondingly, the end-side device receives the second video.

[0259] 606. The end-side device displays the second video.

[0260] Using the above method, the edge device only needs to provide a real portrait photo and multiple full-body photos to obtain a highly realistic digital human video.

[0261] like Figure 10 As shown, one embodiment of the video processing method provided in this application includes:

[0262] 1001. The edge device sends a first image and a description of the first video to the cloud device. Correspondingly, the cloud device receives the first image and the description of the first video.

[0263] If the user cannot provide a second image from multiple angles, or does not require the generation of a digital human based on a specific object, they can input a description of the digital human video they wish to generate on the edge device, so that the cloud device can generate the digital human video.

[0264] For example, users can input text on the device to describe the characteristics of the digital human they want to generate in the video, such as height, gender, age, body type, and clothing, as well as the actions they want the digital human to perform, such as moving limbs, speaking, or making specific facial expressions. Additionally, they can describe information about the scene in which they want the digital human to be located.

[0265] In addition, the edge device can provide several preset digital human models for users to choose from. Users can select their desired digital human model and describe the adjustments they wish to make to it. Furthermore, preset templates can be provided for the digital human's actions and the scene in which the digital human is located. Users can simply describe the content they want the digital human to say, and the cloud-based device can generate the corresponding digital human video.

[0266] 1002. The cloud-based device generates the first video based on the description of the first video.

[0267] Based on the different descriptions of the first video, the cloud-based device can select the appropriate method to generate the digital human video according to the description.

[0268] 1003. The cloud-based device processes the facial features of the first digital object in the first video based on the facial features of the target object in the first image to obtain the second video.

[0269] 1004. The cloud-side device sends the second video to the end-side device. Correspondingly, the end-side device receives the second video.

[0270] 1005. The end-side device displays the second video.

[0271] By generating the second video using the above method, users only need to provide the first image and a description of the first video, which offers great flexibility.

[0272] Reference Figure 11 The following describes the implementation method of video processing of the local first video in the embodiments of this application.

[0273] 1101. The edge device sends the first image to the cloud device. Correspondingly, the cloud device receives the first image.

[0274] If the user has no specific requirements for the generated digital human video, they can simply upload the first image, and the edge device will send the first image to the cloud device.

[0275] 1102. The cloud-side device acquires the first video from the local device.

[0276] If the edge device only provides the first image, the cloud device can obtain a default first video, such as a digital human video, from its local storage. Furthermore, the cloud device can also perform recognition on the first image to identify the gender, facial expressions, etc., of the person in the image, thereby selecting the most suitable digital human video as the basis for subsequent processing.

[0277] 1103. The cloud-based device processes the facial features of the first digital object in the first video based on the facial features of the target object in the first image to obtain the second video.

[0278] 1104. The cloud-side device sends the second video to the end-side device. Correspondingly, the end-side device receives the second video.

[0279] 1105. The end-side device displays the second video.

[0280] Using the above method, the end device only needs to provide a first image, such as a real portrait photo, to obtain a realistic digital human video, which is relatively easy to operate.

[0281] like Figure 12 As shown, Figure 12 This is a schematic diagram of a video processing method in one embodiment of this application.

[0282] 1201. The end-side device receives the first image uploaded by the user.

[0283] The device can receive the first image directly uploaded by the user. Furthermore, uploading includes importing the first image to the device, such as saving the first image to a cloud drive or downloading it from the cloud drive. Alternatively, the first image can be imported via an external device such as a portable hard drive or memory card. Or, the device can directly capture the first image.

[0284] 1202. The end-side device acquires the first video.

[0285] The first video can be obtained in several ways, such as: receiving multiple second images uploaded by the user and generating the first video based on the second images; or generating the first video based on the user's description of the first video; or determining the first video in response to a selection of a preset first video.

[0286] 1203. The end-side device processes the facial features of the first digital object in multiple video frames of the first video based on the facial features of the target object in the first image to obtain the second video.

[0287] Using the above method, the end device can process the first video locally by combining the first image and the first video to obtain the second video.

[0288] The video processing methods have been introduced above. The corresponding devices will be introduced below with reference to the accompanying drawings.

[0289] like Figure 13 As shown, this application embodiment also provides a video processing apparatus 130, including:

[0290] The first acquisition module 1301 is used to acquire a first image, the first image containing a facial image of the target object;

[0291] The second acquisition module 1302 is used to acquire a first video, which includes multiple first video frames. Each of the multiple first video frames contains a first digital object, which is a digital object of the target object.

[0292] The processing module 1303 is used to extract facial features of the target object's facial image in the first image, adjust the facial features of the first digital object contained in each of the multiple first video frames based on the facial features of the target object's facial image, and obtain a second video. The multiple second video frames in the second video correspond to multiple first video frames, and each second video frame contains a second digital object, which is a first digital object whose facial features have been processed.

[0293] The function of the video processing device can be understood in conjunction with the above method embodiments. The video processing device can be a cloud-based device or a device with certain video processing capabilities, such as a graphics workstation.

[0294] It should be noted that, in other embodiments, the first acquisition module, the second acquisition module, and the processing module can all be used to perform any step in the video processing method. The steps that the first acquisition module, the second acquisition module, and the processing module are responsible for implementing can be specified as needed. The first acquisition module, the second acquisition module, and the processing module can respectively implement different steps in the video processing method to realize all the functions of the video processing device.

[0295] like Figure 14 As shown, this application embodiment also provides a video display device 140, including:

[0296] The sending module 1401 is used to send a first image. The first image contains a facial image of the target object.

[0297] The receiving module 1402 is used to receive a second video. The second video is obtained by the video processing device in the following manner: acquiring a first video, the first video including multiple first video frames, each of the multiple first video frames containing a first digital object, the first digital object being a digital object of a target object; extracting facial features of the target object's facial image from a first image; adjusting the facial features of the first digital objects contained in each of the multiple first video frames based on the facial features of the target object's facial image; obtaining a second video, the multiple second video frames in the second video each corresponding to multiple first video frames, each second video frame containing a second digital object, the second digital object being a first digital object whose facial features have been processed.

[0298] Display module 1403 is used to display the second video.

[0299] The function of the video display device can be understood in conjunction with the above method embodiments. The video display device can be an end-side device.

[0300] The first acquisition module, the second acquisition module, the processing module, the sending module, the receiving module, and the display module can all be implemented in software or in hardware. For example, the implementation of the first acquisition module will be described below. Similarly, the implementation of the second acquisition module, the processing module, the sending module, the receiving module, and the display module can refer to the implementation of the first acquisition module.

[0301] It should be noted that in other embodiments, the sending module, receiving module, and display module can all be used to perform any step in the video processing method. The steps that the sending module, receiving module, and display module are responsible for implementing can be specified as needed. By implementing different steps in the video processing method through the sending module, receiving module, and display module, all functions of the video display device can be realized.

[0302] As an example of a software functional unit, the first acquisition module module may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the aforementioned computing instance may be one or more. For example, the first acquisition module module may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions.

[0303] This application also provides a computing device 100. For example... Figure 15 As shown, the computing device 100 includes a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other via the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 100.

[0304] Bus 102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 15 The bus 104 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 104 may include a path for transmitting information between various components of the computing device 100 (e.g., memory 106, processor 104, communication interface 108).

[0305] The processor 104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0306] Memory 106 may include volatile memory, such as random access memory (RAM). Processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0307] The memory 106 stores executable program code, which the processor 104 executes to implement the functions of the aforementioned first acquisition module, second acquisition module, and processing module (or sending module, receiving module, and display module), thereby realizing the video processing method. That is, the memory 106 stores instructions for executing the video processing method.

[0308] Alternatively, the memory 106 stores executable code, which the processor 104 executes to implement the functions of the aforementioned video processing device and video display device, thereby realizing the video processing method. That is, the memory 106 stores instructions for executing the video processing method.

[0309] The communication interface 103 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 100 and other devices or communication networks.

[0310] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0311] like Figure 16 As shown, the computing device cluster includes at least one computing device 100. The memory 106 of one or more computing devices 100 in the computing device cluster may store the same instructions for performing video processing methods.

[0312] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing video processing methods. In other words, a combination of one or more computing devices 100 can jointly execute instructions for executing video processing methods.

[0313] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the video processing device. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more modules among the first acquisition module, the second acquisition module, and the processing module (or the sending module, the receiving module, and the display module).

[0314] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN), a local area network (LAN), or similar. Figure 17 One possible implementation is shown. For example... Figure 17 As shown, the two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this possible implementation, the memory 106 in computing device 100A stores instructions for executing the functions of the first acquisition module. Simultaneously, the memory 106 in computing device 100B stores instructions for executing the functions of the second acquisition module and the processing module.

[0315] Figure 17 The connection method between the computing device clusters shown can be such that, considering the video processing method provided in this application requires processing a large amount of video data, the functions implemented by the second acquisition module and the processing module are delegated to the computing device 100B for execution.

[0316] It should be understood that Figure 17 The functions of the computing device 100A shown can also be performed by multiple computing devices 100. Similarly, the functions of the computing device 100B can also be performed by multiple computing devices 100.

[0317] This application also provides another computing device cluster. The connection relationships between the computing devices in this computing device cluster can be similarly referred to... Figure 16 and Figure 17 The connection method of the computing device cluster. The difference is that the memory 106 of one or more computing devices 100 in the computing device cluster can store the same instructions for executing video processing methods.

[0318] In some possible implementations, the memory 106 of one or more computing devices 100 in the computing device cluster may also store partial instructions for executing video processing methods. In other words, a combination of one or more computing devices 100 can jointly execute instructions for executing video processing methods.

[0319] It should be noted that the memory 106 in different computing devices 100 within the computing device cluster can store different instructions for executing some functions of the video processing system. That is, the instructions stored in the memory 106 of different computing devices 100 can implement the functions of one or more devices among the video processing device and the video display device.

[0320] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a video processing method.

[0321] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a video processing method, or instruct the computing device to perform a video processing method.

[0322] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video processing method, characterized in that, include: Obtain a first image, which contains a facial image of the target object; Acquire a first video, the first video including a plurality of first video frames, each of the plurality of first video frames containing a first digital object, the first digital object being the digital object of the target object; Extract facial features from the facial image of the target object in the first image, and adjust the facial features of the first digital object contained in each of the plurality of first video frames based on the facial features of the target object's facial image; A second video is obtained, wherein multiple second video frames in the second video each correspond to the multiple first video frames, and each second video frame contains a second digital object, wherein the second digital object is the first digital object whose facial features have been processed.

2. The method according to claim 1, characterized in that, Prior to acquiring the first video, the method further includes: Acquire multiple second images, which are images of the target object from different perspectives; The multiple second images are reconstructed in three dimensions to obtain a three-dimensional model of the first digital object. The first video is generated by rendering a three-dimensional model of the first digital object, and the first video contains dynamic images of the first digital object.

3. The method according to claim 1 or 2, characterized in that, Before obtaining the second video, the method further includes: The first digital object contained in the plurality of second video frames and the first video frame corresponding to each of the plurality of second video frames is subjected to consistency processing and fusion processing; wherein, the consistency processing includes at least one of the following: pose consistency processing, style consistency processing, or background consistency processing; the fusion processing includes stitching the facial image in the second video frame with the non-facial image of the first digital object contained in the first video frame.

4. The method according to claim 3, characterized in that, Before performing consistency processing and fusion processing on the second video frame corresponding to each first video frame and the first digital object contained in each first video frame, the method further includes: The system responds to adjustment information sent by the user, the adjustment information being used to adjust at least one of skin color and skin texture; According to the adjustment information, the facial features of the second digital object contained in the plurality of second video frames are adjusted to update the plurality of second video frames; the updated plurality of second video frames are used for the consistency processing and the fusion processing.

5. The method according to claim 3 or 4, characterized in that, The pose consistency processing includes: Establish a mapping relationship between a first key point and a second key point; wherein, the first key point is the facial key point of the second digital object contained in each second video frame, the second key point is the facial key point of the first digital object contained in each first video frame, and the mapping relationship is the correspondence between the position of the first key point in each second video frame and the position of the second key point in each corresponding first video frame; According to the mapping relationship, the positions of the first key points in the corresponding second video frames are aligned according to the positions of the second key points.

6. The method according to any one of claims 3-5, characterized in that, The style consistency processing includes: The style of the first digital object in the first video frame is adjusted based on the style of the facial features of the second digital object; wherein the style includes lighting or skin tone.

7. The method according to any one of claims 3-6, characterized in that, The background consistency processing includes: Eliminate color difference between foreground and background regions; wherein, the foreground region is the region where the first digital object is located in each first video frame, and the background region is the region in each first video frame excluding the foreground region.

8. The method according to any one of claims 1-7, characterized in that, The step of extracting facial features from the facial image of the target object in the first image, and adjusting the facial features of the first digital object contained in each of the plurality of first video frames based on the facial features of the target object's facial image, includes: Extract the facial mask of the target object in the first image, and the facial mask of the first digital object in each first video frame; The facial mask of the first digital object in each first video frame is replaced with the facial mask of the target object, and the mask gap is eliminated; wherein, the mask gap is the gap between the replaced mask and the adjacent area of ​​the mask.

9. A video processing apparatus, characterized in that, include: The first acquisition module is used to acquire a first image, wherein the first image contains a facial image of the target object; The second acquisition module is used to acquire a first video, the first video including a plurality of first video frames, each of the plurality of first video frames containing a first digital object, the first digital object being the digital object of the target object; The processing module is used to extract facial features of the target object's facial image in the first image, and adjust the facial features of the first digital object contained in each of the plurality of first video frames based on the facial features of the target object's facial image; A second video is obtained, wherein multiple second video frames in the second video each correspond to the multiple first video frames, and each second video frame contains a second digital object, wherein the second digital object is the first digital object whose facial features have been processed.

10. The apparatus according to claim 9, characterized in that, Before acquiring the first video, The first acquisition module is further configured to acquire multiple second images, wherein the multiple second images are images of the target object from different perspectives; The processing module is further configured to perform three-dimensional reconstruction on the plurality of second images to obtain a three-dimensional model of the first digital object; and to render the three-dimensional model of the first digital object to generate the first video, wherein the first video contains dynamic images of the first digital object.

11. The apparatus according to claim 9 or 10, characterized in that, Before obtaining the second video, The processing module is further configured to perform consistency processing and fusion processing on the first digital objects contained in the plurality of second video frames and the first video frames corresponding to the plurality of second video frames; wherein, the consistency processing includes at least one of the following: pose consistency processing, style consistency processing, or background consistency processing; the fusion processing includes stitching together the facial image in the second video frame with the non-facial image of the first digital object contained in the first video frame.

12. The apparatus according to claim 11, characterized in that, Before performing consistency processing and fusion processing on the second video frame corresponding to each first video frame and the first digital object contained in each first video frame. The processing module is also used to respond to adjustment information sent by the user, the adjustment information being used to adjust at least one of skin color and skin texture; The processing module is further configured to adjust the facial features of the second digital object contained in the plurality of second video frames according to the adjustment information, so as to update the plurality of second video frames; the updated plurality of second video frames are used for the consistency processing and the fusion processing.

13. The apparatus according to claim 11 or 12, characterized in that, The pose consistency processing includes: Establish a mapping relationship between a first key point and a second key point; wherein, the first key point is the facial key point of the second digital object contained in each second video frame, the second key point is the facial key point of the first digital object contained in each first video frame, and the mapping relationship is the correspondence between the position of the first key point in each second video frame and the position of the second key point in each corresponding first video frame; According to the mapping relationship, the positions of the first key points in the corresponding second video frames are aligned according to the positions of the second key points.

14. The apparatus according to any one of claims 11-13, characterized in that, The style consistency processing includes: The style of the first digital object in the first video frame is adjusted based on the style of the facial features of the second digital object; wherein the style includes lighting or skin tone.

15. The apparatus according to any one of claims 11-14, characterized in that, The background consistency processing includes: Eliminate color difference between foreground and background regions; wherein, the foreground region is the region where the first digital object is located in each first video frame, and the background region is the region in each first video frame excluding the foreground region.

16. The apparatus according to any one of claims 9-15, characterized in that, In the aspect of extracting facial features of the target object's facial image from the first image, and adjusting the facial features of the first digital object contained in each of the plurality of first video frames based on the facial features of the target object's facial image, the processing module is specifically used for: Extract the facial mask of the target object in the first image, and the facial mask of the first digital object in each first video frame; The facial mask of the first digital object in each first video frame is replaced with the facial mask of the target object, and the mask gap is eliminated; wherein, the mask gap is the gap between the replaced mask and the adjacent area of ​​the mask.

17. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-8.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a cluster of computing devices, cause the cluster of computing devices to perform the method as described in any one of claims 1-8.

19. A computer program product, characterized in that, Includes program instructions that, when run on a computing device cluster, cause the computing device cluster to perform the method as described in any one of claims 1-8.