Image processing model training method, image processing method, video processing model training method, and video processing method

By training an image processing model, predictive images are generated using reference images and pose features from multiple perspectives, solving the problem of insufficient detail in portrait video generation in existing technologies and achieving high-quality multi-view image generation.

WO2026007667A1PCT designated stage Publication Date: 2026-01-08ALIBABA (CHINA) CO LTD

Patent Information

Application Number
PCT/CN2025/100773
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-03
Filing Date
2025-06-12
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing technologies generate portrait videos using a single reference image, but the resulting video details are poor and the video quality is low. In particular, when generating from multiple perspectives, it is difficult to accurately handle image content without a perspective.

Method used

By identifying multiple reference images of the target object from different perspectives, the features of each reference image and pose image are obtained. The reference network features and target pose features are used to generate a predicted image, and the image processing model is trained to improve the accuracy and detail reproduction of the generated image.

Benefits of technology

It achieves accurate capture of the target object image from multiple perspectives, avoids the random generation of invisible areas, improves the accuracy and quality of the generated image, and ensures that every detail of the target object in the predicted image is restored.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025100773_08012026_PF_FP_ABST
    Figure CN2025100773_08012026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an image processing model training method, an image processing method, a video processing model training method, and a video processing method. The image processing model training method comprises: determining a plurality of reference images of a target object at different viewing angles, a reference pose image of the target object in each reference image, a target image of the target object, a target pose image of the target object in the target image, and a noise image corresponding to the target image; acquiring a reference image feature of each reference image, a reference pose feature of each reference pose image, and a target pose feature of the target pose image; obtaining a reference network feature on the basis of the reference image features and the reference pose features, and obtaining a target predicted feature on the basis of the reference network feature, the target pose feature, and the noise image; obtaining a predicted image on the basis of the target predicted feature; and training an image processing model on the basis of the predicted image and the target image to obtain a trained image processing model. Thus, the quality of image generation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing model training method, image processing method, video processing model training method and video processing method

[0001] The present disclosure claims priority to Chinese Patent Application No. 202410891514.4, filed on July 3, 2024, entitled "Image processing model training method, image processing method, video processing model training method and video processing method", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] Embodiments of the present disclosure relate to the technical field of computer, and particularly relate to an image processing model training method, an image processing method, a video processing model training method and a video processing method. BACKGROUND

[0003] In the implementation of portrait video generation by using a machine learning model, only a single reference image is usually supported. However, in the implementation of portrait video generation by using a single reference image, only a single-view video can be generated, or in the case of generating a multi-view video, the machine learning model needs to randomly generate or infer the image content corresponding to the view that does not exist in the reference image.

[0004] Therefore, the generated video usually has poor details and poor overall video quality. SUMMARY

[0005] Therefore, the generated video usually has poor details and poor overall video quality.

[0006] According to a first aspect of embodiments of the present disclosure, an image processing model training method is provided, comprising:

[0007] determining a plurality of reference images of a target object under different views, a reference pose image of the target object in each reference image, a target image of the target object, a target pose image of the target object in the target image, and a noise image corresponding to the target image, wherein the poses of the target object in the reference images and the target image are different;

[0008] obtaining reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature of the target pose image;

[0009] obtain a reference network feature according to the reference image feature and the reference pose feature, and obtain a target prediction feature according to the reference network feature, the target pose feature, and the noise image;

[0010] obtain a prediction image according to the target prediction feature, wherein a pose of the target object in the prediction image is determined by referring to the pose of the target object in the target pose image;

[0011] train the image processing model according to the prediction image and the target image, and obtain a trained image processing model.

[0012] According to a second aspect of the embodiments of the present disclosure, an image processing model training apparatus is provided, comprising:

[0013] an image determination module configured to determine a plurality of reference images of a target object under different perspectives, a reference pose image of the target object in each reference image, a target image of the target object, a target pose image of the target object in the target image, and a noise image corresponding to the target image, wherein the poses of the target object in the reference images and the target image are different;

[0014] a feature acquisition module configured to acquire reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature of the target pose image;

[0015] a feature obtaining module configured to obtain a reference network feature according to the reference image features and the reference pose features, and obtain a target prediction feature according to the reference network feature, the target pose feature, and the noise image;

[0016] an image obtaining module configured to obtain a prediction image according to the target prediction feature, wherein a pose of the target object in the prediction image is determined by referring to the pose of the target object in the target pose image;

[0017] a model training module configured to train an image processing model according to the prediction image and the target image, and obtain a trained image processing model.

[0018] According to a third aspect of the embodiments of the present disclosure, another image processing model training method is provided, comprising:

[0019] determining a plurality of reference images of a target object under different perspectives, a reference pose image of the target object in each reference image, a target image of the target object, a target pose image of the target object in the target image, and a noise image corresponding to the target image, and inputting the plurality of reference images, the reference pose image, the target pose image, and the noise image into an image processing model, wherein the pose of the target object in each reference image is different from the pose of the target object in the target image, and the image processing model comprises a feature extraction unit, a reference network unit, a denoising network unit, and a prediction unit;

[0020] obtaining reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature of the target pose image by using the feature extraction unit;

[0021] obtaining reference network features according to the reference image features and the reference pose features by using the reference network unit, and obtaining a target prediction feature according to the reference network features, the target pose feature, and the noise image by using the denoising network unit;

[0022] obtaining a prediction image according to the target prediction feature by using the prediction unit, wherein the pose of the target object in the prediction image is determined by referring to the pose of the target object in the target pose image;

[0023] training the image processing model according to the prediction image and the target image to obtain a trained image processing model.

[0024] According to a fourth aspect of the embodiments of the present disclosure, a video processing model training method is provided, comprising:

[0025] determining a plurality of reference images of a target object under different perspectives, a reference pose image of the target object in each reference image, a target video of the target object, a target pose image sequence of the target object in the target video, and a noise image sequence corresponding to the target video;

[0026] obtaining reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature sequence of the target pose image sequence;

[0027] obtaining reference network features according to the reference image features and the reference pose features, obtaining a candidate prediction feature sequence according to the reference network features, the target pose feature sequence, and the noise image sequence, and obtaining a target prediction feature sequence according to the candidate prediction feature sequence and a time sequence relationship between each candidate prediction feature;

[0028] obtain a prediction video according to the target prediction feature sequence, wherein a pose of the target object in the prediction video is determined by referring to a pose of the target object in the target video;

[0029] train a video processing model according to the prediction video and the target video, and obtain a trained video processing model, wherein the video processing model is determined based on the trained image processing model in the image processing model training method.

[0030] According to a fifth aspect of the embodiments of the present disclosure, a video processing model training apparatus is provided, comprising:

[0031] an image determination module configured to determine a plurality of reference images of a target object under different perspectives, a reference pose image of the target object in each reference image, a target video of the target object, a target pose image sequence of the target object in the target video, and a noise image sequence corresponding to the target video;

[0032] a feature acquisition module configured to acquire reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature sequence of the target pose image sequence;

[0033] a feature obtaining module configured to obtain reference network features according to the reference image features and the reference pose features, obtain a candidate prediction feature sequence according to the reference network features, the target pose feature sequence, and the noise image sequence, and obtain a target prediction feature sequence according to the candidate prediction feature sequence and a time sequence relationship between each candidate prediction feature;

[0034] a video obtaining module configured to obtain a prediction video according to the target prediction feature sequence, wherein a pose of the target object in the prediction video is determined by referring to a pose of the target object in the target video;

[0035] a model training module configured to train a video processing model according to the prediction video and the target video, and obtain a trained video processing model, wherein the video processing model is determined based on the trained image processing model in the image processing model training method.

[0036] According to a sixth aspect of the embodiments of the present disclosure, another video processing model training method is provided, comprising:

[0037] determining a plurality of reference images of the target object under different viewing angles, a reference pose image of the target object in each reference image, a target video of the target object, a target pose image sequence of the target object in the target video, a noise image sequence corresponding to the target video, and inputting the plurality of reference images, the reference pose image in each reference image, the target pose image sequence, and the noise image sequence into a video processing model, wherein the video processing model is determined based on the trained image processing model in the image processing model training method, and the video processing model comprises a feature extraction unit, a reference network unit, a denoising network unit, and a prediction unit;

[0038] using the feature extraction unit, obtaining reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature sequence of the target pose image sequence;

[0039] using the reference network unit, performing attention processing on the reference image features and the reference pose features to obtain reference network features, using the denoising network unit, performing spatial attention and cross-attention processing on the reference network features, the target pose feature sequence, and the noise image sequence to obtain a candidate prediction feature sequence, and using the denoising network unit, performing temporal attention processing on the candidate prediction feature sequence to obtain a target prediction feature sequence;

[0040] using the prediction unit, obtaining a predicted video according to the target prediction feature sequence, wherein the pose of the target object in the predicted video is determined by referring to the pose of the target object in the target video;

[0041] training the video processing model according to the predicted video and the target video.

[0042] According to a seventh aspect of the embodiments of the present disclosure, an image processing method is provided, comprising:

[0043] determining a reference image and a target image, obtaining a reference pose image of a reference object in the reference image according to the reference image, and obtaining a target pose image of a target object in the target image and a noise image corresponding to the target image according to the target image;

[0044] using an image processing model to perform image generation according to the reference image, the reference pose image, the target pose image, and the noise image, to obtain a generated image, wherein the image processing model is trained by the image processing model training method, the appearance of the object in the generated image is determined by the appearance of the reference object, and the action pose of the object in the generated image is determined by the action pose of the target object.

[0045] According to an eighth aspect of embodiments of the present disclosure, a video processing method is provided, comprising:

[0046] determining a reference image and a target video, obtaining a reference pose image of a reference object in the reference image according to the reference image, and obtaining a target pose sequence image of a target object in the target video and a noise sequence image corresponding to the target video according to the target video;

[0047] generating a generated video by using a video processing model according to the reference image, the reference pose image, the target pose sequence image, and the noise sequence image, wherein the video processing model is obtained by training the video processing model training method, an appearance of an object in the generated video is determined by an appearance of the reference object, and a motion pose of the object in the generated video is determined by a motion pose of the target object.

[0048] According to a ninth aspect of embodiments of the present disclosure, a computing device is provided, comprising:

[0049] a memory and a processor;

[0050] The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, which implement the steps of the image processing model training method, the video processing model training method, the image processing method, and the video processing method.

[0051] According to a tenth aspect of embodiments of the present disclosure, a computer readable storage medium is provided, which stores computer programs / instructions, which implement the steps of the image processing model training method, the video processing model training method, the image processing method, and the video processing method when executed by a processor.

[0052] According to an eleventh aspect of embodiments of the present disclosure, a computer program product is provided, comprising computer programs / instructions, which implement the steps of the image processing model training method, the video processing model training method, the image processing method, and the video processing method when executed by a processor.

[0053] This disclosure provides an image processing model training method that, by determining multiple reference images of a target object from different perspectives, can accurately capture the image of the target object from multiple perspectives, avoiding random generation of invisible areas. By obtaining the reference image features of each reference image and the reference pose features of each reference pose image, reference network features related to the reference images are obtained. Using the reference network features and the target pose features of the target pose image, the image processing model can guide the process of restoring the image from a noisy image, thereby obtaining accurate target prediction features. That is, the target prediction features include the reference image features related to the reference images and the target pose features of the target pose image. Based on the prediction image obtained from the target prediction features, the accuracy of the generated prediction image is improved. Furthermore, by accurately capturing the image of the target object from multiple perspectives and referencing the image of the target object from each perspective, every detail of the target object is restored in the prediction image, ensuring the generation quality of the target object in the prediction image. Attached Figure Description

[0054] Figure 1 is a schematic diagram of a scenario of an image processing model training method provided in an embodiment of this disclosure;

[0055] Figure 2 is a flowchart of an image processing model training method provided in an embodiment of this disclosure;

[0056] Figure 3 is a flowchart of another image processing model training method provided in an embodiment of this disclosure;

[0057] Figure 4 is a network framework diagram of an image processing model provided in an embodiment of this disclosure;

[0058] Figure 5 is a training schematic diagram of an image processing model training method provided in an embodiment of this disclosure;

[0059] Figure 6 is a flowchart of an image processing method provided in an embodiment of this disclosure;

[0060] Figure 7 is a flowchart of a video processing model training method provided in an embodiment of this disclosure;

[0061] Figure 8 is a flowchart of another video processing model training method provided in an embodiment of this disclosure;

[0062] Figure 9 is a network framework diagram of a video processing model provided in an embodiment of this disclosure;

[0063] Figure 10 is a training schematic diagram of a video processing model training method provided in an embodiment of this disclosure;

[0064] Figure 11 is a flowchart of a video processing method provided in an embodiment of this disclosure;

[0065] FIG. 12 is a structural schematic diagram of an image processing model training apparatus according to an embodiment of the present disclosure;

[0066] FIG. 13 is a structural schematic diagram of a video processing model training apparatus according to an embodiment of the present disclosure;

[0067] FIG. 14 is a structural block diagram of a computing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0068] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, the present disclosure can be practiced without the specific details, which are not described in the present disclosure, and it is understood that the scope of the present disclosure is not limited to the details of the embodiments described herein. In other instances, well-known methods associated with computing, software development, and / or data analytics have not been described in detail in order to avoid unnecessarily obscuring aspects of the present disclosure.

[0069] The terminology used in one or more embodiments of the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present disclosure. As used in one or more embodiments of the present disclosure and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in one or more embodiments of the present disclosure, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0070] It is to be understood that the terms first, second, etc. can be employed in one or more embodiments of the present disclosure to describe various information. Such information should not be limited by these terms. These terms are only used to distinguish one category of information from another. For example, without departing from the scope of one or more embodiments of the present disclosure, first can be termed second, and similarly, second can be termed first. Depending on the context, the word "if' as used herein can be interpreted to mean "when" or "in response to determining."

[0071] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal.

[0072] First, the nomenclature involved in one or more embodiments of the present disclosure is explained.

[0073] Diffusion model: a generative model based on diffusion process.

[0074] Unet network structure: a symmetrical U-shaped network structure, commonly used in deep learning.

[0075] CLIP (Contrastive Language-Image Pre-Training) model: a pre-trained neural network model, specifically a model for matching image-text correlation.

[0076] In the prior art, some researchers try to generate portrait videos from multi-view pictures, but this method requires accurate camera parameters and more multi-view pictures (more than 5), and the overall realism of the rendered images is poor.

[0077] In the present disclosure, an image processing model training method is provided, and the present disclosure also relates to another image processing model training method, an image processing method, a video processing model training method, another video processing model training method, a video processing method, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail in the following embodiments.

[0078] Referring to FIG. 1, FIG. 1 shows a scene schematic diagram of an image processing model training method according to an embodiment of the present disclosure.

[0079] Specifically, the image processing model training method is implemented by an end-side device 102 and a server 104. The end-side device 102 is configured to send a plurality of reference images of a target object under different perspectives and a target image of the target object to the server 104. Of course, in the case where a pose estimation model and an image noise adding model are deployed on the client side, the pose estimation model is used to obtain a reference pose image of the target object in each reference image and a target pose image of the target object in the target image, and the image noise adding model is used to obtain a noise image corresponding to the target image. Therefore, the end-side device 102 can be configured to send the plurality of reference images of the target object under different perspectives, the plurality of reference pose images, the target pose image, and the noise image to the server 104, and the present disclosure does not limit this.

[0080] The image processing model is trained in the server 104, and the image processing model includes an image feature extraction unit, a pose feature extraction unit, a reference network unit, and a denoising network unit. When the server 104 receives the plurality of reference images, the plurality of reference pose images, the target pose image, and the noise image sent by the end-side device 102, the plurality of reference images are input into the image feature extraction unit to obtain reference image features of each reference image, the reference pose features of each reference pose image and the target pose feature of the target pose image are obtained by using the pose feature extraction unit, the reference network unit is used to process the reference image features and the reference pose features to obtain reference network features, and the denoising network unit is used to process the reference network features, the target pose feature, and the noise image to obtain a target prediction feature. According to the target prediction feature, a prediction image is obtained, so that the image processing model is trained according to the prediction image and the target image, and the trained image processing model is deployed to the end-side device 102, or the model interface information corresponding to the trained image processing model is returned to the end-side device 102, so that the end-side device 102 can call the trained image processing model through the model interface information.

[0081] The end-side device 102 can include a browser, an APP (Application), or a web application such as an H5 (Hyper Text Markup Language 5) application, or a light application (also known as a small program, a lightweight application), or a cloud application, etc. The end-side device can be developed based on the software development kit (SDK) of the corresponding service provided by the server, such as based on the real-time communication (RTC) SDK development, etc. The end-side device can be deployed in an electronic device and needs to be run in dependence on the device or some APP in the device, etc. The electronic device can have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant communication tools, mailbox clients, social platform software, etc.

[0082] The server 104 can be understood as a server providing various services, including a physical server, a cloud server, for example, a server providing a communication service for multiple clients, a server for background training supporting a model used on a client, a server processing data sent by a client, and the like. It should be noted that the server 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server 104 can also be a server of a distributed system, or a server combined with a blockchain. The server 104 can also be a cloud server of a cloud service, a cloud database, cloud computing, cloud functions, cloud storage, a network service, cloud communication, middleware service, domain name service, security service, content delivery network (CDN), and a big data and artificial intelligence platform, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0083] It should be noted that the image processing model training method provided in the embodiments of the present disclosure can be executed by the server 104. In other embodiments of the present disclosure, the image processing model can be deployed in the end-side device 102, so that the end-side device 102 also has similar functions as the server 104, thereby executing the image processing model training method provided in the embodiments of the present disclosure. In other embodiments, the image processing model training method provided in the embodiments of the present disclosure can also be executed by the end-side device 102 and the server 104 together.

[0084] The image processing model training method provided in one embodiment of the present disclosure can achieve accurate capture of the image of the target object from multiple perspectives by determining multiple reference images of the target object under different perspectives, avoiding random generation of invisible areas, obtaining reference network features related to the reference images by obtaining reference image features of each reference image and reference pose features of each reference pose image, and using the reference network features and target pose features of the target pose image to guide the process of restoring the noise image to the image in the image processing model, thereby obtaining accurate target prediction features, i.e., the target prediction features include features related to the reference images in the reference image features and the target pose features of the target pose image. On the basis of obtaining the prediction image based on the target prediction features, the accuracy of the generated prediction image is improved, and by accurately capturing the image of the target object from multiple perspectives, the image of the target object under each perspective is used to restore each detail of the target object in the prediction image, ensuring the generation quality of the target object in the prediction image.

[0085] Referring to FIG. 2, FIG. 2 shows a flowchart of an image processing model training method according to an embodiment of the present disclosure, which specifically includes the following steps.

[0086] Step 202: determining a plurality of reference images of the target object under different viewing angles, a reference pose image of the target object in each reference image, a target image of the target object, a target pose image of the target object in the target image, and a noise image corresponding to the target image.

[0087] Wherein, the target object can be understood as an object capable of exhibiting diversified poses or posture performances, such as a human or a fictional animated character; the plurality of viewing angles can be understood as viewing angles of the target object from a plurality of different specific positions or directions, and the plurality of viewing angles include but are not limited to a front view, a left side view, a right side view, and a back view; the reference image can be understood as an image used as a reference in the image generation process, and the reference image includes but is not limited to providing a background reference of an image and providing an appearance reference of a target object, etc. The reference image in the embodiment of the present disclosure is used to provide an appearance reference of a target object in image generation.

[0088] The target image can be understood as a generated target of an image processing model, which is a label image corresponding to a predicted image output by the image processing model, and the target image provides a pose reference of a target object in the image generation process. For example, the pose of the target object in the target image is “crossing hands with waist”, and the pose of the target object in the generated predicted image should also be “crossing hands with waist”.

[0089] The reference pose image and the target pose image can be understood as pose images containing the action pose of the target object and not containing the appearance of the target object, which are obtained by performing pose detection on the target object in the reference image and the target image; and the target pose image can be used as a control signal in the image generation process, so that the pose of the target object in the generated predicted image tends to be consistent with the pose in the target pose image.

[0090] The noise image can be understood as an image obtained by adding noise to the target image, for example, using a DDIM model (Denoising Diffusion Implicit Models, an algorithm model for generating high-fidelity images) to add random Gaussian noise to the target image to obtain the noise image.

[0091] Specifically, a plurality of reference images of the target object under different viewing angles are determined to obtain the appearance image of the target object from multiple directions, and a label image corresponding to the generated image is obtained by determining the target image; in addition, although the target object in the reference image and the target image implies a pose, the pose and the appearance of the target object are coupled here, and the reference pose image and the target pose image are obtained separately to decouple the appearance image of the target object and the action pose, which helps to enhance the consistency of the generation result and the input pose.

[0092] The purpose of obtaining the noise image corresponding to the target image is to make the noise image as an input of the image processing model, to perform denoising processing on the noise image, and to combine the appearance of the target object in the reference image and the action posture of the target object in the target image in the process of denoising processing, so as to generate an image of the appearance of the target object in the reference image and the action posture of the target object in the target image.

[0093] In one or more embodiments of the present disclosure, the reference image and the target image can be determined first, and on this basis, the reference posture image, the target posture image and the noise image are accurately and efficiently obtained by using the machine learning model, and the specific implementation manner is as follows:

[0094] The determining of the plurality of reference images of the target object under different visual angles, the reference posture image of the target object in each reference image, the target image of the target object, the target posture image of the target object in the target image, and the noise image corresponding to the target image comprises:

[0095] The plurality of reference images of the target object under different visual angles and the target image of the target object are determined.

[0096] The plurality of reference images and the target image are respectively input into a posture estimation model to obtain the reference posture image of the target object in each reference image and the target posture image of the target object in the target image.

[0097] The target image is input into an image noise adding model to obtain the noise image corresponding to the target image, wherein the posture estimation model and the image noise adding model are machine learning models.

[0098] The posture estimation model is used to identify and extract the action posture of the object in the input image, thereby generating a posture image containing the action posture of the object but not containing the appearance features of the object. The image noise adding model is used to add noise to the input image, thereby obtaining the noise image corresponding to the input image, such as adding random Gaussian noise to the input image to obtain the noise image corresponding to the input image.

[0099] For example, in the case of the target object being a human, the posture estimation model is used, such as using ViTPose (Vision Transformer Pose, a human pose estimation method) to perform human key point detection on the human in the reference image, marking the positions of the head, shoulders, hands, knees and feet of the human, thereby obtaining the corresponding reference posture image. Similarly, the target posture image corresponding to the target image is obtained, and the image noise adding model is used, such as using the DDIM model to perform the noise adding process on the target image, and by adding random Gaussian noise, the noise image corresponding to the target image is obtained.

[0100] In actual application, the pose estimation model can realize the recognition and extraction of the action pose of the object in the input image through different kinds of pose detectors. For example, the pose detector can be a two-dimensional skeleton key point detector, the two-dimensional skeleton key point detector is used to detect the skeleton key points, and the pose image of the object is obtained through the connection lines of the skeleton key points. Alternatively, the pose detector can also be a three-dimensional action detector, the three-dimensional action detector can fit the three-dimensional pose of the object in the input image, and then the pose image of the object is obtained through rendering.

[0101] It should be noted that the purpose of obtaining the reference pose image and the target pose image in the embodiment of the present disclosure is to obtain the pose information of the target object (the pose information describes the action pose of the target object in space, that is, the information of how the target object positions the body parts, joints and the shape and direction of the whole), and the reference pose image and the target pose image are actually a specific way to obtain the pose information of the target object. With the development of technology, the pose information of the target object can be obtained by using other technical means without being limited to using the pose image.

[0102] The image processing model training method provided by the embodiment of the present disclosure can quickly and accurately obtain the reference pose image, the target pose image and the noise image based on the determination of the reference image and the target image by using the pose estimation model and the image noise adding model.

[0103] In one or more embodiments of the present disclosure, the reference images under different viewing angles can be obtained in various implementation manners, such as from a reference video of the target object, or directly by photographing the target object under different viewing angles. The specific implementation manners are as follows:

[0104] The method further includes:

[0105] determining a reference video of the target object, and obtaining a plurality of video frames of the target object under different viewing angles from the reference video as the reference images; or

[0106] photographing a plurality of reference images of the target object under different viewing angles.

[0107] Specifically, the plurality of reference images of the target object under different viewing angles can be obtained in various implementation manners, such as obtaining a plurality of video frames of the target object under different viewing angles from a reference video of the target object, determining the plurality of video frames as the plurality of reference images, or directly photographing the plurality of reference images of the target object under different viewing angles.

[0108] In the case that the target object is a fictitious animation character, the multiple reference images of the target object in different perspectives can be obtained from a design drawing of the animation character. A design drawing of an animation character usually includes a front view of the character (showing facial expression, hairstyle, clothing details and body proportion of the character), a side view of the character (emphasizing head contour, height of the nose, thickness of the body and side level of the clothing), and a back view of the character (showing back design of the character, including back style of the hairstyle, details of the back of the clothing), so that the multiple reference images in different perspectives can be obtained from the design drawing of the animation character.

[0109] The image processing model training method provided by the embodiments of the present disclosure can obtain multiple reference images of a target object through various implementation manners, so that the multiple reference images of the target object can be reasonably obtained through appropriate manners according to actual application requirements.

[0110] Step 204: obtaining reference image features of the reference images, reference pose features of the reference pose images, and target pose features of the target pose image.

[0111] On the basis of training the image processing model, reference image features corresponding to the reference images, reference pose features of the reference pose images, and target pose features of the target pose image are obtained, so that the features are processed in the image processing model.

[0112] In one or more embodiments of the present disclosure, by inputting the determined multiple reference images, reference pose images, and target pose image into the image processing model, image features corresponding to each image are obtained in the image processing model, and the specific implementation manners are as follows:

[0113] The obtaining of the reference image features of the reference images, the reference pose features of the reference pose images, and the target pose features of the target pose image includes:

[0114] The multiple reference images, the reference pose images, and the target pose image are input into the image processing model.

[0115] The image processing model extracts features from each reference image, the reference pose images, and the target pose image, and obtains reference image features of the reference images, reference pose features of the reference pose images, and target pose features of the target pose image.

[0116] Specifically, by inputting the reference image, the reference pose image, and the target pose image into the image processing model, feature extraction can be performed on these images in the image processing model, thereby obtaining reference image features of each reference image, reference pose features of each reference pose image, and target pose features of the target pose image. The reference image features include image information corresponding to the reference image, which can represent the content and structure of the image. The reference pose features are used to represent the pose information of the target object in the reference image, and the target pose features are used to represent the pose information of the target object in the target image.

[0117] The image processing model training method provided by the embodiments of the present disclosure can perform feature extraction on the reference image, the reference pose image, and the target pose image in the image processing model, so that the image processing model can perform in-depth analysis and processing based on the corresponding features, thereby improving the processing speed of the image processing model.

[0118] In one or more embodiments of the present disclosure, the image processing model includes an image feature extraction unit and a pose feature extraction unit. Image feature extraction is performed on the image in the image feature extraction unit, and pose feature extraction is performed on the image in the pose feature extraction unit. The specific implementation is as follows:

[0119] The feature extraction on the reference image, the reference pose image, and the target pose image in the image processing model includes:

[0120] The image feature extraction unit is used to perform image feature extraction on each reference image to obtain reference image features of the reference image.

[0121] The pose feature extraction unit is used to perform pose feature extraction on each reference pose image and the target pose image respectively to obtain reference pose features of each reference pose image and target pose features of the target pose image.

[0122] The image feature extraction unit can be understood as a network unit that focuses on extracting image information from an image. The image feature extraction unit can be a specific layer or a group of layers in a deep learning model, or it can exist as an independent model. When the image feature extraction unit is an independent model, it can be a pre-trained feature extraction model that is trained separately for extracting image features and can be embedded into the image processing model as a fixed feature extractor.

[0123] The pose feature extraction unit can be understood as a network unit focusing on extracting pose information from an image containing a pose. For example, the pose feature extraction unit is a 4-layer convolutional network, each layer has a convolution kernel size of 4x4 and a step size of 2x2, and the corresponding channel number is 16, 32, 64, and 128.

[0124] Specifically, the plurality of reference images are input into the image feature extraction unit, and each reference image is encoded by the image feature extraction unit to obtain the reference image features corresponding to each reference image. In the embodiment of the present disclosure, the reference image features mainly include the appearance features of the target object in the reference image. The reference pose image and the target pose image are input into the pose feature extraction unit, the pose feature extraction unit extracts the reference pose features of each reference pose image, and the pose feature extraction unit extracts the target pose features of the target pose image.

[0125] The image processing model training method provided by the embodiment of the present disclosure extracts image features from reference images based on the image feature extraction unit and extracts pose features from reference pose images and target pose images based on the pose feature extraction unit, that is, different types of images can be based on appropriate feature extraction units to more accurately obtain the key features corresponding to the images.

[0126] Step 206: obtaining reference network features according to the reference image features and the reference pose features, and obtaining target prediction features according to the reference network features, the target pose features, and the noise image.

[0127] Specifically, the reference network features are used to enhance the realism, detail richness, and similarity to the reference image in the image generation process. In the case where the reference network features are obtained based on the reference image features and the reference pose features, the reference network features contain rich image information of the reference image and pose information of the target object in the reference image. The image information includes but is not limited to color distribution, texture, object shape, and the like of the reference image. The pose information of the target object includes but is not limited to the position and direction of the target object.

[0128] By combining the reference network features with the target pose features and the noise image, the obtained target prediction features contain the image information of the reference image in the reference network features and the pose information in the target pose features.

[0129] In one or more embodiments of the present disclosure, the image processing model further comprises a reference network unit and a denoising network unit, wherein the reference network unit and the denoising network unit have the same network structure and share the same initial weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer;

[0130] The reference network feature is obtained according to the reference image feature and the reference pose feature, and the target prediction feature is obtained according to the reference network feature, the target pose feature and the noisy image, comprising:

[0131] The reference network feature is obtained by performing spatial attention processing on the reference image feature and the reference pose feature by using the spatial attention layer of the reference network unit;

[0132] The target prediction feature is obtained by performing spatial attention processing on the reference network feature, the target pose feature and the noisy image by using the spatial attention layer of the denoising network unit.

[0133] The reference network unit is used to process the reference images to obtain the related information of the reference images, and the related information of the reference images can be transmitted to the denoising network unit; the denoising network unit is the main network for generating images, and is used to denoise the noisy image to generate a prediction image.

[0134] The spatial attention mechanism of the spatial attention layer can give different weights to different positions of the input data, thereby emphasizing the regions more important to the image generation task. For example, when generating a prediction image of a target object in a front view, the spatial attention layer of the denoising network unit can increase the weight of the reference image in which the target object is in a front view among multiple reference images. In the case that the spatial attention layer can dynamically adjust its attention focus according to different inputs, the image processing model can automatically adapt to the changes of key features in different scenes. The spatial attention layer dynamically adjusts the attention degree to different parts of the input data, which can realize more efficient and more accurate information processing.

[0135] In actual application, the reference network unit and the denoising network unit can both be Unet network structures. On the basis that the reference network unit and the denoising network unit have the same network structure and share the initialized weights, the reference network features of the reference network unit can be integrated into the denoising network unit through the spatial attention, and in the case that the reference network features are obtained through the reference image features corresponding to the plurality of reference images and the reference pose features corresponding to the plurality of reference pose images, the denoising network unit can refer to the appearance of the target object in the reference images under a plurality of viewing angles in the denoising process, thereby avoiding the generation of the invisible area by imagination when using one reference image as a reference, and reducing the accuracy of the generated image.

[0136] Specifically, when processing the plurality of reference images, the reference pose features corresponding to the reference images are also input into the reference network unit, and after the spatial attention processing by the spatial attention layer of the reference network unit, the reference image features and the reference pose features contained in each reference network feature in the reference network features are focused on more important parts.

[0137] In the case that the reference network features are obtained, the reference network features are input into the denoising network unit, the target pose features and the noise image are added in feature, the denoising network features are obtained, and the reference network features and the denoising network features are fused in feature, for example, the reference network features and the denoising network features are spliced in the width dimension, and the spliced features are input into the spatial attention layer of the denoising network unit. In the case that the denoising network features also contain pose information, based on the spatial attention layer of the denoising network unit, the denoising network unit can adaptively assign different weights to the plurality of reference images according to the correlation between the poses and the correlation between the pictures, so as to obtain appropriate target prediction features.

[0138] For example, in the plurality of reference images, there is one reference image in which the target object is front-facing and has both hands on the waist, one reference image in which the target object is front-facing and has one hand on the waist, and one reference image in which the target object is side-facing and has both hands on the waist; on the basis of aiming to generate one prediction image in which the target object is side-facing and has one hand on the waist, the reference image in which the target object is front-facing can be given a lower weight, while the reference image in which the target object is side-facing can be given a higher weight, and at the same time, the reference image in which the target object has one hand on the waist can be given a higher weight than the reference image in which the target object has both hands on the waist.

[0139] The feature maps of each reference image are fused according to the assigned weights to obtain target prediction features, and the target prediction features will contain information from all the reference images, but focus on the features that are more relevant and more important to the prediction image.

[0140] The image processing model training method provided by the embodiments of the present disclosure integrates the reference network features of the reference network unit into the denoising network unit, so that the denoising network unit can adaptively assign weights to the reference network features of each reference image according to the spatial attention mechanism of the spatial attention layer, and obtain more accurate target prediction features.

[0141] In one or more embodiments of the present disclosure, the image feature extraction unit includes a first image feature extraction layer and a second image feature extraction layer; image features of each reference image are extracted by different image feature extraction units, features of the reference image are extracted from multiple directions, and details of the reference image are better preserved in the generated image, and the specific implementation is as follows:

[0142] The image feature extraction unit extracts image features of each reference image, and obtains reference image features of the reference image, including:

[0143] The first image feature extraction layer extracts image features of each reference image, and obtains first reference image features of the reference image.

[0144] The second image feature extraction layer extracts image features of each reference image, and obtains second reference image features of the reference image, wherein the second reference image features include semantic features of the reference image.

[0145] The first image feature extraction layer can be understood as an encoder layer, which extracts image features of the reference image; and the second image feature extraction layer can be understood as a CLIP feature extraction layer, that is, the second image feature extraction layer extracts features through a CLIP model, and the second reference image features obtained by extraction can be understood as CLIP features, and the second reference image features include semantic features of the reference image.

[0146] Specifically, when multiple reference images are input into the image processing model and image features are extracted by the image feature extraction unit, image features of the reference image can be extracted by multiple image feature extraction layers, so that different levels of feature representations of the image can be learned in the case that different image feature extraction layers are trained by different training methods.

[0147] The image processing model training method provided by the embodiments of the present disclosure can extract image features of the reference image from multiple directions and obtain different levels of reference image features in the case that different image feature extraction layers are used to extract image features of the reference image, so that the generated prediction image can better preserve the details of the reference image from multiple dimensions, and the accuracy of the prediction image is improved.

[0148] In one or more embodiments of the present disclosure, the image processing model further comprises a reference network unit and a denoising network unit, wherein the reference network unit and the denoising network unit have the same network structure and share the same initial weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer and a cross-attention layer; in the cross-attention layer, the features output by the spatial attention layer of the denoising network unit are cross-attentively processed with the second reference image features, and the specific implementation is as follows:

[0149] The reference network features are obtained according to the reference image features and the reference pose features, and the target prediction features are obtained according to the reference network features, the target pose features and the noisy image, comprising:

[0150] The spatial attention layer of the reference network unit is used to perform spatial attention processing on the first reference image features and the reference pose features to obtain the reference network features;

[0151] The spatial attention layer of the denoising network unit is used to perform spatial attention processing on the reference network features, the target pose features and the noisy image to obtain target key features;

[0152] The cross-attention layer of the denoising network unit is used to perform cross-attention processing on the second reference image features and the target key features to obtain the target prediction features.

[0153] The target key features can be understood as the features output by the spatial attention layer of the denoising network unit, and the target key features are the features of the key part of the reference network features selected by the spatial attention mechanism of the spatial attention layer. The cross-attention mechanism is a bidirectional attention mechanism that establishes a connection between two input sequences to obtain a better representation.

[0154] Specifically, the first reference image features are input into the spatial attention layer of the reference network unit to perform spatial attention processing on the first reference image features and the reference pose features by the spatial attention layer of the reference network unit to obtain the reference network features, and the reference network features are integrated into the denoising network unit. The reference network features, the target pose features and the noisy image are spatially attentively processed in the spatial attention layer of the denoising network unit to obtain target key features. The processing in the spatial attention layer of the denoising network unit is similar to the above-mentioned embodiments, and will not be described here.

[0155] The target key feature output by the spatial attention layer in the denoising network unit and the second reference image feature are input into the cross attention layer of the denoising network unit, so as to perform cross attention processing on the second reference image feature and the target key feature by the cross attention layer of the denoising network unit, and obtain a target prediction feature, which contains richer image features of the reference image.

[0156] The image processing model training method provided in the embodiments of the present disclosure can make the target prediction feature obtained by performing cross attention processing on the second reference image feature and the target key feature output by the spatial attention layer of the denoising network unit contain richer image features in the reference image, so that the generated prediction image can better retain the details of the reference image from multiple dimensions, and the accuracy of the prediction image is improved.

[0157] In one or more embodiments of the present disclosure, the reference network unit includes a plurality of spatial attention layers and a plurality of cross attention layers; one spatial attention layer and one cross attention layer can be regarded as a group of layers, the reference network feature is obtained by each spatial attention layer of the reference network unit, and the second reference image feature and the reference network feature are cross-attention processed by each cross-attention layer, and the specific implementation manner is as follows:

[0158] The spatial attention processing of the first reference image feature and the reference pose feature by the spatial attention layer of the reference network unit to obtain the reference network feature includes:

[0159] The first reference image feature and the reference pose feature are fused to obtain a reference fusion feature, and the reference fusion feature is spatially attention processed by each spatial attention layer of the reference network unit to obtain the reference network feature.

[0160] After the spatial attention processing of the first reference image feature and the reference pose feature by the spatial attention layer of the reference network unit to obtain the reference network feature, the method further includes:

[0161] The second reference image feature and the reference network feature are cross-attention processed by each cross-attention layer of the reference network unit to obtain a reference cross feature.

[0162] The reference cross feature is determined as the reference fusion feature, and the step of spatially attention processing the reference fusion feature by each spatial attention layer of the reference network unit to obtain the reference network feature is continued.

[0163] The denoising network unit comprises a plurality of spatial attention layers and a plurality of cross-attention layers.

[0164] The spatial attention layers of the denoising network unit are used to perform spatial attention processing on the reference network feature, the target pose feature and the noisy image to obtain a target key feature, comprising:

[0165] The target pose feature and the noisy image are fused to obtain a denoising fusion feature, and the reference network feature and the denoising fusion feature are spatially attention processed at each spatial attention layer of the denoising network unit to obtain the target key feature.

[0166] The cross-attention layers of the denoising network unit are used to perform cross-attention processing on the second reference image feature and the target key feature to obtain the target prediction feature, comprising:

[0167] At each cross-attention layer of the denoising network unit, the second reference image feature and the target key feature are cross-attention processed to obtain the target prediction feature, and the target prediction feature is determined as the denoising fusion feature, and the step of spatially attention processing the reference network feature and the denoising fusion feature at each spatial attention layer of the denoising network unit to obtain the target key feature is continued.

[0168] In actual application, the first reference image feature and the reference pose feature are fused to obtain a fused reference fusion feature, which is input into the first spatial attention layer of the reference network unit. The reference fusion feature is spatially attention processed at the first spatial attention layer of the reference network unit to obtain a reference network feature. The target pose feature and the noisy image are fused to obtain a denoising fusion feature. The reference network feature obtained by using the first spatial attention layer of the reference network unit is feature-spliced with the denoising fusion feature in the width dimension (specifically, in the case that the reference network unit and the denoising network unit are both Unet structures, the features are transmitted in the form of feature maps in the reference network unit and the denoising network unit, and the feature maps have three different dimensions of height, width and channel (also referred to as depth). By feature-splicing in the width dimension, it is ensured that the spliced feature maps remain consistent in the height dimension, and the sum of the widths of all input feature maps in the width dimension). The spliced feature is input into the first spatial attention layer of the denoising network unit, and the spliced feature is spatially attention processed to obtain a target key feature.

[0169] The second reference image feature and the target key feature output by the first layer spatial attention layer in the denoising network unit are input into the first layer cross-attention layer in the reference network unit, and the first cross-attention layer of the denoising network unit is used to perform cross-attention processing on the second reference image feature and the target key feature, to obtain a target prediction feature; in order to ensure the consistency of the reference network unit and the denoising network unit, the second reference image feature and the reference network feature output by the first layer spatial attention layer in the reference network unit are input into the first layer cross-attention layer in the reference network unit, and the reference cross-feature output by the first layer cross-attention layer in the reference network unit is used as the input feature of the second spatial attention layer in the reference network unit, that is, the above-mentioned reference fusion feature, and so on, the reference fusion feature is subjected to spatial attention processing in each spatial attention layer of the reference network unit, to obtain the reference network feature, and the second reference image feature and the reference network feature are subjected to cross-attention processing in each cross-attention layer of the reference network unit, to obtain the reference cross-feature.

[0170] Similarly, in the denoising network unit, the target prediction feature output by the current cross-attention layer is determined as a denoising fusion feature, and the denoising fusion feature is used as the input of the next spatial attention layer, so as to perform spatial attention processing on the reference network feature and the denoising fusion feature in each spatial attention layer of the denoising network unit, to obtain the target key feature, and perform cross-attention processing on the second reference image feature and the target key feature in each cross-attention layer of the denoising network unit, to obtain the target prediction feature.

[0171] It should be noted that, although the spatial attention layers of the reference network unit are all used to process the reference fusion feature, and the spatial attention layers of the denoising network unit are all used to process the spliced feature of the reference network feature and the denoising fusion feature, since the reference network unit and the denoising network unit are Unet network structures, the features in the Unet are subjected to up-sampling and down-sampling processing, therefore, the feature sizes of the reference fusion features processed in the spatial attention layers of each reference network unit are different, and similarly, the feature sizes of the spliced features processed in the spatial attention layers of each denoising network unit are different.

[0172] The image processing model training method provided in the embodiments of the present disclosure integrates the reference network features of each spatial attention layer in the reference network unit into each spatial attention layer of the corresponding denoising network unit, which can improve the accuracy of the target prediction feature.

[0173] Step 208: obtaining a predicted image according to the target prediction feature, wherein the pose of the target object in the predicted image is determined by referring to the pose of the target object in the target pose image.

[0174] The predicted image can be understood as an image generated by the image processing model, in which the pose of the target object tends to be the pose of the target object in the target image, the appearance of the target object tends to be the appearance of the target object in the reference image, and the background and other regions of the predicted image tend to be the corresponding regions of the reference image.

[0175] In one or more embodiments of the present disclosure, the image processing model further comprises a prediction unit; the prediction unit is used to decode and process the target prediction feature, so as to generate a predicted image, and the specific implementation manner is as follows:

[0176] The predicted image is obtained according to the target prediction feature, and the obtaining comprises:

[0177] The prediction unit is used to decode and process the target prediction feature, so as to obtain a predicted image.

[0178] Specifically, the input of the reference network unit and the denoising network unit is the feature extracted from the image or the pose image by the image feature extraction unit and the pose feature extraction unit, so that the output of the reference network unit and the denoising network unit is the target prediction feature of the same dimension as the feature. In order to obtain the predicted image by the target prediction feature, the target prediction feature needs to be decoded and processed, so as to obtain the predicted image.

[0179] The image processing model training method provided by the embodiments of the present disclosure is used to reconstruct the predicted image with the same dimension as the target image by the target prediction feature, decode and process the target prediction feature, and thus reasonably obtain the predicted image.

[0180] Step 210: training the image processing model according to the predicted image and the target image, to obtain a trained image processing model.

[0181] In one or more embodiments of the present disclosure, the training of the image processing model according to the predicted image and the target image, to obtain a trained image processing model, comprises:

[0182] The loss function of the predicted image and the target image is determined, and the image processing model is trained according to the loss function, to obtain a trained image processing model.

[0183] Specifically, in the case of obtaining the predicted image, the image processing model is trained by the predicted image and the target image to obtain the trained image processing model, such as training the image processing model by calculating the loss function (mean square error, cross-entropy loss, etc.) of the predicted image and the target image; it should be noted that the training of the image processing model is mainly to adjust the network parameters of the pose feature extraction unit, the reference network unit and the denoising network unit of the image processing model.

[0184] In actual application, after training the image processing model, the image processing model can be deployed on the client, so that the client can generate images locally using the trained image processing model, or in the case of sending the model interface information corresponding to the image processing model to the client, the client can save the computing resources, the client interacts with the image processing model deployed on the server through the model interface information, and the image processing model can send the generated image to the client after generating the image.

[0185] The image processing model training method provided by one embodiment of the present disclosure can achieve accurate capture of the image of the target object from multiple reference images by determining multiple reference images of the target object under different viewing angles, avoiding random generation of invisible areas, and obtaining reference image features of each reference image, reference pose features of each reference pose image, and target pose features of the target pose image, so as to adaptively select these features in the image processing model, thereby obtaining accurate target prediction features, improving the accuracy of the generated predicted image based on the target prediction features, and restoring each detail of the target object in the predicted image.

[0186] Referring to FIG. 3, FIG. 3 shows a flowchart of another image processing model training method provided by one embodiment of the present disclosure, which specifically includes the following steps.

[0187] Step 302: determining multiple reference images of a target object under different viewing angles, reference pose images of the target object in each reference image, a target image of the target object, a target pose image of the target object in the target image, and a noise image corresponding to the target image, and inputting the multiple reference images, each reference pose image, the target pose image, and the noise image into an image processing model.

[0188] Wherein, the poses of the target object in the reference images and the target image are different, and the image processing model includes a feature extraction unit, a reference network unit, a denoising network unit, and a prediction unit.

[0189] Step 304: acquiring, by using the feature extraction unit, reference image features of the reference images, reference pose features of the reference pose images, and target pose features of the target pose image.

[0190] Step 306: obtaining, by using the reference network unit, reference network features according to the reference image features and the reference pose features, and obtaining, by using the denoising network unit, target prediction features according to the reference network features, the target pose features, and the noisy image.

[0191] Step 308: obtaining, by using the prediction unit, a prediction image according to the target prediction features, wherein a pose of a target object in the prediction image is determined by referring to a pose of the target object in the target pose image.

[0192] Step 310: training the image processing model according to the prediction image and the target image, to obtain a trained image processing model.

[0193] The specific implementation can refer to the above embodiments, which will not be described here.

[0194] The image processing model training method provided by one embodiment of the present disclosure can achieve accurate capture of the image of the target object from multiple reference images by determining multiple reference images of the target object under different viewing angles, avoiding random generation of invisible areas, and obtaining reference image features of the reference images, reference pose features of the reference pose images, and target pose features of the target pose image, so as to adaptively select these features in the image processing model, thereby obtaining accurate target prediction features. On the basis of obtaining the prediction image based on the target prediction features, the accuracy of the generated prediction image is improved, and each detail of the target object is restored in the prediction image.

[0195] Referring to FIG. 4, FIG. 4 shows a network framework diagram of an image processing model according to one embodiment of the present disclosure. According to the network framework diagram, the training of the image processing model is achieved by the following content.

[0196] Specifically, multiple reference images of the target object under different viewing angles are determined, and a target image of the target object is determined. In the case that the object poses of the target object in the reference images are different from the object pose of the target object in the target image, the target object performs action display in the action pose in the target image, and the appearance of the target object is in the appearance image in the multiple reference images. A prediction image of the target object is generated, in which the appearance image is in the reference image, and the action pose is the action pose of the target object in the reference image.

[0197] In actual application, on the basis of determining the target image, the target image can be denoised by using an image denoising model to obtain a denoised noise image (i.e., the noise in FIG. 4), such as using a DDIM model (Denoising Diffusion Implicit Models) to perform the denoising process of the target image, and a noise image corresponding to the target image can be obtained by adding random Gaussian noise.

[0198] The image processing model comprises a pose feature extraction layer, which takes a pose image as input and outputs a pose feature corresponding to the pose image. Therefore, before processing by the pose feature extraction layer, a reference pose image corresponding to the reference image and a target pose image corresponding to the target image are obtained. That is, the image processing model comprises a pose estimation unit, which detects the pose of the object in the input image and constructs a corresponding pose image.

[0199] That is, when the input of the image processing model is the target image, the noise image of the target image, and the plurality of reference images, the target image and the plurality of reference images are input into the pose estimation unit of the image processing model, a plurality of reference pose images corresponding to the plurality of reference images and a target pose image corresponding to the target image (i.e., a pose image corresponding to the target image) are constructed, and the plurality of reference pose images and the target pose image are input into the pose feature extraction layer to obtain a plurality of reference pose features corresponding to the plurality of reference pose images and a target pose feature corresponding to the target pose image.

[0200] Of course, in actual application, the pose estimation unit can also be an independent pose estimation model independent of the image processing model. In this case, the plurality of reference images and the target image are input into the pose estimation model to construct a plurality of reference pose images corresponding to the plurality of reference images and a target pose image corresponding to the target image, and the target pose image, the noise image of the target image, the plurality of reference images, and the plurality of reference pose images are input into the image processing model to obtain a plurality of reference pose features corresponding to the plurality of reference pose images and a target pose feature corresponding to the target pose image by using the pose feature extraction layer of the image processing model.

[0201] Specifically, the image processing model comprises two image processing units, a reference image processing unit (i.e., the reference network unit in the above embodiment) and a denoising image processing unit (i.e., the denoising network unit in the above embodiment). The denoising image processing unit and the reference image processing unit are two Unet networks with similar structures, share initialized weights, and both comprise a spatial attention layer and a cross-attention layer. Therefore, the denoising image processing unit can learn the features associated with the same feature space from the reference image processing unit.

[0202] The first image feature extraction layer (i.e., the encoder in FIG. 4) is used to extract image features of multiple reference images, and multiple reference image features containing appearance information of the target object are obtained; the multiple reference image features are fused (e.g., added) with the multiple reference pose features extracted above, and the fused reference fusion features are input into the spatial attention layer of the reference image processing unit to obtain multiple reference network features.

[0203] The target pose feature is fused with the noise image to obtain a denoising fusion feature. After the denoising fusion feature is fused with the multiple reference network features obtained above (e.g., the multiple reference network features and the denoising fusion feature are spliced in the width dimension), the spatial attention layer of the denoising image processing unit is input. The spatial attention layer can provide an adaptive mechanism that can automatically adjust the focus according to the specific content of each sample to obtain a target key feature in which different weights are assigned to each reference network feature and denoising network feature.

[0204] In the case where the reference network features and the denoising network features both contain pose information, the spatial attention layer of the denoising image processing unit can adaptively assign different weights to the reference network features and the denoising network features according to the correlation of the pose information between the two features and the image correlation between the reference image and the target image, so that the image processing model can focus on a specific key area; for example, in the case where the target image is an image of user a with hands crossed and right inclined at an angle of 45 degrees, the weights of the features corresponding to the front view reference image, the right view reference image, and the back view reference image in the multiple reference image processing features can be increased, and the weight of the feature corresponding to the left view reference image can be reduced.

[0205] It should be noted that, in order to extract image features of multiple reference images from multiple perspectives and make the generated predicted image better retain the details of the reference image, the image processing model can include a second image feature extraction layer (i.e., the CLIP feature extraction layer in FIG. 4), which is used to extract CLIP features (i.e., the second reference image features in the above embodiments) of multiple reference images. Different image feature extraction layers capture and express image features in different ways. By extracting features of reference images through different image feature extraction layers, the diversity of image features can be increased, which helps the image processing model to capture more dimensional information.

[0206] The extracted CLIP feature and the target key feature output by the spatial attention layer in the above denoising image processing unit are input into the cross-attention layer of the denoising image processing unit to obtain a target prediction feature, and the target prediction feature is input into the next layer of spatial attention layer. For details of the training process, refer to FIG. 5, which shows a training schematic diagram of an image processing model training method according to an embodiment of the present disclosure. In the case of five reference images, there are four reference network features, that is, N is 4, w represents the width of the feature, and h represents the height of the feature.

[0207] In the case where the Unet structure includes multiple spatial attention layers (spatial attention modules) and cross-attention layers (cross-attention modules), in each spatial attention layer, multiple reference network features of the reference image processing unit are integrated into the spatial attention layer of the denoising image processing unit, and in each cross-attention layer, the CLIP feature is integrated into the cross-attention layer of the denoising image processing unit.

[0208] In actual application, to ensure the consistency of the features in the reference image processing unit and the denoising image processing unit, so that the features of the reference image processing unit can be applied to the denoising image processing unit, the network feature is also input into the cross-attention layer of the reference image processing unit in each cross-attention layer of the reference image processing unit.

[0209] The target prediction feature output by the denoising image processing unit is input into the decoder (i.e., the prediction unit in the above embodiment), and a prediction image is output by the decoder, so as to train the image processing model according to the prediction image and the target image. In the embodiment of the present disclosure, the posture feature extraction layer, the reference image processing unit and the denoising image processing unit of the image processing model are mainly trained.

[0210] The image processing model training method provided by the embodiment of the present disclosure can obtain multiple features corresponding to multiple reference images by determining multiple reference images of a target object under multiple perspectives, accurately capture the image of the target object from multiple perspectives, avoid random generation of invisible areas, and restore each detail of the target object in the generated prediction image.

[0211] Referring to FIG. 6, which shows a flowchart of an image processing method according to an embodiment of the present disclosure, the method includes the following steps.

[0212] Step 602: Determine a reference image and a target image, obtain a reference posture image of a reference object in the reference image according to the reference image, and obtain a target posture image of a target object in the target image and a noise image corresponding to the target image according to the target image.

[0213] In one or more embodiments of the present disclosure, a user interacts with a user interaction interface of a client, causing the client to send a reference image and a target image to a server, and the implementation is as follows:

[0214] Determining the reference image and the target image includes:

[0215] Receiving the reference image and the target image sent by the client, wherein the reference image and the target image are sent in the case that a specific operation is triggered on the user interaction interface of the client.

[0216] The user can pre-set several target images in the user interaction interface of the client. On the basis that the user uploads any reference image or obtains the reference image by shooting, the user can select a target image, so that the client sends the reference image and the target image to the server.

[0217] Step 604: According to the reference image, the reference pose image, the target pose image and the noise image, an image processing model is used for image generation to obtain a generated image, wherein the image processing model is obtained by training the above image processing model training method, the appearance of the object in the generated image is determined by the appearance of the reference object, and the action pose of the object in the generated image is determined by the action pose of the target object.

[0218] Specifically, taking the application of the image processing method in the animation field as an example, the user can provide any image of an animation image under different perspectives, and the images provided by the user are reference images, and a target image is selected, aiming to make the animation image display the action pose by using the pose of the object in the target image. The image processing model is inputted with the plurality of reference images provided by the user and the selected target image, and the image processing model is used to generate an image in which the animation image displays the action pose by using the pose of the object in the target image.

[0219] The image processing method provided by the embodiments of the present disclosure can provide reasonable supplements and prior knowledge for image generation based on the pre-training weights of the reference network unit and the denoising network unit, and by introducing a technology capable of processing any number of reference image inputs, key information can be adaptively extracted from the reference images, accurate capture of the user image from multiple perspectives can be realized, brain supplement generation for invisible areas can be avoided, and the accuracy of image replication can be significantly enhanced.

[0220] Referring to FIG. 7, FIG. 7 shows a flowchart of a video processing model training method according to an embodiment of the present disclosure, which specifically includes the following steps.

[0221] Step 702: determining a plurality of reference images of the target object under different view angles, a reference pose image of the target object in each reference image, a target video of the target object, a target pose image sequence of the target object in the target video, and a noise image sequence corresponding to the target video.

[0222] Step 704: obtaining reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature sequence of the target pose image sequence.

[0223] In one or more embodiments of the present disclosure, the obtaining of the reference image features of the reference images, the reference pose features of the reference pose images, and the target pose feature sequence of the target pose image sequence comprises:

[0224] inputting the plurality of reference images, the reference pose images, and the target pose image sequence into the image processing model;

[0225] extracting features of the plurality of reference images, the reference pose images, and the target pose image sequence in the image processing model to obtain the reference image features of the reference images, the reference pose features of the reference pose images, and the target pose feature sequence of the target pose image sequence.

[0226] Step 706: obtaining a reference network feature according to the reference image features and the reference pose features, obtaining a candidate prediction feature sequence according to the reference network feature, the target pose feature sequence, and the noise image sequence, and obtaining a target prediction feature sequence according to the candidate prediction feature sequence and a time sequence relationship between each candidate prediction feature.

[0227] In one or more embodiments of the present disclosure, the obtaining of the reference network feature according to the reference image features and the reference pose features, and the obtaining of the candidate prediction feature sequence according to the reference network feature, the target pose feature sequence, and the noise image sequence comprises:

[0228] obtaining a reference network feature according to the reference image features and the reference pose features, and obtaining a candidate prediction feature sequence according to the reference network feature, the target pose feature sequence, and the noise image sequence, by using the image processing model.

[0229] The candidate prediction feature sequence can be understood as a feature output by a cross-attention layer of a denoising network unit.

[0230] Specifically, the video processing model takes the image processing model trained as a base model, and trains the video processing model based on the image processing model. Based on the target video, a target pose image sequence of a target object in the target video is obtained, and a target pose feature sequence of the target pose image sequence is further obtained. The target pose feature sequence contains time information. For details, refer to the above embodiments, which will not be described here.

[0231] In one or more embodiments of the present disclosure, a time attention layer is arranged in the denoising network unit of the image processing model. The target prediction feature sequence is obtained according to the candidate prediction feature sequence and the time sequence relationship between each candidate prediction feature, which includes:

[0232] The time attention layer is used to perform time attention processing on the candidate prediction features in the candidate prediction feature sequence and the time sequence relationship between each candidate prediction feature, so as to obtain the target prediction feature sequence.

[0233] Specifically, the time attention layer is introduced into the image processing model trained in the previous stage, and the time attention layer is arranged behind the cross attention layer in each denoising network unit, and the spatial attention layer and the cross attention layer in the image processing model are kept unchanged.

[0234] The video processing model still selects multiple reference images as input, and selects multiple continuous target images and continuous control signals (target pose image sequence) as input. In actual application, in order to align the multiple frame denoising network features of the denoising network unit, the reference network features are copied in the time dimension, and then spliced with the multiple frame denoising network features for subsequent training. In the time attention layer, the attention operation is performed on the candidate prediction feature sequence output by the cross attention layer in the denoising network unit in the time dimension, that is, the time attention processing is performed on the candidate prediction features in the candidate prediction feature sequence and the time sequence relationship between each candidate prediction feature, so as to learn the correlation between different frames and obtain the target prediction feature sequence.

[0235] The video processing model training method provided by the embodiments of the present disclosure can learn the correlation between different frames by performing attention operation on the candidate prediction feature sequence output by the cross attention layer in the denoising network unit through the time attention layer, so that the video processing model can have the ability to generate continuous pictures, and the generated video is smoother.

[0236] Step 708: obtaining a prediction video according to the target prediction feature sequence, wherein the pose of the target object in the prediction video is determined by referring to the pose of the target object in the target video.

[0237] Step 710: training a video processing model according to the predicted video and the target video, to obtain a trained video processing model, wherein the video processing model is determined based on the trained image processing model in the image processing model training method.

[0238] In one or more embodiments of the present disclosure, the training of the video processing model according to the predicted video and the target video to obtain the trained video processing model comprises:

[0239] According to the predicted video and the target video, the network parameters of the time attention layer in the denoising network unit of the image processing model are adjusted to obtain the trained video processing model.

[0240] In actual application, since the video processing model takes the image processing model as a base model, the network parameters of the newly set time attention layer are trained, and the network parameters of other network units are not changed.

[0241] For specific implementation, reference can be made to the above embodiments, which will not be described here again.

[0242] The video processing model training method provided by the embodiments of the present disclosure sets a time attention layer related to a video on the basis of the image processing model obtained by training, learns the correlation between different frames in the target video by using the time attention layer, so that the video processing model has the ability to generate continuous pictures, the generated video is smoother, and on the basis of inputting multiple reference pictures, the user image can be accurately captured from multiple perspectives, and each detail of the user can be restored in the generated video.

[0243] Referring to FIG. 8, FIG. 8 shows a flowchart of another video processing model training method provided by an embodiment of the present disclosure, which specifically comprises the following steps.

[0244] Step 802: determining multiple reference images of a target object under different perspectives, reference pose images of the target object in each reference image, a target video of the target object, a target pose image sequence of the target object in the target video, a noise image sequence corresponding to the target video, and inputting the multiple reference images, the reference pose images, the target pose image sequence, and the noise image sequence into a video processing model.

[0245] The video processing model is determined based on the trained image processing model in the image processing model training method, and the video processing model comprises a feature extraction unit, a reference network unit, a denoising network unit, and a prediction unit.

[0246] Step 804: acquiring, by using the feature extraction unit, reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature sequence of the target pose image sequence.

[0247] Step 806: obtaining, by using the reference network unit, reference network features according to the reference image features and the reference pose features, and obtaining, by using the denoising network unit, a candidate prediction feature sequence according to the reference network features, the target pose feature sequence, and the noisy image sequence, and obtaining a target prediction feature sequence according to the candidate prediction feature sequence.

[0248] Step 808: obtaining, by using the prediction unit, a prediction video according to the target prediction feature sequence, wherein the pose of the target object in the prediction video is determined by referring to the pose of the target object in the target video.

[0249] Step 810: training a video processing model according to the prediction video and the target video to obtain a trained video processing model.

[0250] The specific implementation can refer to the above embodiments, which will not be repeated here.

[0251] The video processing model training method provided by the embodiments of the present disclosure sets a time attention layer related to a video on the basis of the image processing model obtained by training, learns the correlation between different frames in the target video by using the time attention layer, so that the video processing model has the ability to generate continuous pictures, and the generated video is smoother, and also on the basis of inputting multiple reference images, the user image can be accurately captured from multiple perspectives, and each detail of the user can be restored in the generated video.

[0252] Referring to FIG. 9, FIG. 9 shows a network framework diagram of a video processing model according to an embodiment of the present disclosure, and according to the network framework diagram, the training of the video processing model is realized by the following content.

[0253] The video processing model is determined on the basis of the above-mentioned image processing model obtained by training; specifically, the time attention layer is introduced in the denoising image processing unit in the above-mentioned image processing model, and the spatial attention layer and the cross-attention layer remain unchanged, and the video processing model is obtained.

[0254] In actual application, the video processing model needs to determine multiple reference images of a target object and a target video of the target object, and obtain multiple continuous target images from the target video.

[0255] In actual applications, on the basis of determining the target image, the image noise adding model can be used to add noise to multiple target images to obtain noise-added noise images. For example, the DDIM model (Denoising Diffusion Implicit Models) is used to add noise to multiple target images, and multiple noise images corresponding to the multiple target images (i.e., the noise image sequence in the above embodiment) can be obtained by adding random Gaussian noise.

[0256] In the case where the input of the image processing model is multiple target images, multiple noise images of the multiple target images, and multiple reference images, the multiple target images and the multiple reference images are input into the pose estimation network layer of the image processing model, multiple reference pose images corresponding to the multiple reference images and a target pose image sequence (i.e., the pose sequence in the figure) corresponding to the multiple target images are constructed, and the multiple reference pose images and the target pose image sequence are input into the pose feature extraction layer to obtain multiple reference pose features corresponding to the multiple reference pose images and a target pose feature sequence corresponding to the target pose image sequence, wherein the target pose feature sequence contains time information.

[0257] Of course, in actual applications, the pose estimation network layer can also be an independent pose estimation model independent of the image processing model. In this case, the multiple reference images and the multiple target images are input into the pose estimation model to construct multiple reference pose images corresponding to the multiple reference images and a target pose image sequence corresponding to the multiple target images, and the target pose image sequence, multiple noise images of the multiple target images, multiple reference images, and multiple reference pose images are input into the image processing model to obtain multiple reference pose features corresponding to the multiple reference pose images and a target pose feature sequence corresponding to the target pose image sequence by using the pose feature extraction layer of the image processing model.

[0258] Specifically, the video processing model includes two video processing units, one being a reference video processing unit (i.e., the reference image processing unit in the above embodiment, the reference network in FIG. 9), and the other being a denoising video processing unit (the time attention layer is introduced into the denoising image processing unit in the above embodiment, and the denoising network in FIG. 9).

[0259] By using the above embodiment, the multiple reference network features can be obtained by using the spatial attention layer of the reference image processing unit.

[0260] The target posture feature sequence is fused with multiple frames of noise images to obtain a target video processing feature (i.e., the multi-frame denoising network feature in the above embodiment). On the basis that the target posture feature sequence contains time information, the target video processing feature also contains time information. In order to align the reference network feature with the target video processing feature, the multiple reference network features are copied in the time dimension, and then fused with the target video processing feature, and input into the spatial attention layer of the denoising video processing unit to obtain a target reference feature in which different weights are assigned to each reference network feature and target image processing feature.

[0261] Similarly to the above embodiment, CLIP image features corresponding to multiple reference images are obtained and integrated into the cross-attention layer of the denoising video processing unit. In each spatial attention layer, the multiple reference network features of the reference video processing unit are integrated into the spatial attention layer of the denoising video processing unit. In each cross-attention layer, the CLIP image features are integrated into the cross-attention layer of the denoising image processing unit.

[0262] A time attention layer is introduced after each cross-attention layer of the denoising video processing unit. The features output by the cross-attention layer are input into the time attention layer. The time attention layer performs attention operation in the time dimension to obtain correlation information between multiple frames of target images in the target posture image sequence according to the time information contained in the features, and outputs a candidate prediction feature sequence containing time sequence information. The specific training process can be referred to FIG. 10. FIG. 10 shows a training schematic diagram of a video processing model training method according to an embodiment of the present disclosure. In the denoising network feature, t represents the time dimension.

[0263] A time attention module (i.e., the time attention layer in the above embodiment) is introduced after the cross-attention module to perform time attention processing and obtain the correlation between frames of the target video.

[0264] The features output by the denoising image processing unit are input into a decoder, and a predicted video is output by the decoder, so as to train the video processing model according to the predicted image and the target video. In the embodiment of the present disclosure, the time attention layer of the video processing model is trained to enable the video processing model to have the ability to generate continuous pictures, and the generated video is smoother.

[0265] The video processing model training method provided by the embodiment of the present disclosure introduces a technology capable of processing any number of reference images input to generate high-quality videos. By adaptively extracting key information from these reference images, not only the accuracy of image replication is significantly enhanced, but also the expressiveness of video generation is improved.

[0266] Referring to FIG. 11, FIG. 11 shows a flowchart of a video processing method according to an embodiment of the present disclosure, which specifically includes the following steps.

[0267] Step 1102: determining a reference image and a target video, obtaining a reference pose image of a reference object in the reference image according to the reference image, and obtaining a target pose sequence image of a target object in the target video and a noise sequence image corresponding to the target video according to the target video.

[0268] In one or more embodiments of the present disclosure, the determination of the reference image and the target video includes:

[0269] receiving a reference image and a target video sent by a client, wherein the reference image and the target video are sent in a case where a specific operation is triggered on a user interaction interface of the client.

[0270] Step 1104: performing video generation using a video processing model according to the reference image, the reference pose image, the target pose sequence image, and the noise sequence image, to obtain a generated video, wherein the video processing model is obtained by training the video processing model using the video processing model training method, the appearance of an object in the generated video is determined by the appearance of the reference object, and the motion pose of the object in the generated video is determined by the motion pose of the target object.

[0271] Specifically, taking the application of the video processing method to a short video platform as an example, a user can provide multiple images of himself / herself from different perspectives, and select a short video as a target video, aiming to use the pose of the object in the target video to demonstrate the motion pose. The multiple images provided by the user and the selected target video are input into the video processing model, and the video processing model is used to generate a video in which the user demonstrates the motion pose using the pose of the object in the target video.

[0272] Moreover, in the case where the user provides multiple images from different perspectives, even if the object in the target video selected by the user is converted from a front perspective to a back perspective, the video processing model can avoid brain-storming generation for invisible areas, but instead adaptively increase the weight of the features corresponding to the back image in the images provided by the user, so as to ensure that the video processing model can still generate a smooth and high-quality video.

[0273] For another example, taking the application of the video processing method to a sports scene as an example, the video processing method can provide an immersive experience for participants, as if they were actually participating in sports. The video processing method can not only capture every subtle motion of the characters in sports activities, but also vividly display these details when making cool sports videos.

[0274] The video processing method provided by the embodiments of the present disclosure can support any number of reference images, greatly improving the flexibility of use and opening up a broad imagination space for future product design.

[0275] Corresponding to the method embodiments described above, the present disclosure also provides image processing model training device embodiments. FIG. 12 shows a structural schematic diagram of an image processing model training device according to an embodiment of the present disclosure. As shown in FIG. 12, the device comprises:

[0276] The image determination module 1202 is configured to determine a plurality of reference images of a target object under different viewing angles, reference pose images of the target object in each reference image, a target image of the target object, a target pose image of the target object in the target image, and a noise image corresponding to the target image, wherein the poses of the target object in each reference image and the target image are different.

[0277] The feature acquisition module 1204 is configured to acquire reference image features of each reference image, reference pose features of each reference pose image, and a target pose feature of the target pose image.

[0278] The feature acquisition module 1204 is configured to acquire reference image features of each reference image, reference pose features of each reference pose image, and a target pose feature of the target pose image.

[0279] The image acquisition module 1208 is configured to acquire a predicted image according to the target prediction feature, wherein the pose of the target object in the predicted image is determined by referring to the pose of the target object in the target pose image.

[0280] The model training module 1210 is configured to train the image processing model according to the predicted image and the target image, and obtain a trained image processing model.

[0281] Optionally, the image determination module 1202 is further configured to:

[0282] determine a plurality of reference images of a target object under different viewing angles and a target image of the target object.

[0283] input the plurality of reference images and the target image into a pose estimation model respectively to obtain reference pose images of the target object in each reference image and a target pose image of the target object in the target image.

[0284] input the target image into an image noise adding model to obtain a noise image corresponding to the target image, wherein the pose estimation model and the image noise adding model are machine learning models.

[0285] Optionally, the image determination module 1202 is further configured to:

[0286] determine a reference video of the target object, and obtain a plurality of video frames of the target object at different perspectives as reference images from the reference video; or

[0287] photograph a plurality of reference images of the target object at different perspectives.

[0288] Optionally, the feature obtaining module 1204 is further configured to:

[0289] input the plurality of reference images, the reference pose images, and the target pose image into the image processing model;

[0290] perform feature extraction on the reference images, the reference pose images, and the target pose image in the image processing model to obtain reference image features of the reference images, reference pose features of the reference pose images, and target pose features of the target pose image.

[0291] Optionally, the feature obtaining module 1204 is further configured to:

[0292] perform image feature extraction on the reference images by using the image feature extraction unit to obtain reference image features of the reference images;

[0293] perform pose feature extraction on the reference pose images and the target pose image by using the pose feature extraction unit respectively to obtain reference pose features of the reference pose images and target pose features of the target pose image.

[0294] Optionally, the feature obtaining module 1206 is further configured to:

[0295] perform spatial attention processing on the reference image features and the reference pose features by using a spatial attention layer of the reference network unit to obtain the reference network features;

[0296] perform spatial attention processing on the reference network features, the target pose features, and the noise image by using a spatial attention layer of the de-noising network unit to obtain the target prediction features.

[0297] Optionally, the feature obtaining module 1204 is further configured to:

[0298] perform image feature extraction on the reference images by using the first image feature extraction layer to obtain first reference image features of the reference images;

[0299] perform image feature extraction on the reference images by using the second image feature extraction layer to obtain second reference image features of the reference images, wherein the second reference image features comprise semanticized features of the reference images.

[0300] Optionally, the feature obtaining module 1206 is further configured to:

[0301] perform spatial attention processing on the first reference image features and the reference pose features by using a spatial attention layer of the reference network unit to obtain the reference network features;

[0302] perform spatial attention processing on the reference network features, the target pose features and the noisy image by using a spatial attention layer of the denoising network unit to obtain the target key features;

[0303] perform cross-attention processing on the second reference image features and the target key features by using a cross-attention layer of the denoising network unit to obtain the target prediction features.

[0304] Optionally, the feature obtaining module 1206 is further configured to:

[0305] perform feature fusion on the first reference image features and the reference pose features to obtain reference fusion features, and perform spatial attention processing on the reference fusion features by using each spatial attention layer of the reference network unit to obtain the reference network features;

[0306] perform cross-attention processing on the second reference image features and the reference network features by using each cross-attention layer of the reference network unit to obtain reference cross features;

[0307] determine the reference cross features as the reference fusion features, and continue to perform the spatial attention processing on the reference fusion features by using each spatial attention layer of the reference network unit to obtain the reference network features.

[0308] Optionally, the feature obtaining module 1206 is further configured to:

[0309] perform feature fusion on the target pose features and the noisy image to obtain denoising fusion features, and perform spatial attention processing on the reference network features and the denoising fusion features by using each spatial attention layer of the denoising network unit to obtain the target key features;

[0310] In each cross attention layer of the denoising network unit, the second reference image feature and the target key feature are subjected to cross attention processing to obtain the target prediction feature, and the target prediction feature is determined as the denoising fusion feature, and the step of performing spatial attention processing on the reference network feature and the denoising fusion feature in each spatial attention layer of the denoising network unit to obtain the target key feature is continued.

[0311] Optionally, the image obtaining module 1208 is further configured to:

[0312] The target prediction feature is subjected to decoding processing by the prediction unit to obtain a prediction image.

[0313] Optionally, the model training module 1210 is further configured to:

[0314] A loss function of the prediction image and the target image is determined, and the image processing model is trained according to the loss function to obtain a trained image processing model.

[0315] The image processing model training device provided by one embodiment of the present disclosure can realize accurate capturing of the image of the target object from multiple perspectives by using multiple reference images under different perspectives, avoids random generation of invisible areas, obtains reference network features related to reference images by obtaining reference image features of each reference image and reference pose features of each reference pose image, and uses the reference network features and target pose features of the target pose image to guide the process of restoring a noise image to an image in the image processing model, thereby obtaining accurate target prediction features, i.e., the target prediction features contain features related to reference images in the reference image features and target pose features of the target pose image. On the basis of obtaining a prediction image based on the target prediction features, the accuracy of the generated prediction image is improved, and by accurately capturing the image of the target object from multiple perspectives, each detail of the target object in the prediction image is restored by referring to the image of the target object under each perspective, and the generation quality of the target object in the prediction image is ensured.

[0316] The above is a schematic scheme of the image processing model training device of the present embodiment. It should be noted that the technical scheme of the image processing model training device belongs to the same concept as the technical scheme of the image processing model training method described above, and the details of the technical scheme of the image processing model training device that are not described in detail can be referred to the description of the technical scheme of the image processing model training method.

[0317] Corresponding to the method embodiments described above, the disclosure also provides video processing model training device embodiments. FIG. 13 shows a structural schematic diagram of a video processing model training device according to an embodiment of the disclosure. As shown in FIG. 13, the device comprises:

[0318] An image determining module 1302 configured to determine a plurality of reference images of a target object under different perspectives, a reference pose image of the target object in each reference image, a target video of the target object, a target pose image sequence of the target object in the target video, and a noise image sequence corresponding to the target video.

[0319] A feature obtaining module 1304 configured to obtain reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature sequence of the target pose image sequence.

[0320] A feature obtaining module 1306 configured to obtain reference network features according to the reference image features and the reference pose features, obtain a candidate prediction feature sequence according to the reference network features, the target pose feature sequence, and the noise image sequence, and obtain a target prediction feature sequence according to the candidate prediction feature sequence and a time sequence relationship between each candidate prediction feature.

[0321] A video obtaining module 1308 configured to obtain a prediction video according to the target prediction feature sequence, wherein a pose of the target object in the prediction video is determined by referring to a pose of the target object in the target video.

[0322] A model training module 1310 configured to train a video processing model according to the prediction video and the target video, and obtain a trained video processing model, wherein the video processing model is determined based on a trained image processing model in the image processing model training method described above.

[0323] Optionally, the feature obtaining module 1304 is further configured to:

[0324] input the plurality of reference images, each reference pose image, and the target pose image sequence into the image processing model;

[0325] extract features of the plurality of reference images, each reference pose image, and the target pose image sequence in the image processing model to obtain reference image features of each reference image, reference pose features of each reference pose image, and a target pose feature sequence of the target pose image sequence.

[0326] Optionally, the feature obtaining module 1306 is further configured to:

[0327] The image processing model is used to obtain reference network features according to the reference image features and the reference pose features, and a candidate prediction feature sequence is obtained according to the reference network features, the target pose feature sequence, and the noise image sequence.

[0328] Optionally, the feature obtaining module 1306 is further configured to:

[0329] The time attention layer is used to perform time attention processing on the candidate prediction features in the candidate prediction feature sequence and the time sequence relationship between the candidate prediction features, to obtain a target prediction feature sequence.

[0330] Optionally, the model training module 1310 is further configured to:

[0331] According to the prediction video and the target video, the network parameters of the time attention layer in the denoising network unit of the image processing model are adjusted to obtain the video processing model.

[0332] The video processing model training apparatus provided by the embodiments of the present disclosure can learn the correlation between different frames by performing attention operation on the candidate prediction feature sequence output by the cross attention layer in the denoising network unit through the time attention layer, so that the video processing model can have the ability to generate continuous pictures, and the generated video is smoother.

[0333] The above is a schematic scheme of the video processing model training apparatus of the present embodiment. It should be noted that the technical scheme of the video processing model training apparatus belongs to the same concept as the technical scheme of the video processing model training method described above, and the details of the technical scheme of the video processing model training apparatus that are not described in detail can be referred to the description of the technical scheme of the video processing model training method.

[0334] FIG. 14 shows a structural block diagram of a computing device 1400 according to an embodiment of the present disclosure. The components of the computing device 1400 include but are not limited to a memory 1410 and a processor 1420. The processor 1420 is connected to the memory 1410 through a bus 1430, and a database 1450 is used to save data.

[0335] The computing device 1400 also includes an access device 1440 that enables the computing device 1400 to communicate via one or more networks 1460. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of such networks, such as the Internet. The access device 1440 can include one or more of any type of network interface (for example, a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, Near Field Communication (NFC).

[0336] In one embodiment of the present disclosure, the above-mentioned components of the computing device 1400 and other components not shown in FIG. 14 can also be connected to each other, for example, through a bus. It should be understood that the computing device structure block diagram shown in FIG. 14 is only for the purpose of example, and is not a limitation on the scope of the present disclosure. Those skilled in the art can add or replace other components as needed.

[0337] The computing device 1400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (for example, a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (for example, a smartphone), a wearable computing device (for example, a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1400 can also be a mobile or stationary server.

[0338] The processor 1420 is configured to execute computer program / instructions, which when executed by the processor, implement the steps of the above-mentioned image processing model training method, video processing model training method, image processing method, and video processing method.

[0339] The various embodiments in the present disclosure are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the computing device embodiment is described simply because it is basically similar to the image processing model training method, the video processing model training method, the image processing method, and the video processing method embodiments. The relevant parts can be referred to the description of the image processing model training method, the video processing model training method, the image processing method, and the video processing method embodiments.

[0340] An embodiment of the present disclosure further provides a computer readable storage medium storing computer programs / instructions, which are executed by a processor to implement the steps of the image processing model training method, the video processing model training method, the image processing method, and the video processing method.

[0341] The various embodiments in the present disclosure are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, the computer readable storage medium embodiment is described simply because it is basically similar to the image processing model training method, the video processing model training method, the image processing method, and the video processing method embodiments. The relevant parts can be referred to the description of the image processing model training method, the video processing model training method, the image processing method, and the video processing method embodiments.

[0342] An embodiment of the present disclosure further provides a computer program product, which includes computer programs / instructions, which are executed by a processor to implement the steps of the image processing model training method, the video processing model training method, the image processing method, and the video processing method.

[0343] The above is a schematic scheme of the computer program product of the present embodiment. It should be noted that the technical scheme of the computer program product is the same as the technical scheme of the image processing model training method, the video processing model training method, the image processing method, and the video processing method, and the details of the technical scheme of the computer program product not described in detail can be referred to the description of the technical scheme of the image processing model training method, the video processing model training method, the image processing method, and the video processing method.

[0344] The above describes particular embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still accomplish the desired results. Additionally, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.

[0345] The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate additions or subtractions according to the requirements of patent practice, for example, in some regions, according to the patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0346] It should be noted that for the foregoing method embodiments, in order to facilitate description, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the embodiments of the present disclosure are not limited by the order of the described actions, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the present disclosure are all preferred embodiments, and the actions and modules involved are not necessarily necessary for the embodiments of the present disclosure.

[0347] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0348] The preferred embodiments of the present disclosure disclosed above are only used to help explain the present disclosure. The alternative embodiments do not describe all the details and do not limit the present disclosure to the specific embodiments described. Obviously, according to the content of the embodiments of the present disclosure, many modifications and changes can be made. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can well understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents.

Claims

1. An image processing model training method, comprising: determining a plurality of reference images of a target object under different perspectives, a reference pose image of the target object in each reference image, a target image of the target object, a target pose image of the target object in the target image, and a noise image corresponding to the target image, wherein the poses of the target object in the reference images and the target image are different; obtaining reference image features of the reference images, reference pose features of the reference pose images, and target pose features of the target pose image; obtaining reference network features according to the reference image features and the reference pose features, and obtaining target prediction features according to the reference network features, the target pose features, and the noise image; obtaining a prediction image according to the target prediction features, wherein the pose of the target object in the prediction image is determined by referring to the pose of the target object in the target pose image; training an image processing model according to the prediction image and the target image to obtain a trained image processing model.

2. The image processing model training method of claim 1, wherein the determining the plurality of reference images of the target object under different perspectives, the reference pose image of the target object in each reference image, the target image of the target object, the target pose image of the target object in the target image, and the noise image corresponding to the target image comprises: determining the plurality of reference images of the target object under different perspectives and the target image of the target object; inputting the plurality of reference images and the target image into a pose estimation model to obtain the reference pose image of the target object in each reference image and the target pose image of the target object in the target image; inputting the target image into an image noise adding model to obtain the noise image corresponding to the target image, wherein the pose estimation model and the image noise adding model are machine learning models.

3. The image processing model training method of claim 1 or 2, wherein the determining the plurality of reference images of the target object under different perspectives comprises: determining a reference video of the target object, and obtaining a plurality of video frames of the target object under different perspectives from the reference video as the reference images; or capturing a plurality of reference images of the target object under different perspectives.

4. The image processing model training method of any one of claims 1-3, wherein the obtaining the reference image features of the reference images, the reference pose features of the reference pose images, and the target pose features of the target pose image comprises: inputting the plurality of reference images, the reference pose images, and the target pose image into the image processing model; extracting features of each reference image, the reference pose images, and the target pose image in the image processing model to obtain the reference image features of the reference images, the reference pose features of the reference pose images, and the target pose features of the target pose image. ​ 5. The image processing model training method of claim 4, wherein the image processing model comprises an image feature extraction unit and a pose feature extraction unit; and wherein the feature extraction of each reference image, each reference pose image, and the target pose image in the image processing model to obtain reference image features of each reference image, reference pose features of each reference pose image, and target pose features of the target pose image comprises: image feature extraction of each reference image by the image feature extraction unit to obtain reference image features of each reference image; and pose feature extraction of each reference pose image and the target pose image by the pose feature extraction unit to obtain reference pose features of each reference pose image and target pose features of the target pose image. The reference network unit and the denoising network unit have the same network structure and share the initialized weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer; The reference network unit and the denoising network unit have the same network structure and share the initialized weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer and a cross-attention layer; The reference network unit and the denoising network unit have the same network structure and share the initialized weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer and a cross-attention layer; 6. The image processing model training method of any one of claims 1-5, the image processing model further comprising a reference network unit and a denoising network unit, wherein, The reference network unit and the denoising network unit have the same network structure and share the initialized weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer and a cross-attention layer; 7. The image processing model training method of claim 5, wherein the image feature extraction unit comprises a first image feature extraction layer and a second image feature extraction layer; and wherein the image feature extraction of each reference image by the image feature extraction unit to obtain reference image features of each reference image comprises: image feature extraction of each reference image by the first image feature extraction layer to obtain first reference image features of each reference image; and image feature extraction of each reference image by the second image feature extraction layer to obtain second reference image features of each reference image, wherein the second reference image features comprise semantic features of the reference images. The reference network unit and the denoising network unit have the same network structure and share the initialized weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer and a cross-attention layer; The reference network unit and the denoising network unit have the same network structure and share the initialized weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer and a cross-attention layer; The reference network unit and the denoising network unit have the same network structure and share the initialized weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer and a cross-attention layer; The reference network unit and the denoising network unit have the same network structure and share the initialized weights, and the reference network unit and the denoising network unit each comprise a spatial attention layer and a cross-attention layer; ​ ​ 8. The image processing model training method of claim 7, the image processing model further comprising a reference network unit and a denoising network unit, wherein, ​ ​ ​ The spatial attention layer of the denoising network unit is used for performing spatial attention processing on the reference network feature, the target attitude feature and the noisy image to obtain a target key feature; The cross-attention layer of the denoising network unit is used for performing cross-attention processing on the second reference image feature and the target key feature to obtain the target prediction feature.

9. The image processing model training method according to claim 8, wherein the reference network unit comprises a plurality of spatial attention layers and a plurality of cross-attention layers; The spatial attention layer of the reference network unit is used for performing spatial attention processing on the first reference image feature and the reference attitude feature to obtain the reference network feature, comprising: The first reference image feature and the reference attitude feature are fused to obtain a reference fusion feature, and the reference fusion feature is subjected to spatial attention processing at each spatial attention layer of the reference network unit to obtain the reference network feature; After the spatial attention layer of the reference network unit is used for performing spatial attention processing on the first reference image feature and the reference attitude feature to obtain the reference network feature, the method further comprises: The second reference image feature and the reference network feature are subjected to cross-attention processing at each cross-attention layer of the reference network unit to obtain a reference cross feature; The reference cross feature is determined as the reference fusion feature, and the step of performing spatial attention processing on the reference fusion feature at each spatial attention layer of the reference network unit to obtain the reference network feature is continued.

10. The image processing model training method according to claim 8 or 9, wherein the denoising network unit comprises a plurality of spatial attention layers and a plurality of cross-attention layers; The spatial attention layer of the denoising network unit is used for performing spatial attention processing on the reference network feature, the target attitude feature and the noisy image to obtain a target key feature, comprising: The target attitude feature and the noisy image are fused to obtain a denoising fusion feature, and the reference network feature and the denoising fusion feature are subjected to spatial attention processing at each spatial attention layer of the denoising network unit to obtain the target key feature; The cross-attention layer of the denoising network unit is used for performing cross-attention processing on the second reference image feature and the target key feature to obtain the target prediction feature, comprising: The second reference image feature and the target key feature are subjected to cross-attention processing at each cross-attention layer of the denoising network unit to obtain the target prediction feature, and the target prediction feature is determined as the denoising fusion feature, and the step of performing spatial attention processing on the reference network feature and the denoising fusion feature at each spatial attention layer of the denoising network unit to obtain the target key feature is continued. ​ 11. The image processing model training method of any one of claims 1-10, wherein the image processing model comprises a prediction unit; and wherein the obtaining a predicted image according to the target prediction feature comprises: decoding the target prediction feature using the prediction unit to obtain the predicted image.

12. The image processing model training method of any one of claims 1-11, wherein the training the image processing model according to the predicted image and the target image to obtain a trained image processing model comprises: determining a loss function of the predicted image and the target image, and training the image processing model according to the loss function to obtain the trained image processing model.

13. An image processing model training method, comprising: determining a plurality of reference images of a target object under different viewing angles, a reference pose image of the target object in each reference image, a target image of the target object, a target pose image of the target object in the target image, and a noise image corresponding to the target image, and inputting the plurality of reference images, the reference pose image, the target pose image, and the noise image into an image processing model, wherein the pose of the target object in each reference image is different from the pose of the target object in the target image, and the image processing model comprises a feature extraction unit, a reference network unit, a denoising network unit, and a prediction unit; obtaining reference image features of the plurality of reference images, reference pose features of the reference pose image, and target pose features of the target pose image using the feature extraction unit; obtaining reference network features according to the reference image features and the reference pose features using the reference network unit, and obtaining target prediction features according to the reference network features, the target pose features, and the noise image using the denoising network unit; obtaining a predicted image according to the target prediction features using the prediction unit, wherein the pose of the target object in the predicted image is determined by referring to the pose of the target object in the target pose image; and training the image processing model according to the predicted image and the target image to obtain a trained image processing model.

14. A video processing model training method, comprising: determining a plurality of reference images of a target object under different viewing angles, a reference pose image of the target object in each reference image, a target video of the target object, a sequence of target pose images of the target object in the target video, and a sequence of noise images corresponding to the target video; obtaining reference image features of the plurality of reference images, reference pose features of the reference pose image, and a sequence of target pose features of the sequence of target pose images; obtaining reference network features according to the reference image features and the reference pose features, a sequence of candidate prediction features according to the reference network features, the sequence of target pose features, and the sequence of noise images, and a sequence of target prediction features according to the sequence of candidate prediction features and a time sequence relationship between each candidate prediction feature in the sequence of candidate prediction features; and training the image processing model according to the sequence of target prediction features and the target video to obtain a trained video processing model. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ obtaining a predicted video according to the target predicted feature sequence, wherein a pose of a target object in the predicted video is determined by referring to a pose of the target object in the target video; training a video processing model according to the predicted video and the target video to obtain a trained video processing model, wherein the video processing model is determined based on the trained image processing model in any one of the image processing model training methods in claims 1-12.

15. The video processing model training method of claim 14, wherein the obtaining of the reference image features of the reference images, the reference pose features of the reference pose images, and the target pose feature sequence of the target pose image sequence comprises: inputting the plurality of reference images, the reference pose images, and the target pose image sequence into the image processing model; extracting features of the plurality of reference images, the reference pose images, and the target pose image sequence in the image processing model to obtain the reference image features of the reference images, the reference pose features of the reference pose images, and the target pose feature sequence of the target pose image sequence.

16. The video processing model training method of claim 14 or 15, wherein the obtaining of the reference network features according to the reference image features and the reference pose features, and the obtaining of the candidate predicted feature sequence according to the reference network features, the target pose feature sequence, and the noise image sequence comprises: obtaining reference network features according to the reference image features and the reference pose features, and obtaining a candidate predicted feature sequence according to the reference network features, the target pose feature sequence, and the noise image sequence by using the image processing model.

17. The video processing model training method of any one of claims 14-16, wherein a time attention layer is arranged in a denoising network unit of the image processing model. The obtaining of the target predicted feature sequence according to the candidate predicted feature sequence and the time sequence relationship between the candidate predicted features comprises: performing time attention processing on the candidate predicted features in the candidate predicted feature sequence and the time sequence relationship between the candidate predicted features by using the time attention layer to obtain the target predicted feature sequence.

18. The video processing model training method of claim 17, wherein the training of the video processing model according to the predicted video and the target video to obtain the trained video processing model comprises: adjusting network parameters of the time attention layer in the denoising network unit of the image processing model according to the predicted video and the target video to obtain the trained video processing model.

19. A video processing model training method, comprising: determining a plurality of reference images of the target object under different perspectives, a reference pose image of the target object in each of the reference images, a target video of the target object, a target pose image sequence of the target object in the target video, a noise image sequence corresponding to the target video, and inputting the plurality of reference images, the reference pose image in each of the reference images, the target pose image sequence, and the noise image sequence into a video processing model, wherein the video processing model is determined based on the image processing model trained by any one of the image processing model training methods in claims 1-12, and the video processing model comprises a feature extraction unit, a reference network unit, a denoising network unit, and a prediction unit; obtaining reference image features of the reference images, reference pose features of the reference pose images, and a target pose feature sequence of the target pose image sequence by using the feature extraction unit; obtaining a reference network feature based on the reference image features and the reference pose features by using the reference network unit, obtaining a candidate prediction feature sequence based on the reference network feature, the target pose feature sequence, and the noise image sequence by using the denoising network unit, and obtaining a target prediction feature sequence based on the candidate prediction feature sequence; obtaining a prediction video based on the target prediction feature sequence by using the prediction unit, wherein the pose of the target object in the prediction video is determined by referring to the pose of the target object in the target video; training the video processing model based on the prediction video and the target video.

20. An image processing method, comprising: determining a reference image and a target image, obtaining a reference pose image of a reference object in the reference image based on the reference image, and obtaining a target pose image of a target object in the target image and a noise image corresponding to the target image based on the target image; generating an image based on the reference image, the reference pose image, the target pose image, and the noise image by using an image processing model to obtain a generated image, wherein the image processing model is trained by any one of the image processing model training methods in claims 1-12, the appearance of an object in the generated image is determined by the appearance of the reference object, and the motion pose of the object in the generated image is determined by the motion pose of the target object.

21. The image processing method of claim 20, wherein the determining the reference image and the target image comprises: receiving a reference image and a target image sent by a client, wherein the reference image and the target image are sent in a case where a specific operation is triggered on a user interaction interface of the client.

22. A video processing method, comprising: determining a reference image and a target video, obtaining a reference pose image of a reference object in the reference image based on the reference image, and obtaining a target pose sequence image of a target object in the target video and a noise sequence image corresponding to the target video based on the target video; According to the reference image, the reference pose image, the target pose sequence image and the noise sequence image, a video is generated by using a video processing model, to obtain a generated video, wherein the video processing model is trained by any one of the video processing model training methods in claims 14-18, the appearance of the object in the generated video is determined by the appearance of the reference object, and the action pose of the object in the generated video is determined by the action pose of the target object.

23. The video processing method of claim 22, wherein the determining the reference image and the target video comprises: receiving the reference image and the target video sent by a client, wherein the reference image and the target video are sent in a case that a specific operation is triggered on a user interaction interface of the client.

24. A computing device, comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, and the computer programs / instructions, when executed by the processor, implement the steps of the method in any one of claims 1-23.

25. A computer readable storage medium storing computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method in any one of claims 1-23.

26. A computer program product comprising computer programs / instructions, and the computer programs / instructions, when executed by a processor, implement the steps of the method in any one of claims 1-23.

Citation Information

Patent Citations

  • Image generation method and device based on attitude guidance

    CN117576248A

  • Method and device for generating posture of object in image

    CN117894038A

  • Image processing model training method, image processing method, video processing model training method and video processing method

    CN118982592A

Cited By

  • Reliability-aware multi-modal fusion method for fine-grained behavior detection

    CN122433020A