Image generation method and device, electronic equipment and storage medium

By using noise vectors and attitude control parameters in the image generation method to generate three-dimensional representations and mesh models, sampling and obtaining target features, and finally rendering and generating target images, the problems of poor image generation effect and insufficient attitude control in the prior art are solved, and high-fidelity and multi-pose image generation are achieved.

CN119991894APending Publication Date: 2025-05-13FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311503038.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to generate high-fidelity three-dimensional virtual human images, especially in local areas such as face and hands, and it is impossible to control the subject and local posture at the same time.

Method used

By determining the three-dimensional representation of each preset area in the target object based on the noise vector, generating a three-dimensional grid model with the posture control parameters, and obtaining the target features through camera pose sampling, and finally rendering and generating the target image.

Benefits of technology

It realizes the generation of high-fidelity images for different sizes in the target object, and has the ability to control the subject and local posture at the same time, improving the authenticity and control accuracy of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991894A_ABST
    Figure CN119991894A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image generation method and device, electronic equipment and a storage medium, and the method comprises the steps: determining the three-dimensional representation of each preset region in a target object according to a noise vector; wherein the three-dimensional representation is used for representing the features of the space midpoint; the size proportions of the preset areas in the target object are different; determining a three-dimensional grid model in a target attitude according to the attitude control parameters of the preset areas; according to the camera poses of the preset areas, sampling corresponding areas in the three-dimensional grid model to obtain sampling points corresponding to the preset areas; according to the three-dimensional representation of each preset area, determining a target feature corresponding to each sampling point; rendering each preset area according to each target feature to generate a target image; wherein the target image comprises a target object in a target posture. High-fidelity image generation can be realized, and attitude control of a main body and local parts can be realized at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of computer technology, and in particular, to an image generation method, device, electronic device, and storage medium. Background Art

[0002] In the prior art, a 3D virtual human image can be generated by generating a model. However, since the face, hands and other parts only occupy a small area of ​​the human body, the generation effect of these parts is often poor, which seriously affects the authenticity of the image. In addition, it is currently impossible to control the posture of these small areas while controlling the posture of the virtual human body. Summary of the invention

[0003] The embodiments of the present disclosure provide an image generation method, device, electronic device and storage medium, which can not only realize high-fidelity image generation, but also realize simultaneous posture control of the main body and the local parts.

[0004] In a first aspect, an embodiment of the present disclosure provides an image generation method, comprising:

[0005] Determine, according to the noise vector, a three-dimensional representation of each preset area in the target object; wherein the three-dimensional representation is used to characterize the characteristics of a point in space; and the size proportion of each preset area in the target object is different;

[0006] Determining a three-dimensional mesh model with a target posture according to the posture control parameters of each preset area;

[0007] According to the camera pose of each preset area, sampling the corresponding area in the three-dimensional grid model respectively to obtain sampling points corresponding to each preset area;

[0008] Determining target features corresponding to each of the sampling points according to the three-dimensional representations of each of the preset areas;

[0009] The preset areas are rendered according to the target features to generate a target image; wherein the target image includes the target object in the target posture.

[0010] In a second aspect, the present disclosure also provides an image generating device, including:

[0011] A three-dimensional representation determination module, used to determine the three-dimensional representation of each preset area in the target object according to the noise vector; wherein the three-dimensional representation is used to represent the characteristics of a point in space; and the size proportion of each preset area in the target object is different;

[0012] A mesh model determination module, used to determine a three-dimensional mesh model with a target posture according to the posture control parameters of each preset area;

[0013] A sampling module, used to sample the corresponding areas in the three-dimensional grid model according to the camera pose of each preset area, so as to obtain sampling points corresponding to each preset area;

[0014] A target feature determination module, used to determine the target feature corresponding to each sampling point according to the three-dimensional representation of each preset area;

[0015] An image generation module is used to render each of the preset areas according to each of the target features to generate a target image; wherein the target image includes the target object in the target posture.

[0016] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:

[0017] one or more processors;

[0018] a storage device for storing one or more programs,

[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the image generating method as described in any one of the embodiments of the present disclosure.

[0020] In a fourth aspect, the embodiments of the present disclosure further provide a storage medium comprising computer executable instructions, which, when executed by a computer processor, are used to execute the image generation method as described in any one of the embodiments of the present disclosure.

[0021] The technical solution of the embodiment of the present disclosure is to determine the three-dimensional representation of each preset area in the target object according to the noise vector; wherein the three-dimensional representation is used to represent the characteristics of a point in space; the size proportion of each preset area in the target object is different; according to the posture control parameters of each preset area, a three-dimensional grid model with a target posture is determined; according to the camera posture of each preset area, the corresponding area in the three-dimensional grid model is sampled respectively to obtain the sampling points corresponding to each preset area; according to the three-dimensional representation of each preset area, the target feature corresponding to each sampling point is determined; according to each target feature, each preset area is rendered to generate a target image; wherein the target image includes the target object with the target posture.

[0022] In the technical solution of the embodiment of the present disclosure, the corresponding three-dimensional representation can be determined for each preset area with different size ratios in the target object, so that each preset area can be rendered based on each three-dimensional representation, which can ensure high-fidelity image generation for each preset area with different size ratios in the target object. In addition, the three-dimensional grid model can be generated in combination with the posture control parameters of each preset area, and sampling can be performed on this basis to obtain the target features corresponding to each sampling point for rendering each preset area, which can achieve posture control of each preset area with different size ratios in the target object. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0024] Figure 1 A flowchart of an image generation method provided by an embodiment of the present disclosure;

[0025] Figure 2 A schematic diagram of three-plane features in an image generation method provided by an embodiment of the present disclosure;

[0026] Figure 3 A schematic block diagram of a network structure in an image generation method provided by an embodiment of the present disclosure;

[0027] Figure 4 A schematic diagram of the structure of an image generating device provided by an embodiment of the present disclosure;

[0028] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0030] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0031] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0032] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0033] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0034] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0035] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.

[0036] Figure 1 The present invention is a flowchart of an image generation method provided by an embodiment of the present invention. The present invention is applicable to the case of rendering an image containing a virtual object, and the virtual object in the rendered image has a high degree of realism and each preset area can achieve gesture control. The method can be executed by an image generation device, which can be implemented in the form of software and / or hardware, and the device can be configured in an electronic device, such as a computer.

[0037] like Figure 1 As shown, the image generation method provided in this embodiment may include:

[0038] S110. Determine a three-dimensional representation of each preset area in the target object according to the noise vector; wherein the three-dimensional representation is used to represent the characteristics of a point in space; and each preset area has a different size proportion in the target object.

[0039] In the disclosed embodiment, the noise vector may include a random noise vector, for example, a random noise vector sampled based on a Gaussian distribution. The target object may be considered as a three-dimensional virtual object that needs to be included in the target image to be generated, and at least two preset areas in the target object may be pre-divided. It can be considered that the target object can be constituted by each preset area. Among them, the size proportion of each preset area in the target object is different, and it can be considered that each preset area may include a main area with a larger size proportion in the target object, and may also include a local area with a smaller size proportion in the target object.

[0040] Among them, the three-dimensional representation is used to represent the features of points in space, and it can be considered that the features of each spatial point in the target object can be stored in the three-dimensional representation. Among them, any feature representation method that can store spatial point features can be used as a three-dimensional representation, and no specific limitation is made here.

[0041] Since the size proportion of each preset area in the target object is different, the size of the three-dimensional representation used to store its features may also be different. When the size proportion of the preset area in the target object is large, a correspondingly larger three-dimensional representation may be set, and the size of the three-dimensional representation of each preset area may be set based on experience or experiments.

[0042] The three-dimensional representation of each preset area in the target object can be generated based on the noise vector based on the existing feature determination method. For example, the three-dimensional representation of each preset area in the target object can be generated based on the random noise vector through the constructed neural network.

[0043] By determining three-dimensional representations of different sizes for each preset area, each preset area can be rendered based on the corresponding three-dimensional representation, which is beneficial to improving the generation capability of preset areas with smaller sizes, improving the image quality of rendered images, and achieving high-fidelity image generation.

[0044] S120: Determine a three-dimensional mesh model with a target posture according to the posture control parameters of each preset area.

[0045] In the disclosed embodiment, the posture control parameters may include parameters such as position and rotation, which can be used to control the posture of the corresponding preset area. The posture control parameters of each preset area can be set in advance, and it is necessary to ensure that each posture control parameter is reasonable so that the posture of each preset area is natural and reasonable. The three-dimensional mesh model can be considered as an unmapped model for reference. The three-dimensional mesh model presents a target posture, and it can be considered that the areas corresponding to each preset area in the three-dimensional mesh model respectively present postures corresponding to the posture control parameters of each preset area.

[0046] The three-dimensional mesh model can be generated based on an existing neural network model conditioned on the posture control parameters. For example, the pre-configured standard model can be deformed according to the posture control parameters of each preset area through the Skinned Multi-Person Linear eXpressive (SMPL-X) model to generate a three-dimensional mesh model with a target posture.

[0047] S130 . Sample the corresponding areas in the three-dimensional grid model according to the camera pose of each preset area to obtain sampling points corresponding to each preset area.

[0048] In the disclosed embodiment, the camera pose may include camera external parameters (such as rotation, translation and other parameters), and the camera pose of each preset area can be pre-set, and different camera poses must be reasonable. For example, the camera perspective of a higher preset area in the target object is usually an upward perspective, and the camera perspective of a lower preset area in the target object is usually a downward perspective, etc.

[0049] First, according to the camera posture of each preset area, multiple rays for sampling the three-dimensional mesh model can be emitted from the camera viewpoint. Then, the points closest to the rays can be determined from the areas of the three-dimensional mesh model corresponding to each preset area, and these points can be called sampling points corresponding to each preset area.

[0050] S140: Determine the target feature corresponding to each sampling point according to the three-dimensional representation of each preset area.

[0051] Among them, based on the inverse linear blend Skinning (IS) method, the corresponding spatial point in the three-dimensional representation can be determined according to the spatial position of the sampling point in the three-dimensional mesh model; the characteristics of the spatial point are used as the target features corresponding to the sampling point, so as to perform subsequent image rendering steps according to the target features.

[0052] Among them, determining the corresponding spatial points in the three-dimensional representation according to the spatial positions of the sampling points in the three-dimensional grid model may include: according to the deformation amount of the standard model to the three-dimensional grid model and the spatial positions of each sampling point in the three-dimensional grid model, reversely mapping each sampling point to the standard model to obtain a standard sampling point; according to the spatial correspondence between the standard model and the three-dimensional representation, determining the spatial points in the three-dimensional representation corresponding to the standard sampling points, and these spatial points can be considered as the spatial points corresponding to each sampling point.

[0053] By determining the target features corresponding to the sampling points of each preset area according to the camera posture of each preset area, the image generation quality of each preset area can be further improved.

[0054] In some optional implementations, the three-dimensional representation includes three-plane features; wherein the three-plane features are composed of three orthogonal plane features. Figure 2 A schematic diagram of three-plane features in an image generation method provided by an embodiment of the present disclosure. Figure 2 The three-plane features include feature maps of three planes, namely, the feature map of plane A, the feature map of plane B, and the feature map of plane C, and plane A, plane B, and plane C are orthogonal to each other. Among them, the three-dimensional space formed by the three-plane features has a spatial correspondence with the standard model before the three-dimensional mesh model is deformed.

[0055] Correspondingly, determining the target features corresponding to each sampling point based on the three-dimensional representation of each preset area can include: mapping each sampling point to the corresponding three-plane feature according to the posture control parameters of each preset area to obtain each mapping point; and determining each target feature according to the feature components of each plane feature in the three-plane feature to which each mapping point belongs.

[0056] Among them, the deformation amount of the standard model to the three-dimensional mesh model can be determined according to the posture control parameters of each preset area. Then, according to the deformation amount and the spatial position of each sampling point in the three-dimensional mesh model, each sampling point can be reversely mapped to the standard model to obtain the standard sampling point; and according to the spatial correspondence between the standard model and the three-plane feature, the spatial point corresponding to the standard sampling point in the three-plane feature is determined, that is, the mapping point is obtained.

[0057] Since different mapping points correspond to different preset areas, the three-plane features of the preset area corresponding to the mapping point can be used as the three-plane features to which the mapping point belongs. Figure 2 Taking the mapping point a in as an example, the process of determining its target feature may include: projecting the mapping point a onto plane A, plane B and screen C respectively to obtain projection points a1, a2 and a3; obtaining features at a1, a2 and a3 as feature components of the mapping point; determining the target feature according to each feature component, for example, obtaining the target feature by weighted summing up each feature component.

[0058] In these optional implementations, the three-dimensional representation may be a three-plane feature; furthermore, the target feature may be determined according to the feature components of the mapping points corresponding to each sampling point on the three-plane feature.

[0059] S150, rendering each preset area according to each target feature to generate a target image; wherein the target image includes a target object in a target posture.

[0060] After determining the target features corresponding to each sampling point, the color and geometric structure of each preset area can be rendered according to the target features based on the existing rendering method. For example, the target features can be encoded into colors and geometric shapes through a network layer composed of a multilayer perceptron (MLP); wherein the geometric modeling can be performed through a signed distance field (SDF).

[0061] In addition, in the process of rendering each preset area, for the connecting area of ​​at least two preset areas, the target features of the sampling points of the area in the three-dimensional representation corresponding to the at least two preset areas can also be combined for rendering. For example, the target features of the sampling points of the area in the three-dimensional representation corresponding to the at least two preset areas can be weighted summed to obtain the final target features of the sampling points of the area for rendering the area. Based on this rendering method, a smooth transition of the connecting area can be achieved, and the image quality can be improved.

[0062] In some optional implementations, rendering each preset area according to each target feature to generate a target image may include: rendering each preset area according to each target feature to obtain an initial image; and super-resolution reconstructing the initial image to obtain a target image. In this case, super-resolution reconstruction can be performed on the rendered initial image based on an existing super-resolution reconstruction network, which can improve image clarity and further improve image generation quality.

[0063] The technical solution of the embodiment of the present disclosure determines the three-dimensional representation of each preset area in the target object according to the noise vector; wherein the three-dimensional representation is used to represent the characteristics of a point in space; the size proportion of each preset area in the target object is different; according to the posture control parameters of each preset area, a three-dimensional grid model with a target posture is determined; according to the camera posture of each preset area, the corresponding area in the three-dimensional grid model is sampled respectively to obtain the sampling points corresponding to each preset area; according to the three-dimensional representation of each preset area, the target feature corresponding to each sampling point is determined; according to each target feature, each preset area is rendered to generate a target image; wherein the target image includes the target object with the target posture.

[0064] In the technical solution of the embodiment of the present disclosure, the corresponding three-dimensional representation can be determined for each preset area with different size ratios in the target object, so that each preset area can be rendered based on each three-dimensional representation, which can ensure high-fidelity image generation for each preset area with different size ratios in the target object. In addition, the three-dimensional grid model can be generated in combination with the posture control parameters of each preset area, and sampling can be performed on this basis to obtain the target features corresponding to each sampling point for rendering each preset area, which can achieve posture control of each preset area with different size ratios in the target object.

[0065] The various optional schemes in the image generation method provided in the embodiment of the present disclosure and the above embodiment can be combined. The image generation method provided in this embodiment describes in detail the network structure corresponding to the image generation method. The generator network can generate each three-dimensional representation according to the noise vector, and the neural rendering network can obtain the target features corresponding to each sampling point, and render each preset area according to the target features. In addition, the generator network and the neural rendering network can be constructed by generative adversarial training with the discriminator network, so as to ensure the authenticity of the rendered image.

[0066] In the image generation method provided in this embodiment, a three-dimensional representation of each preset area in the target object is determined according to a noise vector, including: determining the three-dimensional representation of each preset area in the target object according to the noise vector through a generator network; rendering each preset area according to each target feature, including: rendering each preset area according to each target feature through a neural rendering network; wherein the generator network and the neural rendering network are constructed by performing generative adversarial training with a discriminator network of each preset area.

[0067] For example, Figure 3 The following is a schematic block diagram of a network structure in an image generation method provided by an embodiment of the present disclosure. Figure 3 As shown, the target object may include a virtual human object; accordingly, each preset area may include a torso area, a face area, and a hand area. Figure 3 The target object and the preset area shown are taken as examples to explain the image generation process and the network construction process. In the case of other target objects and / or preset areas, reference can be made to the corresponding explanations in the embodiments of the present disclosure, which are not exhaustive here.

[0068] See also Figure 3 , the process of generating a target image containing a virtual human object may include:

[0069] Generator network execution steps:

[0070] Receive the noise vector Z and the control parameter c of the torso area b ; Through the mapping network in the generator network, Z and c b Encode and obtain three identical encoding results W; through the generator in the generator network, the three-dimensional representation F of the torso area is synthesized respectively based on the input W. b , 3D representation of the facial region F f and the three-dimensional representation of the hand region F h .

[0071] In the process of network construction, the control parameter c of the trunk area is used bAs a generation condition for 3D representation, it can ensure that the rendering results output by the network have good stability. In addition, during the construction process, the c of the input generator network b The weight of can be gradually reduced to gradually optimize the network parameters. Correspondingly, when generating images based on the constructed network, c b Input to the generator network for the generation of 3D representations.

[0072] In order to balance the rendering accuracy and network computing amount of the face and hand regions, the size of the three-plane features of the face and hand regions can be set to half of the three-plane features of the torso region. In addition, to further save computing costs, the symmetry of the hand can be used to use a three-plane feature to represent the left and right hands through a horizontal flip operation.

[0073] Neural rendering network execution steps:

[0074] According to the posture control parameter p of the torso area b , the posture control parameter p of the facial region f and the posture control parameter p of the hand area h , determine the 3D mesh model with the target posture; according to the camera posture c of the torso area b , the camera pose c of the facial region f and the camera pose c of the hand region h , respectively sample the corresponding areas in the three-dimensional mesh model to obtain the sampling points corresponding to each preset area; based on the reverse skinning ( Figure 3 ISp h The invention relates to a method for representing a virtual human body by a three-dimensional representation, wherein the mapping point corresponding to the sampling point in the three-dimensional representation is determined according to the spatial position of the sampling point in the three-dimensional mesh model; each target feature is determined according to the feature component of each plane feature in the three plane features to which the mapping point belongs; and the corresponding torso region, facial region or hand region is rendered according to each target feature to generate an initial image; wherein the initial image includes a virtual human object in a target posture.

[0075] Among them, the bounding boxes of the face area, the left hand area and the right hand area can be defined in the standard model before the three-dimensional mesh model is deformed. After the sampling points are mapped to the standard model, if they fall into these defined bounding boxes, the mapping points and target features can be determined from the three-plane features corresponding to the bounding boxes.

[0076] In addition, the initial image can be super-resolved and reconstructed through a super-resolution reconstruction network to obtain a target image with high clarity.

[0077] See also Figure 3In the process of building the generator network and the neural rendering network, the face discriminator, the torso discriminator, and the hand discriminator can also be connected after the super-resolution reconstruction network, so that the generator network and the neural rendering network can be trained against each discriminator network. By supervising the generation results of each region through each discriminator network, the quality of image generation and control ability can be guaranteed. After the network is built, each discriminator network can be removed from the overall network structure to output the target image.

[0078] In existing solutions for generating images containing 3D virtual human objects, the controllable area in the generated image is limited to the torso area, while the face area and hand area cannot be controlled. In addition, since the face, hands and other areas only occupy a small area of ​​the human body, the generated details are often less realistic.

[0079] The image generation method provided by the present disclosure, which includes a virtual human object, can render the torso area, facial area, and hand area separately through multi-part, multi-scale three-dimensional representations, which can improve the image generation capabilities of the facial area and the hand area. In addition, by performing multi-part rendering based on multiple posture control parameters and camera postures, it is not only possible to simultaneously control the torso area, facial area, and hand area, but also to improve the image quality of the hand area and the facial area. Experiments have shown that the image generation method provided by the present disclosure has very good image generation effects and control capabilities of virtual portrait objects on public data sets.

[0080] In some optional implementations, the construction process of the generator network and the neural rendering network may include:

[0081] Through the generator network and the neural rendering network, the image of each preset area is rendered according to the sample noise vector, the sample camera pose of each preset area and the sample pose control parameters of each preset area; through the discriminator network of each preset area, the image of each preset area is discriminated respectively to obtain each discrimination result; the generative adversarial loss is determined according to each discrimination result, and the generator network, the neural rendering network and the discriminator network are constructed based on the generative adversarial loss.

[0082] Among them, the process of generating a target image containing a virtual human object in the embodiment of the present disclosure can be referred to, and the image of each preset area can be rendered according to the sample noise vector, the sample camera pose of each preset area and the sample pose control parameters of each preset area.

[0083] See again Figure 3, when the target object is a virtual human object and each preset region includes a trunk region, a facial region and a hand region, the discriminator network may include a facial discriminator, a trunk discriminator and a hand discriminator. The rendered target image may be used as the input of the trunk discriminator, the facial region in the target image may be used as the input of the facial discriminator, and the hand region in the target image may be used as the input of the hand discriminator. In some implementations, the posture control parameter p of the trunk region may be b and camera pose c b At the same time, as the input of the torso discriminator, the posture control parameter p of the facial area f and camera pose c f At the same time, as the input of the face discriminator, the posture control parameter p of the hand area is h and camera pose c h At the same time, it serves as the input of the hand discriminator to improve the discrimination accuracy of the discriminator.

[0084] Through the discriminator network of each preset area, the image of each preset area is discriminated respectively to obtain the score of the image being real, which is used to determine the generative adversarial loss. Then, based on the generative adversarial loss, the generator network, neural rendering network and discriminator network can be constructed.

[0085] Exemplarily, the generative adversarial loss may be determined based on the following formula:

[0086]

[0087]

[0088] Among them, L G The image generation loss of the generator network and the neural rendering network can be characterized. This loss can be used to adjust the network parameters of the generator network and the neural rendering network to make the generated image closer to the real image; L D The discriminant loss of the discriminator network can be characterized, and the loss can be used to adjust the network parameters of the discriminator network so that the discriminator network can accurately identify the output images of the generator network and the neural rendering network as fake;

[0089] in, It can represent the generation or discrimination loss of the torso area, face area and hand area respectively, and when the upper right corner is G, it belongs to the image generation loss, and when it is D, it belongs to the discrimination loss; λ b , f , h The loss weighting factors of the torso area, face area and hand area can be represented respectively and can be obtained through learning; M f 、M hIt can be set to zero when the face and hand regions are not visible in the target image to balance the image generation loss; The regularization losses of the torso region, face region, and hand region can be represented respectively to regularize the discriminator network; ⊙ can represent instance multiplication.

[0090] In addition, in some implementations, in order to improve the rationality and smoothness of each preset area in the geometric dimension, it is also possible to G Add the minimum surface loss L minsorf , optical path function loss (Eikonal loss) L Eik and the prior regularization loss L Prior wait.

[0091] In these optional implementations, by setting a discriminator network corresponding to multiple regions to supervise the rendering results of each region, the image generation quality and control capability of each region can be guaranteed.

[0092] The technical solution of the embodiment of the present disclosure describes in detail the network structure corresponding to the image generation method. The generator network can generate each three-dimensional representation according to the noise vector, and the neural rendering network can obtain the target features corresponding to each sampling point, and render each preset area according to the target features. In addition, the generator network and the neural rendering network can be constructed by generative adversarial training with the discriminator network, so as to ensure the authenticity of the rendered image. In addition, the image generation method provided in the embodiment of the present disclosure and the image generation method provided in the above embodiment belong to the same public concept. The technical details not described in detail in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0093] The various optional solutions in the image generation method provided in the embodiment of the present disclosure and the above embodiment can be combined. The image generation method provided in this embodiment describes in detail the actual downstream application scenarios, such as text-driven generation of target objects, or voice-driven posture of target objects.

[0094] The image generation method provided in this embodiment may further include: generating a noise vector based on a text description of the target object. The text description of the target object may be converted into a noise vector based on an existing text conversion method. The text description may be obtained by recognizing speech data or may be text data directly input by a user. By generating a noise vector based on the text description, the generated target object may be aligned with the text description to meet the image generation requirements of the user.

[0095] The image generation method provided in this embodiment may also include: obtaining a posture control parameter sequence for each preset area; accordingly, after rendering a target image corresponding to each posture control parameter in the posture control parameter sequence, it also includes: generating a target video according to each target image.

[0096] Among them, the current posture control parameters of each preset area can be obtained in sequence from the posture control parameter sequence of each preset area, and a three-dimensional grid model can be generated according to the current posture control parameters; then, the corresponding areas in the three-dimensional grid model are sampled according to the camera posture of each preset area to obtain the sampling points corresponding to each preset area; according to the three-dimensional representation of each preset area, the target features corresponding to each sampling point are determined; and each preset area is rendered according to each target feature to generate a target image. After obtaining the target image sequence, a target video can also be generated according to each target image based on the existing method of generating video from an image to realize dynamic control of the posture of each preset area of ​​the target object to meet the user's animation generation needs.

[0097] In some implementations, obtaining the posture control parameter sequence of each preset area may include: determining the posture control parameter sequence of each preset area according to the received voice data. Among them, the posture control parameter sequence of each preset area can be determined from the received voice data based on an existing natural language processing model. Exemplarily, when the voice data is "Raise the right hand quickly and then slowly put it down", a series of posture control parameters of the right hand area can be set to achieve the action effect described by the voice data. Thereby, the posture of the target object can be driven by voice, the threshold for animation generation can be lowered, and the user experience can be improved.

[0098] The technical solution of the embodiment of the present disclosure describes in detail the actual downstream application scenarios, such as the generation of target objects driven by text, and the position and posture of target objects driven by voice, etc. The image generation method provided by the embodiment of the present disclosure and the image generation method provided by the above embodiment belong to the same public concept, and the technical details not described in detail in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.

[0099] Figure 4 The image generating device provided by the embodiment of the present disclosure is a structural schematic diagram. The image generating device provided by the embodiment is suitable for rendering an image containing a virtual object, and the virtual object in the rendered image has high realism and each preset area can realize posture control.

[0100] like Figure 4 As shown, the image generating device provided by the embodiment of the present disclosure may include:

[0101] The three-dimensional representation determination module 410 is used to determine the three-dimensional representation of each preset area in the target object according to the noise vector; wherein the three-dimensional representation is used to represent the characteristics of a point in space; and each preset area has a different size proportion in the target object;

[0102] A mesh model determination module 420, for determining a three-dimensional mesh model in a target posture according to the posture control parameters of each preset area;

[0103] The sampling module 430 is used to sample the corresponding areas in the three-dimensional grid model according to the camera posture of each preset area, so as to obtain sampling points corresponding to each preset area;

[0104] A target feature determination module 440 is used to determine the target feature corresponding to each sampling point according to the three-dimensional representation of each preset area;

[0105] The image generation module 450 is used to render each preset area according to each target feature to generate a target image; wherein the target image includes a target object in a target posture.

[0106] In some optional implementations, the three-dimensional representation includes a three-plane feature; wherein the three-plane feature is composed of three orthogonal plane features;

[0107] Accordingly, the target feature determination module can be used to:

[0108] According to the posture control parameters of each preset area, each sampling point is mapped to the corresponding three-plane feature to obtain each mapping point;

[0109] Each target feature is determined according to the feature components of each plane feature in the three plane features to which each mapping point belongs.

[0110] In some optional implementations, the three-dimensional representation determination module may be used to: determine the three-dimensional representation of each preset area in the target object according to the noise vector through the generator network;

[0111] The image generation module can be used to: render each preset area according to each target feature through a neural rendering network;

[0112] Among them, the generator network and the neural rendering network are constructed by generative adversarial training with the discriminator networks of each preset area.

[0113] In some optional implementations, the image generating device may further include:

[0114] Building blocks that can be used to build generator networks and neural rendering networks based on the following process:

[0115] Through the generator network and the neural rendering network, the image of each preset area is rendered according to the sample noise vector, the sample camera pose of each preset area and the sample pose control parameter of each preset area;

[0116] Through the discriminator network of each preset area, the image of each preset area is discriminated respectively to obtain each discrimination result;

[0117] The generative adversarial loss is determined according to each discrimination result, and the generator network, neural rendering network and discriminator network are constructed based on the generative adversarial loss.

[0118] In some optional implementations, the image generation module may also be used to:

[0119] Render each preset area according to each target feature to obtain an initial image;

[0120] The initial image is reconstructed by super-resolution to obtain the target image.

[0121] In some optional implementations, the image generating device may further include:

[0122] The noise generation module is used to generate a noise vector according to the text description of the target object.

[0123] In some optional implementations, the image generating device may further include:

[0124] A parameter sequence acquisition module is used to acquire the posture control parameter sequence of each preset area;

[0125] Correspondingly, the image generation module can also be used to generate a target video according to each target image after rendering to obtain the target image corresponding to each posture control parameter in the posture control parameter sequence.

[0126] In some optional implementations, the parameter sequence acquisition module may be used to:

[0127] According to the received voice data, a posture control parameter sequence of each preset area is determined.

[0128] In some optional implementations, the target object includes a virtual human object; accordingly, each preset area includes a torso area, a facial area, and a hand area.

[0129] The image generating device provided in the embodiments of the present disclosure can execute the image generating method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0130] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0131] Reference below Figure 5 , which shows an electronic device (eg, Figure 5 The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0132] like Figure 5 As shown, the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 to a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0133] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 500 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0134] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the image generation method of the embodiment of the present disclosure are executed.

[0135] The electronic device provided in the embodiment of the present disclosure and the image generation method provided in the above embodiment belong to the same disclosed concept. The technical details not fully described in this embodiment can be referred to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0136] The embodiment of the present disclosure provides a computer storage medium on which a computer program is stored. When the program is executed by a processor, the image generating method provided by the above embodiment is implemented.

[0137] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or a flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate, or transmit programs for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0138] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0139] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0140] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0141] According to the noise vector, a three-dimensional representation of each preset area in the target object is determined; wherein the three-dimensional representation is used to represent the characteristics of a point in space; each preset area has a different size proportion in the target object; according to the posture control parameters of each preset area, a three-dimensional grid model with a target posture is determined; according to the camera posture of each preset area, the corresponding area in the three-dimensional grid model is sampled respectively to obtain the sampling points corresponding to each preset area; according to the three-dimensional representation of each preset area, the target feature corresponding to each sampling point is determined; according to each target feature, each preset area is rendered to generate a target image; wherein the target image includes the target object with the target posture.

[0142] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0143] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0144] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the names of the units and modules do not, in some cases, limit the units and modules themselves.

[0145] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), etc.

[0146] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0147] According to one or more embodiments of the present disclosure, there is provided an image generating method, the method comprising:

[0148] Determine, according to the noise vector, a three-dimensional representation of each preset area in the target object; wherein the three-dimensional representation is used to characterize the characteristics of a point in space; and the size proportion of each preset area in the target object is different;

[0149] Determining a three-dimensional mesh model with a target posture according to the posture control parameters of each preset area;

[0150] According to the camera pose of each preset area, sampling the corresponding area in the three-dimensional grid model respectively to obtain sampling points corresponding to each preset area;

[0151] Determining target features corresponding to each of the sampling points according to the three-dimensional representations of each of the preset areas;

[0152] The preset areas are rendered according to the target features to generate a target image; wherein the target image includes the target object in the target posture.

[0153] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:

[0154] In some optional implementations, the three-dimensional representation includes three-plane features; wherein the three-plane features are composed of three orthogonal plane features;

[0155] Accordingly, determining the target features corresponding to each sampling point according to the three-dimensional representation of each preset area includes:

[0156] According to the posture control parameters of each preset area, each sampling point is mapped to a corresponding three-plane feature to obtain each mapping point;

[0157] Each target feature is determined according to the feature components of each plane feature in the three plane features to which each mapping point belongs.

[0158] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:

[0159] In some optional implementations, determining the three-dimensional representation of each preset area in the target object according to the noise vector includes: determining the three-dimensional representation of each preset area in the target object according to the noise vector through a generator network;

[0160] Rendering each preset area according to each target feature includes: rendering each preset area according to each target feature through a neural rendering network;

[0161] The generator network and the neural rendering network are constructed by performing generative adversarial training with the discriminator networks of the preset areas.

[0162] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:

[0163] In some optional implementations, the construction process of the generator network and the neural rendering network includes:

[0164] By means of the generator network and the neural rendering network, images of the preset areas are rendered according to the sample noise vector, the sample camera poses of the preset areas and the sample pose control parameters of the preset areas;

[0165] Using the discriminator network of each preset area, the image of each preset area is discriminated respectively to obtain each discrimination result;

[0166] A generative adversarial loss is determined according to each of the discrimination results, and the generator network, the neural rendering network and the discriminator network are constructed based on the generative adversarial loss.

[0167] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:

[0168] In some optional implementations, rendering the preset areas according to the target features to generate a target image includes:

[0169] Rendering each of the preset areas according to each of the target features to obtain an initial image;

[0170] The initial image is reconstructed with super resolution to obtain a target image.

[0171] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:

[0172] In some optional implementations, the noise vector is generated according to a text description of the target object.

[0173] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:

[0174] In some optional implementations, acquiring a posture control parameter sequence of each preset area;

[0175] Correspondingly, after rendering to obtain the target image corresponding to each posture control parameter in the posture control parameter sequence, the method further includes: generating a target video according to each target image.

[0176] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:

[0177] In some optional implementations, the acquiring the posture control parameter sequence of each preset area includes:

[0178] Determine the posture control parameter sequence of each preset area according to the received voice data.

[0179] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:

[0180] In some optional implementations, the target object includes a virtual human object; accordingly, the preset areas include a torso area, a facial area, and a hand area.

[0181] According to one or more embodiments of the present disclosure, there is provided an image generating device, the device comprising:

[0182] A three-dimensional representation determination module, used to determine the three-dimensional representation of each preset area in the target object according to the noise vector; wherein the three-dimensional representation is used to represent the characteristics of a point in space; and the size proportion of each preset area in the target object is different;

[0183] A mesh model determination module, used to determine a three-dimensional mesh model with a target posture according to the posture control parameters of each preset area;

[0184] A sampling module, used to sample the corresponding areas in the three-dimensional grid model according to the camera pose of each preset area, so as to obtain sampling points corresponding to each preset area;

[0185] A target feature determination module, used to determine the target feature corresponding to each sampling point according to the three-dimensional representation of each preset area;

[0186] An image generation module is used to render each of the preset areas according to each of the target features to generate a target image; wherein the target image includes the target object in the target posture.

[0187] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.

[0188] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0189] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.

Claims

1. An image generation method, characterized in that: include: Determine, according to the noise vector, a three-dimensional representation of each preset area in the target object; wherein the three-dimensional representation is used to characterize the characteristics of a point in space; and the size proportion of each preset area in the target object is different; Determining a three-dimensional mesh model with a target posture according to the posture control parameters of each preset area; According to the camera pose of each preset area, sampling the corresponding area in the three-dimensional grid model respectively to obtain sampling points corresponding to each preset area; Determining target features corresponding to each of the sampling points according to the three-dimensional representations of each of the preset areas; The preset areas are rendered according to the target features to generate a target image; wherein the target image includes the target object in the target posture.

2. The method according to claim 1, characterized in that The three-dimensional representation includes three-plane features; wherein the three-plane features are composed of three orthogonal plane features; Accordingly, determining the target features corresponding to each sampling point according to the three-dimensional representation of each preset area includes: According to the posture control parameters of each preset area, each sampling point is mapped to a corresponding three-plane feature to obtain each mapping point; Each target feature is determined according to the feature components of each plane feature in the three plane features to which each mapping point belongs.

3. The method according to claim 1, characterized in that: Determining the three-dimensional representation of each preset area in the target object according to the noise vector includes: determining the three-dimensional representation of each preset area in the target object according to the noise vector through a generator network; Rendering each preset area according to each target feature includes: rendering each preset area according to each target feature through a neural rendering network; The generator network and the neural rendering network are constructed by performing generative adversarial training with the discriminator networks of the preset areas.

4. The method according to claim 3, characterized in that The construction process of the generator network and the neural rendering network includes: By means of the generator network and the neural rendering network, images of the preset areas are rendered according to the sample noise vector, the sample camera poses of the preset areas and the sample pose control parameters of the preset areas; Using the discriminator network of each preset area, the images of each preset area are discriminated respectively to obtain respective discrimination results; A generative adversarial loss is determined according to each of the discrimination results, and the generator network, the neural rendering network and the discriminator network are constructed based on the generative adversarial loss.

5. The method according to claim 1, characterized in that Rendering each of the preset areas according to each of the target features to generate a target image includes: Rendering each of the preset areas according to each of the target features to obtain an initial image; The initial image is reconstructed with super resolution to obtain a target image.

6. The method according to claim 1, characterized in that Also includes: The noise vector is generated according to the text description of the target object.

7. The method according to claim 1, characterized in that Also includes: Acquire a sequence of posture control parameters for each of the preset areas; Correspondingly, after rendering to obtain the target image corresponding to each posture control parameter in the posture control parameter sequence, the method further includes: generating a target video according to each target image.

8. The method according to claim 7, characterized in that The step of obtaining the posture control parameter sequence of each preset area includes: Determine the posture control parameter sequence of each preset area according to the received voice data.

9. The method according to any one of claims 1 to 8, characterized in that: The target object includes a virtual human object; correspondingly, the preset areas include a torso area, a facial area and a hand area.

10. An image generating device, characterized in that: include: A three-dimensional representation determination module, used to determine the three-dimensional representation of each preset area in the target object according to the noise vector; wherein the three-dimensional representation is used to represent the characteristics of a point in space; and the size proportion of each preset area in the target object is different; A mesh model determination module, used to determine a three-dimensional mesh model with a target posture according to the posture control parameters of each preset area; A sampling module, used to sample the corresponding areas in the three-dimensional grid model according to the camera pose of each preset area, so as to obtain sampling points corresponding to each preset area; A target feature determination module, used to determine the target feature corresponding to each sampling point according to the three-dimensional representation of each preset area; An image generation module is used to render each of the preset areas according to each of the target features to generate a target image; wherein the target image includes the target object in the target posture.

11. An electronic device, characterized in that: The electronic device comprises: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image generating method according to any one of claims 1 to 9.

12. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the image generation method according to any one of claims 1 to 9 when executed by a computer processor.