A method for training an image generation model, an image generation method and device
By determining the spatial information and hand features of the gesture, a second image generation model is trained using an image generation model, which solves the problem of strong artificiality in gesture images and achieves image generation with higher realism.
Patent Information
- Application Number
- CN202310014499.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-05
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2043-01-05
AI Technical Summary
The gesture images generated by existing technologies have a strong sense of artificiality and are difficult to closely resemble real images, which affects the effectiveness of image generation.
By determining the spatial information and hand features of the gesture, a second image generation model is generated using an image generation model. The model is then trained using sample images to optimize its realism.
The generated images are more realistic and closer to real images, thus improving the effectiveness of image generation.
Smart Images

Figure CN116012883B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a training method for an image generation model, an image generation method, and an apparatus. Background Technology
[0002] Currently, electronic devices can render gesture images under different actions using computer graphics, and the realism of these gesture images can be enhanced through methods such as texture mapping.
[0003] However, the gesture images generated by the above methods have a strong sense of artificiality and are difficult to closely resemble real images, which may reduce the effectiveness of image generation. Summary of the Invention
[0004] This disclosure provides a training method for an image generation model, an image generation method, and an apparatus, which solves the technical problem in related technologies that the generated gesture images have a strong sense of artificiality, are difficult to closely resemble real images, and may reduce the effectiveness of image generation.
[0005] The technical solution of this disclosure is as follows:
[0006] According to a first aspect of the present disclosure, a method for training an image generation model is provided. The method may include: determining spatial information of at least one gesture and hand features, wherein the spatial information of a gesture is used to characterize the positional relationship between at least two key points included in the gesture, and the hand features are used to characterize attribute information of the hand; inputting the spatial information of each of the at least one gesture and the hand features corresponding to each gesture into a first image generation model to obtain a target image corresponding to the at least one gesture; and training the first image generation model based on the hand features corresponding to each gesture, the target image corresponding to the at least one gesture, and a sample image corresponding to the at least one gesture to generate a second image generation model.
[0007] Optionally, the training method of the image generation model further includes: acquiring sample data, the sample data including at least one sample image and descriptive information of gestures included in each of the at least one sample image, wherein the descriptive information of a gesture included in a sample image is used to characterize the meaning of the gesture included in the sample image, and the at least one sample image is a sample image corresponding to the at least one gesture; when the sample data contains descriptive information of a first gesture, the first sample image is encoded to obtain the hand features corresponding to the first gesture, the first gesture being one of the at least one gesture, and the first sample image being a sample image corresponding to the first gesture.
[0008] Optionally, the training method of the image generation model further includes: when there is no descriptive information of the first gesture in the sample data, determining the preset hand features as the hand features corresponding to the first gesture.
[0009] Optionally, determining the spatial information of at least one gesture specifically includes: rendering the first gesture based on the description information of the first gesture to obtain the spatial information of the first gesture, wherein the spatial information of the first gesture includes the three-dimensional coordinates of each of the at least two key points included in the first gesture and the rotation parameters of each key point, and the first gesture is one of the at least one gesture.
[0010] Optionally, the above-mentioned training of the first image generation model based on the hand features corresponding to each gesture, the target image corresponding to the at least one gesture, and the sample image corresponding to the at least one gesture to generate a second image generation model includes: determining a first loss, which is used to characterize the degree of inconsistency between the probability distribution of the hand features corresponding to each gesture and a preset normal distribution; inputting each target image in the at least one target image into an initial discriminator to obtain a first probability, which is used to characterize the probability that each target image is judged as a first label, the first label being used to characterize a sample image, and the at least one target image being a target image corresponding to the at least one gesture; determining a second loss, which is used to characterize the degree of inconsistency between the pixels of each sample image in the at least one sample image and the pixels of the at least one target image, the at least one sample image being a sample image corresponding to the at least one gesture; determining a third loss based on the first loss, the first probability, and the second loss; and updating the parameters in the first image generation model based on the third loss to generate the second image generation model.
[0011] Optionally, the training method of the image generation model further includes: inputting each target image into the initial discriminator to obtain a second probability, the second probability being used to characterize the probability that each target image is classified as a second label, the second label being used to characterize non-sample images; inputting each sample image into the initial discriminator to obtain a third probability, the third probability being used to characterize the probability that each sample image is classified as the first label; determining a fourth loss based on the second probability and the third probability; and updating the parameters in the initial discriminator based on the fourth loss to generate a target discriminator.
[0012] According to a second aspect of the present disclosure, an image generation method is provided. The method may include: determining spatial information of a preset gesture and preset hand features, wherein the spatial information of the preset gesture characterizes the positional relationship between at least two key points included in the preset gesture, and the preset hand features characterize attribute information of the hand; inputting the spatial information of the preset gesture and the preset hand features into a second image generation model to obtain a target generated image, wherein the second image generation model is trained based on the training method of any of the optional image generation models in the first aspect described above.
[0013] Optionally, the image generation method further includes: obtaining descriptive information of the preset gesture, the descriptive information of the preset gesture being used to characterize the meaning of the preset gesture; rendering the preset gesture based on the descriptive information of the preset gesture to obtain spatial information of the preset gesture, the spatial information of the preset gesture including the three-dimensional coordinates of each of at least two key points included in the preset gesture and the rotation parameters of each key point.
[0014] According to a third aspect of the present disclosure, a training apparatus for an image generation model is provided. The apparatus may include: a determining module and a processing module; the determining module is configured to determine spatial information and hand features of at least one gesture, wherein the spatial information of a gesture is used to characterize the positional relationship between at least two key points included in the gesture, and the hand features are used to characterize attribute information of the hand; the processing module is configured to input the spatial information of each of the at least one gesture and the corresponding hand features into a first image generation model to obtain a target image corresponding to the at least one gesture; the processing module is further configured to train the first image generation model based on the corresponding hand features of each gesture, the target image corresponding to the at least one gesture, and a sample image corresponding to the at least one gesture to generate a second image generation model.
[0015] Optionally, the training apparatus for the image generation model further includes an acquisition module; the acquisition module is configured to acquire sample data, the sample data including at least one sample image and descriptive information of a gesture included in each of the at least one sample image, wherein the descriptive information of a gesture included in a sample image is used to characterize the meaning of the gesture included in the sample image, and the at least one sample image is a sample image corresponding to the at least one gesture; the processing module is further configured to encode the first sample image when the sample data contains descriptive information of a first gesture to obtain the hand features corresponding to the first gesture, wherein the first gesture is one of the at least one gesture, and the first sample image is a sample image corresponding to the first gesture.
[0016] Optionally, the determining module is further configured to determine a preset hand feature as the hand feature corresponding to the first gesture when the sample data does not contain descriptive information of the first gesture.
[0017] Optionally, the processing module is specifically configured to render the first gesture based on the description information of the first gesture to obtain the spatial information of the first gesture. The spatial information of the first gesture includes the three-dimensional coordinates of each of the at least two key points included in the first gesture and the rotation parameters of each key point. The first gesture is one of the at least one gesture.
[0018] Optionally, the determining module is specifically configured to: determine a first loss, which characterizes the degree of inconsistency between the probability distribution of the hand features corresponding to each gesture and a preset normal distribution; the processing module is specifically configured to: input each target image from at least one target image into an initial discriminator to obtain a first probability, which characterizes the probability that each target image is classified as a first label, which characterizes a sample image, and the at least one target image is a target image corresponding to the at least one gesture; the determining module is further configured to: determine a second loss, which characterizes the degree of inconsistency between the pixels of each sample image from at least one sample image and the pixels of the at least one target image, and the at least one sample image is a sample image corresponding to the at least one gesture; the determining module is further configured to: determine a third loss based on the first loss, the first probability, and the second loss; and the processing module is further configured to: update the parameters in the first image generation model based on the third loss to generate the second image generation model.
[0019] Optionally, the processing module is further configured to input each target image into the initial discriminator to obtain a second probability, the second probability being used to characterize the probability that each target image is classified as a second label, the second label being used to characterize non-sample images; the processing module is further configured to input each sample image into the initial discriminator to obtain a third probability, the third probability being used to characterize the probability that each sample image is classified as the first label; the determining module is further configured to determine a fourth loss based on the second probability and the third probability; the processing module is further configured to update the parameters in the initial discriminator based on the fourth loss to generate a target discriminator.
[0020] According to a fourth aspect of the present disclosure, an image generation apparatus is provided. The apparatus may include: a determining module and a processing module; the determining module is configured to determine spatial information of a preset gesture and preset hand features, the spatial information of the preset gesture being used to characterize the positional relationship between at least two key points included in the preset gesture, and the preset hand features being used to characterize attribute information of the hand; the processing module is configured to input the spatial information of the preset gesture and the preset hand features into a second image generation model to obtain a target generated image, the second image generation model being trained based on any of the optional image generation model training methods of the first aspect described above.
[0021] Optionally, the image generation apparatus further includes an acquisition module; the acquisition module is configured to acquire description information of the preset gesture, the description information of the preset gesture being used to characterize the meaning of the preset gesture; the processing module is further configured to perform rendering processing on the preset gesture based on the description information of the preset gesture to obtain spatial information of the preset gesture, the spatial information of the preset gesture including the three-dimensional coordinates of each of the at least two key points included in the preset gesture and the rotation parameters of each key point.
[0022] According to a fifth aspect of the present disclosure, an electronic device is provided, which may include: a processor and a memory configured to store processor-executable instructions; wherein the processor is configured to execute the instructions to implement a training method for any of the optional image generation models in the first aspect above, or to implement any of the optional image generation methods in the second aspect above.
[0023] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, on which instructions are stored, such that when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform either the training method of the optional image generation model in the first aspect or the optional image generation method in the second aspect.
[0024] According to a seventh aspect of the present disclosure, a computer program product is provided, the computer program product including computer instructions that, when executed on a processor of an electronic device, cause the electronic device to perform a training method for any of the optional image generation models in the first aspect, or to perform any of the optional image generation methods in the second aspect described above.
[0025] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:
[0026] Based on any of the above aspects, in this disclosure, an electronic device can determine the spatial information and hand features of at least one gesture, and input the spatial information of each gesture and the corresponding hand features into a first image generation model to obtain a target image corresponding to the at least one gesture; then, the electronic device can train the first image generation model based on the hand features corresponding to each gesture, the target image corresponding to the at least one gesture, and a sample image corresponding to the at least one gesture to generate a second image generation model. In this disclosure, since hand features are used to characterize the attribute information of the hand, this attribute information can characterize the style (or pattern) of the hand; and the target image corresponding to the at least one gesture is a new image generated by the electronic device based on the first image generation model, and the sample image corresponding to the at least one gesture can be understood as a real image. Thus, by training the first image generation model based on the hand features corresponding to each gesture, the new image generated by the first image generation model, and the real image, the electronic device can generate a second image generation model with higher realism. Specifically, the image generated by the electronic device based on the second image generation model has higher realism and is closer to a real image, which can improve the effectiveness of image generation and increase the realism of the image.
[0027] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0028] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0029] Figure 1 A flowchart illustrating a training method for an image generation model provided in an embodiment of this disclosure is shown.
[0030] Figure 2 A flowchart illustrating another image generation model training method provided in this disclosure is shown.
[0031] Figure 3 A flowchart illustrating another image generation model training method provided in this disclosure is shown.
[0032] Figure 4 A flowchart illustrating another image generation model training method provided in this disclosure is shown.
[0033] Figure 5 A flowchart illustrating another image generation model training method provided in this disclosure is shown.
[0034] Figure 6 A flowchart illustrating another image generation model training method provided in this disclosure is shown.
[0035] Figure 7 A schematic flowchart of an image generation method provided in an embodiment of this disclosure is shown;
[0036] Figure 8 A flowchart illustrating yet another image generation method provided in an embodiment of this disclosure is shown;
[0037] Figure 9 A schematic diagram of the structure of a training device for an image generation model provided in an embodiment of this disclosure is shown;
[0038] Figure 10 A schematic diagram of the structure of a training device for another image generation model provided in an embodiment of this disclosure is shown;
[0039] Figure 11 A schematic diagram of the structure of an image generation apparatus provided in an embodiment of this disclosure is shown;
[0040] Figure 12 A schematic diagram of the structure of another image generation apparatus provided in an embodiment of the present disclosure is shown. Detailed Implementation
[0041] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0042] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0043] It should also be understood that the term "comprising" indicates the presence of the described feature, whole, step, operation, element and / or component, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements and / or components.
[0044] It should be noted that the user information (including but not limited to user device information, user personal information, user behavior information, etc.) and data (including but not limited to sample data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0045] In related technologies, gesture images generated based on computer graphics and texture mapping tend to have a strong sense of artificiality and are difficult to closely resemble real images, which may reduce the effectiveness of image generation.
[0046] Based on this, embodiments of this disclosure provide a training method, an image generation method, and an apparatus for an image generation model. Since hand features are used to characterize hand attribute information, this attribute information can characterize the hand's style (or pattern); and the target image corresponding to at least one gesture is a new image generated by the electronic device based on the first image generation model, while the sample image corresponding to at least one gesture can be understood as a real image. Thus, by training the first image generation model based on the hand features corresponding to each gesture, the new image generated by the first image generation model, and the real image, the electronic device can generate a second image generation model with higher realism. Specifically, the images generated by the electronic device based on the second image generation model have higher realism and are closer to real images, thereby improving the effectiveness of image generation and enhancing the realism of the images.
[0047] The image generation model training method, image generation method, and apparatus provided in this disclosure are applied to image generation scenarios (specifically, generating images containing gestures). When an electronic device determines the spatial information and hand features of at least one gesture, it can train a first image generation model according to the method provided in this disclosure to generate a second image generation model. The electronic device can then input the spatial information of a gesture and a hand feature into the second image generation model to obtain the target generated image.
[0048] The training method for the image generation model and the image generation method provided in this disclosure are illustrated below with reference to the accompanying drawings:
[0049] For example, the electronic device executing the training method of the image generation model and the image generation method provided in the embodiments of this disclosure can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) / virtual reality (VR) device, etc. This disclosure does not impose any special limitations on the specific form of the electronic device. It can interact with the user through one or more methods such as keyboard, touchpad, touch screen, remote control, voice interaction or handwriting device.
[0050] like Figure 1 As shown, the training method for the image generation model provided in this embodiment may include S101-S103.
[0051] S101, The electronic device determines spatial information of at least one gesture and hand features.
[0052] Among them, the spatial information of a gesture is used to characterize the positional relationship between at least two key points included in the gesture, and the hand feature is used to characterize the attribute information of the hand.
[0053] It should be understood that a gesture is a posture of the hand, which includes at least two key points that are key points (or skeletal points) of the hand.
[0054] In one implementation of this disclosure, the attribute information of the hand may include at least one of the hand's color, hand brightness, and hand texture.
[0055] S102. The electronic device inputs the spatial information of each gesture and the hand features corresponding to each gesture into the first image generation model to obtain a target image corresponding to at least one gesture.
[0056] It should be understood that the spatial information of a gesture and a hand feature (specifically, a hand feature corresponding to the gesture) correspond to a target image (i.e., the target image corresponding to the gesture).
[0057] In one scenario, a gesture can correspond to a single hand feature. In this case, the electronic device obtains the same number of target images corresponding to the at least one gesture as the number of the at least one gesture, based on the spatial information of each gesture and the corresponding hand feature.
[0058] In another scenario, a gesture can correspond to at least two hand features. In this case, the electronic device obtains a greater number of target images corresponding to the at least one gesture than the number of the at least one gesture, based on the spatial information of each gesture and the corresponding hand features.
[0059] S103. The electronic device trains the first image generation model based on the hand features corresponding to each gesture, the target image corresponding to at least one gesture, and the sample image corresponding to at least one gesture, and generates the second image generation model.
[0060] Based on the description of the above embodiments, it should be understood that the hand features corresponding to a gesture are used to characterize this attribute information. This attribute information can characterize the style (or pattern) of the hand.
[0061] It is understood that the first image generation model (or the second image generation model) is a neural network model used to generate images. The target image corresponding to the at least one gesture is a new image generated by the electronic device based on the first image generation model, and the sample image corresponding to the at least one gesture can be understood as a real image.
[0062] In one alternative implementation, the sample image corresponding to the at least one gesture is an image included in the sample data. The electronic device can obtain the sample image corresponding to the at least one gesture by acquiring the sample data.
[0063] It is understood that, for the target image corresponding to the at least one gesture and the sample image corresponding to the at least one gesture, the gesture included in the target image corresponding to a gesture is the same as the gesture included in the sample image corresponding to the gesture.
[0064] In this embodiment, the electronic device trains the first image generation model based on the hand features corresponding to each gesture, a new image generated by the first image generation model, and a real image, thereby generating a second image generation model with higher realism. Specifically, the images generated by the electronic device based on the second image generation model have higher realism and are closer to real images, which can improve the effectiveness of image generation and enhance the realism of the images.
[0065] Optionally, the first image generation model (or the second image generation model) described above can be a generative adversarial network (GAN).
[0066] The technical solution provided by the above embodiments can bring at least the following beneficial effects: As shown in S101-S103, the electronic device can determine the spatial information and hand features of at least one gesture, and input the spatial information of each gesture and the corresponding hand features into a first image generation model to obtain a target image corresponding to the at least one gesture; then the electronic device can train the first image generation model based on the hand features corresponding to each gesture, the target image corresponding to the at least one gesture, and the sample image corresponding to the at least one gesture to generate a second image generation model. In this disclosure, since hand features are used to characterize the attribute information of the hand, the attribute information can characterize the style (or pattern) of the hand; and the target image corresponding to at least one gesture is a new image generated by the electronic device based on the first image generation model, and the sample image corresponding to at least one gesture can be understood as a real image. Thus, by training the first image generation model based on the hand features corresponding to each gesture, the new image generated by the first image generation model, and the real image, the electronic device can generate a second image generation model with higher realism. Specifically, the image generated by the electronic device based on the second image generation model has higher realism and is closer to the real image, which can improve the effectiveness of image generation and improve the realism of the image.
[0067] Combination Figure 1 ,like Figure 2 As shown, the training method for the image generation model provided in this embodiment may further include S104-S105.
[0068] S104. Electronic devices acquire sample data.
[0069] The sample data includes at least one sample image and descriptive information about the gestures included in each of the at least one sample image. The descriptive information about the gestures included in a sample image is used to characterize the meaning of the gestures included in the sample image, and the at least one sample image is a sample image corresponding to the at least one gesture mentioned above.
[0070] For example, the descriptive information of a gesture can be "victory", "amazing", "heart" or "OK", etc.
[0071] S105. When the sample data contains descriptive information of the first gesture, the electronic device encodes the first sample image to obtain the hand features corresponding to the first gesture.
[0072] Wherein, the first gesture is one of the above-mentioned at least one gesture, and the first sample image is a sample image corresponding to the first gesture.
[0073] It should be understood that when the sample data contains descriptive information of the first gesture, it indicates that the sample data contains a sample image corresponding to the first gesture, that is, a real image corresponding to the first gesture. At this time, the electronic device can encode the real image corresponding to the first gesture (i.e., the first sample image) to obtain the hand features corresponding to the first gesture.
[0074] In one alternative implementation, the electronic device can encode the first sample image based on a certain style encoder to obtain the hand features corresponding to the first gesture.
[0075] The technical solution provided by the above embodiments can bring at least the following beneficial effects: As shown in S104-S105, the electronic device can acquire sample data, which includes at least one sample image and descriptive information of the gesture included in each of the at least one sample image. When the sample data contains descriptive information of a first gesture, it indicates that there is a sample image corresponding to the first gesture in the sample data, that is, there is a real image corresponding to the first gesture. At this time, the electronic device can encode the real image (i.e., the first sample image) corresponding to the first gesture to obtain the hand features corresponding to the first gesture. It can accurately and effectively determine the hand features corresponding to the gesture, thereby improving the accuracy of model training.
[0076] Combination Figure 2 ,like Figure 3 As shown, the training method for the image generation model provided in this embodiment of the present disclosure further includes S106.
[0077] S106. When there is no descriptive information for the first gesture in the sample data, the electronic device will determine the preset hand feature as the hand feature corresponding to the first gesture.
[0078] It should be understood that when the sample data does not contain descriptive information for the first gesture, it means that there is no sample image corresponding to the first gesture in the sample data, that is, there is no real image corresponding to the first gesture. At this time, the electronic device can randomly assign a hand feature (i.e., the preset hand feature) to the first gesture, and can quickly and effectively determine the hand feature corresponding to the gesture.
[0079] The technical solution provided by the above embodiments can bring at least the following beneficial effects: As can be seen from S106, when there is no descriptive information of the first gesture in the sample data, it means that there is no sample image corresponding to the first gesture in the sample data, that is, there is no real image corresponding to the first gesture. At this time, the electronic device can randomly assign a hand feature to the first gesture, which can quickly and effectively determine the hand feature corresponding to the gesture, thereby improving the efficiency of model training.
[0080] Combination Figure 2 ,like Figure 4 As shown, in one implementation of the present disclosure, the electronic device determining at least one potential spatial information may specifically include S1011.
[0081] S1011. The electronic device renders the first gesture based on the description information of the first gesture to obtain the spatial information of the first gesture.
[0082] The spatial information of the first gesture includes the three-dimensional coordinates of each of the at least two key points included in the first gesture and the rotation parameters of each key point, and the first gesture is one of the above-mentioned at least one gesture.
[0083] Alternatively, the electronic device can render the first gesture using a rendering function.
[0084] In one alternative implementation, the spatial information of a gesture can be understood as the result of voxel representation in a preset space, and the spatial information (or voxel) of the gesture needs to be aligned with the image corresponding to the gesture.
[0085] The technical solution provided by the above embodiments can bring at least the following beneficial effects: As can be seen from S1011, the electronic device can render the gesture based on the description information of the gesture. The spatial information of the gesture obtained includes the three-dimensional coordinates of each of the at least two key points included in the gesture and the rotation parameters of each key point. The three-dimensional coordinates of each key point and the rotation parameters of each key point can more accurately characterize the spatial information of the first gesture, thereby improving the accuracy of model training.
[0086] Combination Figure 1 ,like Figure 5 As shown, in one implementation of this disclosure, the electronic device trains a first image generation model based on the hand features corresponding to each gesture, the target image corresponding to at least one gesture, and the sample image corresponding to at least one gesture, and generates a second image generation model, which may specifically include S1031-S1035.
[0087] S1031, The first loss is determined for electronic equipment.
[0088] The first loss is used to characterize the degree of inconsistency between the probability distribution of the hand features corresponding to each gesture in the at least one gesture and the preset normal distribution.
[0089] It should be understood that the probability distribution of the hand features corresponding to each gesture can be a normal distribution (or a Gaussian distribution).
[0090] In one alternative implementation, the preset normal distribution can be a standard normal distribution, that is, a normal distribution with a mean of 0 and a standard deviation of 1, which can be denoted as N(0,1).
[0091] Optionally, the electronic device may determine the first loss as the KL divergence (or relative entropy) between the probability distribution of the hand features corresponding to each gesture and a preset normal distribution.
[0092] That is, the electronic device can determine that the first loss satisfies the following formula:
[0093] L1 = KL(A, B)
[0094] Where L represents the first loss, A represents the probability distribution of the hand features corresponding to each gesture, B represents the preset normal distribution, and KL(A,B) represents the KL divergence between the probability distribution of the hand features corresponding to each gesture and the preset normal distribution.
[0095] S1032. The electronic device inputs each of the at least one target images into the initial discriminator to obtain a first probability.
[0096] The first probability is used to characterize the probability that each target image is identified as the first label, the first label is used to characterize the sample image, and the at least one target image is the target image corresponding to the at least one gesture mentioned above.
[0097] In this embodiment of the disclosure, a sample image can be understood as a real image.
[0098] It should be understood that the initial discriminator is used to discriminate each input image (including each target image and each of the aforementioned sample images) to determine the probability that each image is the first label (i.e., the sample image) and the probability that each image is the second label (i.e., the non-sample image).
[0099] It is understandable that the first probability can be the sum of the probabilities of each target image being identified as the first label.
[0100] S1033, Electronic equipment determines the second loss.
[0101] The second loss is used to characterize the degree of inconsistency between the pixels of each sample image in at least one sample image and the pixels of the at least one target image, wherein the at least one sample image is a sample image corresponding to the at least one gesture mentioned above.
[0102] In one alternative implementation, the electronic device can determine the second loss based on the L1 norm loss function (i.e., the minimum absolute value error). Specifically, the electronic device can determine that the second loss satisfies the following formula:
[0103]
[0104] Where L2 represents the second loss, P i P represents the number of pixels in the i-th sample image. i ' represents the number of pixels in the target image corresponding to the i-th sample image, and n represents the number of at least one sample image, 1≤i≤n.
[0105] S1034. The electronic device determines the third loss based on the first loss, the first probability, and the second loss.
[0106] Optionally, the electronic device may determine the third loss as the sum of the first loss, the first probability, and the second loss.
[0107] S1035. The electronic device updates the parameters in the first image generation model based on the third loss and generates the second image generation model.
[0108] The technical solution provided by the above embodiments can bring at least the following beneficial effects: As shown in S1031-S1035, the electronic device can determine the first loss and the second loss, and input each target image in at least one target image into the initial discriminator to obtain the first probability; then the electronic device can determine the third loss based on the first loss, the second loss and the first probability, and update the parameters in the first image generation model based on the third loss to generate the second image generation model. In this embodiment of the present disclosure, since the first loss is used to characterize the degree of inconsistency between the probability distribution of the hand features corresponding to each gesture and the preset normal distribution, the second loss is used to characterize the degree of inconsistency between the pixels of each sample image in at least one sample image and the pixels of the at least one target image, and the first probability is used to characterize the probability that each target object is judged as the first label (i.e., the sample image), the electronic device can accurately and effectively update the parameters in the first image generation model based on the third loss, and can train a second image generation model with higher accuracy and realism.
[0109] Combination Figure 5 ,like Figure 6 As shown, the training method for the image generation model provided in this embodiment may further include S107-S110.
[0110] S107. The electronic device will input the initial discriminator for each target image to obtain the second probability.
[0111] The second probability is used to characterize the probability that each target image is identified as the second label, and the second label is used to characterize non-sample images.
[0112] Based on the description of the above embodiments, it should be understood that the initial discriminator is used to discriminate each input image (including each target image and each sample image) to determine the probability that each image is a first label (i.e., a sample image) and the probability that each image is a second label (i.e., a non-sample image).
[0113] In this embodiment of the disclosure, a non-sample image can be understood as a newly generated image, specifically a new image generated based on an image generation model (e.g., a first image generation model or a second image generation model).
[0114] It is understandable that the second probability can be the sum of the probabilities of each target image being identified as the second label.
[0115] S108. The electronic device inputs each sample image into the initial discriminator to obtain the third probability.
[0116] The third probability is used to characterize the probability that each sample image is identified as the first label mentioned above.
[0117] It should be understood that the third probability can be the sum of the probabilities of each sample image being classified as the first label.
[0118] S109. The electronic device determines the fourth loss based on the second and third probabilities.
[0119] Optionally, the electronic device may determine the fourth loss as the sum of the second probability and the third probability.
[0120] S110. The electronic device updates the parameters in the initial discriminator based on the fourth loss and generates a target discriminator.
[0121] The technical solution provided by the above embodiments can bring at least the following beneficial effects: As shown in S107-S110, the electronic device can input each target image into the initial discriminator to obtain a second probability, and input each sample image into the initial discriminator to obtain a third probability; then the electronic device can determine a fourth loss based on the second probability and the third probability, and update the parameters in the initial discriminator based on the fourth loss to generate a target discriminator. In this embodiment, since the second probability is used to characterize the probability that each target image is judged as the second label (i.e., a non-sample image), and the third probability is used to characterize the probability that each sample image is judged as the first label (i.e., a sample image), the electronic device can accurately and effectively update the parameters in the initial discriminator based on the fourth loss, and can train a target discriminator with higher accuracy. This target discriminator can more accurately determine whether each input image is a real image.
[0122] In one implementation of this disclosure, the initial discriminator may include a first initial discriminator and a second initial discriminator, wherein the first initial discriminator is an initial discriminator corresponding to hand features, and the second initial discriminator is a discriminator corresponding to spatial information.
[0123] It should be understood that the first initial discriminator is used to supervise the authenticity of the optimized hand features, and the second initial discriminator is used to supervise the optimized hand shape.
[0124] like Figure 7 As shown, the image generation method provided in this embodiment may include S201-S202.
[0125] S201, The electronic device determines the spatial information of the preset gesture and the preset hand features.
[0126] The spatial information of the preset gesture is used to characterize the positional information between at least two key points included in the preset gesture, and the preset hand features are used to characterize the attribute information of the hand. At least one of color, hand features, and hand texture is included.
[0127] Based on the description of the above embodiments, it should be understood that the attribute information of the hand may include at least one of the hand's color, hand brightness, and hand texture, and this attribute information may characterize the hand's style (or pattern).
[0128] In one alternative implementation, the preset gesture can be one of the above-mentioned gestures.
[0129] S202, The electronic device inputs the spatial information of the preset gesture and the preset hand features into the second image generation model to obtain the target generated image.
[0130] The second image generation model is trained based on the image generation model training method provided in the above-described embodiments of this disclosure.
[0131] Specifically, the second image generation model is generated by the electronic device training the first image generation model based on the hand features corresponding to each of the at least one gesture, the target image corresponding to the at least one gesture, and the sample image corresponding to the at least one gesture. The target image corresponding to the at least one gesture is obtained by the electronic device inputting the spatial information of each gesture and the hand features of each gesture into the first image generation model, and the sample image corresponding to the at least one gesture is an image included in the sample data.
[0132] It is understandable that the first image generation model is the image generation model in its initial state, while the second image generation model is the image generation model that has already been trained.
[0133] The technical solution provided by the above embodiments can bring at least the following beneficial effects: As shown in S201-S202, the electronic device can determine the spatial information of a preset gesture and preset hand features; then the electronic device can input the spatial information of the preset gesture and the preset hand features into the second image generation model to obtain a target generated image. In this embodiment of the present disclosure, since a hand feature is used to characterize the attribute information of the hand, the attribute information can characterize the style (or pattern) of the hand, and the second image generation model has high realism; thus, when the electronic device inputs the preset hand features and the spatial information of the preset gesture into the second image generation model, it can generate a target generated image with higher realism, which is closer to the real image, thereby improving the effectiveness of image generation and enhancing the realism of the image.
[0134] Combination Figure 7 ,like Figure 8 As shown, the image generation method provided in this embodiment may further include S203-S204.
[0135] S203, The electronic device acquires description information of the preset gesture.
[0136] The description information of the preset gesture is used to characterize the meaning of the preset gesture.
[0137] S204. The electronic device renders the preset gesture based on the description information of the preset gesture to obtain the spatial information of the preset gesture.
[0138] The spatial information of the preset gesture includes the three-dimensional coordinates of each of the at least two key points included in the preset gesture, as well as the rotation parameters of each key point.
[0139] It should be noted that the explanation of the spatial information of the preset gesture obtained by rendering the preset gesture based on the description information of the preset gesture by the electronic device is the same as or similar to the description of the spatial information of the first gesture obtained by rendering the first gesture based on the description information of the first gesture by the electronic device mentioned above, and will not be repeated here.
[0140] The technical solutions provided by the above embodiments can bring at least the following beneficial effects: As shown in S203-S204, the electronic device can acquire the description information of the preset gesture, and perform rendering processing on the preset gesture based on the description information to obtain the spatial information of the preset gesture. The spatial information of the preset gesture includes the three-dimensional coordinates of each of the at least two key points included in the preset gesture and the rotation parameters of each key point. In this embodiment of the present disclosure, the three-dimensional coordinates of each key point and the rotation parameters of each key point can more accurately characterize the spatial information of the preset gesture, thereby improving the accuracy of image generation.
[0141] It is understood that, in practical implementation, the electronic device described in the embodiments of this disclosure may include one or more hardware structures and / or software modules for implementing the training method and image generation method of the aforementioned corresponding image generation model. These hardware structures and / or software modules can constitute an electronic device. Those skilled in the art should readily recognize that, based on the algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is implemented in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0142] Based on this understanding, the present disclosure also provides a training device for an image generation model. Figure 9 A schematic diagram of the structure of a training apparatus for an image generation model provided in an embodiment of this disclosure is shown. Figure 9 As shown, the training device 10 for the image generation model may include a determination module 101 and a processing module 102.
[0143] The determining module 101 is configured to determine spatial information of at least one gesture and hand features, wherein the spatial information of a gesture is used to characterize the positional relationship between at least two key points included in the gesture, and the hand features are used to characterize the attribute information of the hand.
[0144] The processing module 102 is configured to input the spatial information of each gesture and the hand features corresponding to each gesture into a first image generation model to obtain a target image corresponding to the at least one gesture.
[0145] The processing module 102 is further configured to train the first image generation model based on the hand features corresponding to each gesture, the target image corresponding to the at least one gesture, and the sample image corresponding to the at least one gesture, to generate a second image generation model.
[0146] Optionally, the training device 10 for the image generation model also includes an acquisition module 103.
[0147] The acquisition module 103 is configured to acquire sample data, which includes at least one sample image and descriptive information of the gesture included in each of the at least one sample image, wherein the descriptive information of the gesture included in a sample image is used to characterize the meaning of the gesture included in the sample image, and the at least one sample image is a sample image corresponding to the at least one gesture.
[0148] The processing module 102 is further configured to encode the first sample image when the sample data contains descriptive information of the first gesture, to obtain the hand features corresponding to the first gesture, wherein the first gesture is one of the at least one gesture, and the first sample image is a sample image corresponding to the first gesture.
[0149] Optionally, the determining module 101 is further configured to determine a preset hand feature as the hand feature corresponding to the first gesture when the sample data does not contain descriptive information of the first gesture.
[0150] Optionally, the processing module 102 is specifically configured to perform rendering processing on the first gesture based on the description information of the first gesture to obtain the spatial information of the first gesture. The spatial information of the first gesture includes the three-dimensional coordinates of each of the at least two key points included in the first gesture and the rotation parameters of each key point. The first gesture is one of the at least one gesture.
[0151] Optionally, the determining module 101 is specifically configured to determine a first loss, which is used to characterize the degree of inconsistency between the probability distribution of the hand features corresponding to each gesture and a preset normal distribution.
[0152] The processing module 102 is specifically configured to input each of the at least one target images into an initial discriminator to obtain a first probability. The first probability is used to characterize the probability that each target image is judged as a first label. The first label is used to characterize a sample image. The at least one target image is a target image corresponding to the at least one gesture.
[0153] The determining module 101 is further configured to determine a second loss, which is used to characterize the degree of inconsistency between the pixels of each sample image in at least one sample image and the pixels of the at least one target image, wherein the at least one sample image is a sample image corresponding to the at least one gesture.
[0154] The determination module 101 is further configured to determine the third loss based on the first loss, the first probability, and the second loss.
[0155] The processing module 102 is further configured to update the parameters in the first image generation model based on the third loss, and generate the second image generation model.
[0156] Optionally, the processing module 102 is further configured to input each target image into the initial discriminator to obtain a second probability, the second probability being used to characterize the probability that each target image is judged as a second label, the second label being used to characterize non-sample images.
[0157] The processing module 102 is also configured to input each sample image into the initial discriminator to obtain a third probability, which is used to characterize the probability that each sample image is judged as the first label.
[0158] The determination module 101 is also configured to determine a fourth loss based on the second probability and the third probability.
[0159] The processing module 102 is also configured to update the parameters in the initial discriminator based on the fourth loss, and generate a target discriminator.
[0160] As described above, the embodiments of this disclosure can divide the training device for the image generation model into functional modules according to the above method examples. The integrated modules can be implemented in hardware or as software functional modules. Furthermore, it should be noted that the module division in these embodiments is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into a single processing module.
[0161] The specific methods by which each module performs its operation and the beneficial effects of the training apparatus for the image generation model in the above embodiments have been described in detail in the foregoing method embodiments, and will not be repeated here.
[0162] Figure 10 This is a schematic diagram of the structure of a training device for another image generation model provided in this disclosure. For example... Figure 10 The training apparatus 20 for the image generation model may include at least one processor 201 and a memory 203 for storing processor-executable instructions. The processor 201 is configured to execute the instructions in the memory 203 to implement the training method for the image generation model in the above embodiments.
[0163] In addition, the training device 20 for the image generation model may also include a communication bus 202 and at least one communication interface 204.
[0164] Processor 201 may be a processor (central processing unit, CPU), microprocessor unit, ASIC, or one or more integrated circuits for controlling the execution of programs according to the present disclosure.
[0165] The communication bus 202 may include a path for transmitting information between the aforementioned components.
[0166] Communication interface 204 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.
[0167] The memory 203 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory may exist independently and be connected to the processing unit via a bus. The memory may also be integrated with the processing unit.
[0168] The memory 203 stores instructions for executing the present invention, and the processor 201 controls the execution of these instructions. The processor 201 executes the instructions stored in the memory 203 to implement the functions of the method disclosed herein.
[0169] In a specific implementation, as one embodiment, the processor 201 may include one or more CPUs, for example... Figure 10 CPU0 and CPU1 in the CPU.
[0170] In a specific implementation, as one example, the training device 20 for the image generation model may include multiple processors, such as... Figure 10 Processors 201 and 207 are described herein. Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0171] In a specific implementation, as one embodiment, the training device 20 for the image generation model may further include an output device 205 and an input device 206. The output device 205 communicates with the processor 201 and can display information in various ways. For example, the output device 205 may be a liquid crystal display (LCD), a light-emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 206 communicates with the processor 201 and can accept user input in various ways. For example, the input device 206 may be a mouse, keyboard, touchscreen device, or sensing device, etc.
[0172] Figure 11 This is a structural example diagram of an image generation apparatus provided in this disclosure. Figure 11 As shown, the image generation device 30 may include a determining module 301 and a processing module 302.
[0173] The determining module 301 is configured to determine the spatial information of a preset gesture and preset hand features. The spatial information of the preset gesture is used to characterize the positional relationship between at least two key points included in the preset gesture, and the preset hand features are used to characterize the attribute information of the hand.
[0174] The processing module 302 is configured to input the spatial information of the preset gesture and the preset hand features into the second image generation model to obtain the target generated image. The second image generation model is trained based on the training method of the image generation model provided in the above-described embodiments of this disclosure.
[0175] Optionally, the image generation apparatus 30 further includes an acquisition module 303.
[0176] The acquisition module 303 is configured to acquire the description information of the preset gesture, which is used to characterize the meaning of the preset gesture.
[0177] The processing module 302 is also configured to render the preset gesture based on the description information of the preset gesture to obtain the spatial information of the preset gesture. The spatial information of the preset gesture includes the three-dimensional coordinates of each of the at least two key points included in the preset gesture and the rotation parameters of each key point.
[0178] Figure 12 This is a schematic diagram of another image generation apparatus provided in this disclosure. For example... Figure 12The image generation apparatus 40 may include at least one processor 401 and a memory 403 for storing processor-executable instructions. The processor 401 is configured to execute the instructions in the memory 403 to implement the image generation method described in the above embodiments.
[0179] In addition, the image generation apparatus 40 may also include a communication bus 402 and at least one communication interface 404.
[0180] Processor 401 may be a CPU, a microprocessor unit, an ASIC, or one or more integrated circuits for controlling the execution of programs according to the present disclosure.
[0181] The communication bus 402 may include a path for transmitting information between the aforementioned components.
[0182] Communication interface 404 is used with any transceiver or similar device for communicating with other devices or communication networks, such as Ethernet, RAN, WLAN, etc.
[0183] The memory 403 can be ROM or other types of static storage devices capable of storing static information and instructions, RAM or other types of dynamic storage devices capable of storing information and instructions, or it can be EEPROM, CD-ROM or other optical disc storage, optical disk storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. The memory can exist independently and be connected to the processing unit via a bus. The memory can also be integrated with the processing unit.
[0184] The memory 403 stores instructions for executing the present invention, and the processor 401 controls the execution of these instructions. The processor 401 executes the instructions stored in the memory 403 to implement the functions of the method disclosed herein.
[0185] In a specific implementation, as one example, processor 401 may include one or more CPUs, for example... Figure 12 CPU0 and CPU1 in the CPU.
[0186] In a specific implementation, as one example, the image generation device 40 may include multiple processors, such as... Figure 12 Processors 401 and 407 are described herein. Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).
[0187] In a specific implementation, as one embodiment, the image generation apparatus 40 may further include an output device 405 and an input device 406. The output device 405 communicates with the processor 401 and can display information in various ways. For example, the output device 405 may be an LCD, LED display device, CRT display device, or projector, etc. The input device 406 communicates with the processor 401 and can accept user input in various ways. For example, the input device 406 may be a mouse, keyboard, touchscreen device, or sensing device, etc.
[0188] Those skilled in the art will understand that the above Figure 10 as well as Figure 12 The structure shown does not constitute a limitation on the training device for the image generation model or the image generation device. It may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0189] In addition, this disclosure also provides a computer-readable storage medium including instructions that, when executed by a processor of an electronic device, cause the electronic device to perform the training method for the image generation model and the image generation method as provided in the above embodiments.
[0190] In addition, this disclosure also provides a computer program product, including instructions that, when executed by a processor of an electronic device, cause the electronic device to perform the training method for the image generation model and the image generation method as provided in the above embodiments.
[0191] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
Claims
1. A training method for an image generation model, characterized in that, include: Determine the spatial information of at least one gesture and hand features, wherein the spatial information of a gesture is used to characterize the positional relationship between at least two key points included in the gesture, and the hand features are used to characterize the attribute information of the hand; The spatial information of each gesture and the hand features corresponding to each gesture are input into the first image generation model to obtain the target image corresponding to the at least one gesture; A first loss and a second loss are determined. The first loss is used to characterize the degree of inconsistency between the probability distribution of the hand features corresponding to each gesture and a preset normal distribution. The second loss is used to characterize the degree of inconsistency between the pixels of each sample image in at least one sample image and the pixels of at least one target image. The at least one sample image is a sample image corresponding to the at least one gesture. Each of the at least one target images is input into an initial discriminator to obtain a first probability. The first probability is used to characterize the probability that each target image is judged as a first label. The first label is used to characterize a sample image. The at least one target image is a target image corresponding to the at least one gesture. The sum of the first loss, the first probability, and the second loss is determined as the third loss; Based on the third loss, the parameters in the first image generation model are updated to generate the second image generation model.
2. The training method for the image generation model according to claim 1, characterized in that, The method further includes: Acquire sample data, which includes at least one sample image and descriptive information of the gesture included in each of the at least one sample image, wherein the descriptive information of the gesture included in a sample image is used to characterize the meaning of the gesture included in the sample image, and the at least one sample image is a sample image corresponding to the at least one gesture; When the sample data contains description information of a first gesture, the first sample image is encoded to obtain the hand features corresponding to the first gesture. The first gesture is one of the at least one gesture, and the first sample image is a sample image corresponding to the first gesture.
3. The training method for the image generation model according to claim 2, characterized in that, The method further includes: When the sample data does not contain description information of the first gesture, a preset hand feature is determined as the hand feature corresponding to the first gesture.
4. The training method for the image generation model according to any one of claims 1-3, characterized in that, Determining the spatial information of at least one gesture includes: The first gesture is rendered based on the description information of the first gesture to obtain the spatial information of the first gesture. The spatial information of the first gesture includes the three-dimensional coordinates of each key point in at least two key points included in the first gesture and the rotation parameters of each key point. The first gesture is one of the at least one gesture.
5. The training method for the image generation model according to claim 1, characterized in that, The method further includes: Each target image is input into the initial discriminator to obtain a second probability. The second probability is used to characterize the probability that each target image is judged as a second label. The second label is used to characterize non-sample images. Each sample image is input into the initial discriminator to obtain a third probability, which is used to characterize the probability that each sample image is judged as the first label; The sum of the second probability and the third probability is determined as the fourth loss; Based on the fourth loss, the parameters in the initial discriminator are updated to generate the target discriminator.
6. An image generation method, characterized in that, include: Determine the spatial information of a preset gesture and preset hand features. The spatial information of the preset gesture is used to characterize the positional relationship between at least two key points included in the preset gesture, and the preset hand features are used to characterize the attribute information of the hand. The spatial information of the preset gesture and the preset hand features are input into the second image generation model to obtain the target generated image. The second image generation model is trained based on the training method of the image generation model according to any one of claims 1-5.
7. The image generation method according to claim 6, characterized in that, The method further includes: Obtain the description information of the preset gesture, which is used to characterize the meaning of the preset gesture; The preset gesture is rendered based on its description information to obtain its spatial information. The spatial information of the preset gesture includes the three-dimensional coordinates of each of the at least two key points included in the preset gesture and the rotation parameters of each key point.
8. A training device for an image generation model, characterized in that, include: Determine the module and the processing module; The determining module is configured to determine spatial information of at least one gesture and hand features, wherein the spatial information of a gesture is used to characterize the positional relationship between at least two key points included in the gesture, and the hand features are used to characterize the attribute information of the hand. The processing module is configured to input the spatial information of each gesture and the hand features corresponding to each gesture into a first image generation model to obtain a target image corresponding to the at least one gesture; The processing module is further configured to determine a first loss and a second loss, wherein the first loss is used to characterize the degree of inconsistency between the probability distribution of the hand features corresponding to each gesture and a preset normal distribution; and the second loss is used to characterize the degree of inconsistency between the pixels of each sample image in at least one sample image and the pixels of at least one target image, wherein the at least one sample image is a sample image corresponding to the at least one gesture. The processing module is further configured to input each of the at least one target images into an initial discriminator to obtain a first probability, wherein the first probability is used to characterize the probability that each target image is judged as a first label, the first label is used to characterize a sample image, and the at least one target image is a target image corresponding to the at least one gesture; The processing module is further configured to determine the sum of the first loss, the first probability, and the second loss as the third loss; The processing module is further configured to update the parameters in the first image generation model based on the third loss, and generate a second image generation model.
9. An image generation apparatus, characterized in that, include: Determine the module and the processing module; The determining module is configured to determine the spatial information of a preset gesture and preset hand features. The spatial information of the preset gesture is used to characterize the positional relationship between at least two key points included in the preset gesture, and the preset hand features are used to characterize the attribute information of the hand. The processing module is configured to input the spatial information of the preset gesture and the preset hand features into a second image generation model to obtain a target generated image. The second image generation model is trained based on the training method of the image generation model according to any one of claims 1-5.
10. An electronic device, characterized in that, The electronic device includes: processor; A memory configured to store processor-executable instructions; The processor is configured to execute the instructions to implement the training method of the image generation model as described in any one of claims 1-5, or to implement the image generation method as described in claim 6 or 7.
11. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the training method of the image generation model as described in any one of claims 1-5, or the image generation method as described in claim 6 or 7.
Citation Information
Patent Citations
Virtual human generation method and device based on monocular vision
CN113658303A
Multi-branch deep learning-based 3D face reconstruction model training method and system, and medium
CN114926591A
Hand feature extraction and gesture recognition method, electronic equipment and storage medium
CN115035320A