A method for generating controllable posture and related device

By using the graph diffusion model in pose controllable generation, a variety and correct pose skeleton diagram is generated from natural language descriptions, the problems of limited data and manual adjustment deviation in the prior art are solved, and a high accuracy and diverse pose generation is achieved.

CN119621942BActive Publication Date: 2025-05-13BEIJING SOHU NEW MEDIA INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510154091.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-13
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The existing controllable pose generation methods rely on samples and manual adjustments, which makes it difficult to meet different scenarios and user needs under limited data. Manual dragging of key points is prone to position deviations, resulting in unnatural or deformed image poses.

Method used

A graph diffusion model generates diverse and correct pose skeleton diagrams from natural language descriptions through graph diffusion models. The graph diffusion model consists of a pre-trained stable diffusion model, and a graph convolutional neural network is inserted into the denoising model to process data features in non-Euclidean spaces.

Benefits of technology

It realizes the generation of diverse and correct posture skeleton diagrams from natural language descriptions, improves the accuracy and diversity of controllable posture generation, reduces the need for manual intervention, and improves the efficiency of the generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621942B_ABST
    Figure CN119621942B_ABST
Patent Text Reader

Abstract

The present application discloses a posture controllable generation method and related devices, which relate to the field of generative artificial intelligence technology. The stable diffusion model is trained in advance to obtain a graph diffusion model. The denoising model in the stable diffusion model can process data features of non-Euclidean space (i.e., data features that are relatively scattered but have spatial structure) by inserting a graph convolutional neural network. The first natural language text to be processed is obtained, and a first pure noise heat map is randomly generated; the first natural language text and the first pure noise heat map are input into the graph diffusion model; the first key point heat map output by the graph diffusion model is obtained, and the first key point heat map matches the first natural language text; the first key point heat map is converted into a first posture skeleton map. Based on the graph diffusion model, the present application can generate diverse and correct two-dimensional posture skeleton maps from natural language descriptions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of generative artificial intelligence technology, and in particular to a posture-controllable generation method and related devices. Background Art

[0002] Pose-controllable generation aims to generate character images with specific postures based on the input natural language description, which is widely used in virtual reality, animation production, game development and other fields.

[0003] Existing pose-controllable generation methods usually extract pose skeletons from existing images, or manually drag key points to adjust the skeleton pose, resulting in the pose-controllable generation process being highly dependent on the collected samples and manual quality. Therefore, with limited data, it is difficult to meet different scenarios and user needs. In addition, manually dragging key points is prone to key point position deviations, which may cause the final generated image pose to be unnatural or deformed. Summary of the invention

[0004] In view of the above problems, the present application provides a controllable posture generation method and related devices to achieve the purpose of generating diverse and correct posture skeleton graphs from natural language descriptions. The specific scheme is as follows:

[0005] A first aspect of the present application provides a posture controllable generation method, the posture controllable generation method comprising:

[0006] Obtaining a first natural language text to be processed, and randomly generating a first pure noise heat map;

[0007] Inputting the first natural language text and the first pure noise heat map into a graph diffusion model, wherein the graph diffusion model is obtained by pre-training a stable diffusion model, and a denoising model in the stable diffusion model is inserted into a graph convolutional neural network;

[0008] Obtaining a first key point heat map output by the graph diffusion model, where the first key point heat map matches the first natural language text;

[0009] The first key point heat map is converted into a first posture skeleton map.

[0010] In a possible implementation, the posture controllable generation method further includes:

[0011] A corresponding actual image is generated according to the first posture skeleton image.

[0012] In a possible implementation, the process of pre-training the stable diffusion model to obtain the graph diffusion model includes:

[0013] Acquire a second natural language text and a second posture skeleton graph, wherein the second posture skeleton graph matches the second natural language text;

[0014] Converting the second posture skeleton image into a second key point heat map;

[0015] Performing diffusion training on the stable diffusion model using the second natural language text and the second key point heat map to adjust model parameters of the denoising model in which the graph convolutional neural network is inserted;

[0016] The stable diffusion model after the diffusion training is completed is used as the graph diffusion model.

[0017] In a possible implementation, the step of performing diffusion training on the stable diffusion model using the second natural language text and the second key point heat map includes:

[0018] Entering the forward noisy process, performing multiple noise-adding steps on the second key point heat map to obtain a second pure noise heat map, and recording the first noise-adding amount of each noise-adding step;

[0019] Enter the reverse denoising process to extract the text features of the second natural language text; based on the cross-attention mechanism, fuse the text features with the second pure noise heat map and input them into the stable diffusion model, and trigger the stable diffusion model to output the third key point heat map of this round of training by predicting the second noise amount of each noise step; calculate the loss function value of this round of training according to the first noise amount and the second noise amount; if the loss function value does not meet the corresponding end condition, adjust the network parameters of the denoising model inserted with the graph convolutional neural network according to the loss function value, and return to execute the triggering of the stable diffusion model to output the third key point heat map of this round of training by predicting the second noise amount of each noise step; if the loss function value meets the corresponding end condition, end this round of training.

[0020] In one possible implementation, the graph convolutional neural network is inserted between the downsampling component and the upsampling component of the denoising model.

[0021] A second aspect of the present application provides a posture controllable generation device, the posture controllable generation device comprising:

[0022] A model training module, used for pre-training a stable diffusion model to obtain a graph diffusion model, wherein a denoising model is inserted into a graph convolutional neural network in the stable diffusion model;

[0023] A posture inference module is used to obtain a first natural language text to be processed and randomly generate a first pure noise heat map; input the first natural language text and the first pure noise heat map into the graph diffusion model; obtain a first key point heat map output by the graph diffusion model, the first key point heat map matches the first natural language text; and convert the first key point heat map into a first posture skeleton map.

[0024] In a possible implementation, the model training module is specifically used to:

[0025] Acquire a second natural language text and a second posture skeleton graph, wherein the second posture skeleton graph matches the second natural language text; convert the second posture skeleton graph into a second key point heat map; perform diffusion training on the stable diffusion model using the second natural language text and the second key point heat map to adjust the model parameters of the denoising model in which the graph convolutional neural network is inserted; and use the stable diffusion model after the diffusion training as the graph diffusion model.

[0026] The third aspect of the present application provides a computer program product, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the gesture-controllable generation method of the first aspect or any implementation of the first aspect.

[0027] A fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:

[0028] The memory is used to store computer programs;

[0029] The processor is used to execute the computer program so that the electronic device can implement the gesture-controllable generation method of the above-mentioned first aspect or any implementation manner of the first aspect.

[0030] The fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the posture-controllable generation method of the above-mentioned first aspect or any implementation method of the first aspect.

[0031] By means of the above technical solution, the posture controllable generation method and related devices provided by the present application pre-train the stable diffusion model to obtain the graph diffusion model. The denoising model in the stable diffusion model can process the data features of non-Euclidean space (i.e., relatively scattered data features with spatial structure) by inserting the graph convolutional neural network. Obtain the first natural language text to be processed, and randomly generate a first pure noise heat map; input the first natural language text and the first pure noise heat map into the graph diffusion model; obtain the first key point heat map output by the graph diffusion model, and the first key point heat map matches the first natural language text; convert the first key point heat map into a first posture skeleton map. Based on the graph diffusion model, the present application can generate diverse and correct two-dimensional posture skeleton maps from natural language descriptions, and in the posture controllable generation scenario, increase the additional posture information with diversity and correctness that can be used by the posture controllable generation task. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0033] Figure 1 A flowchart of a posture controllable generation method provided in an embodiment of the present application;

[0034] Figure 2 A partial flow chart of a posture controllable generation method provided in an embodiment of the present application;

[0035] Figure 3 A schematic diagram of another part of the flow chart of a gesture controllable generation method provided in an embodiment of the present application;

[0036] Figure 4 An example of a posture skeleton diagram provided in an embodiment of the present application;

[0037] Figure 5 An example of an actual image provided in an embodiment of the present application;

[0038] Figure 6 Example of pose skeleton graph generated for WGAN-LP R.

[0039] Figure 7 Another example of a pose skeleton graph generated for WGAN-LP R.

[0040] Figure 8 A schematic diagram of the structure of a gesture controllable generation device provided in an embodiment of the present application;

[0041] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation method section of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.

[0043] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0044] The terms "first", "second" etc. in the specification of the application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable in appropriate circumstances, and this is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0045] In the field of image generation, the widespread application of posture-controllable generation technology enables more precise control of the posture of characters when generating images. Usually, these technologies constrain the generation process based on the input posture skeleton data or key point information, so that the generated image has a specific posture. However, in practical applications, there is a problem of insufficient reference samples, especially when it is desired to generate complex or unique postures, the existing posture data resources are difficult to meet the needs.

[0046] To solve the above problems, the present application provides a posture controllable generation method that can flexibly generate diverse and correct posture skeleton graphs from the natural language description of a given posture as reference information for posture controllable generation.

[0047] See also Figure 1 , Figure 1 A flow chart of a posture controllable generation method provided in an embodiment of the present application. Figure 1 As shown, a posture controllable generation method provided in an embodiment of the present application may include steps S101 to S104, and these steps are described in detail below.

[0048] S101, obtaining a first natural language text to be processed, and randomly generating a first pure noise heat map.

[0049] In the embodiment of the present application, after obtaining the natural language text to be processed (i.e., the first natural language text) input by the user, a pure noise heat map can be randomly generated, and the pure noise heat map is a posture skeleton map of pure noise. The first natural language text is the natural language description of the posture skeleton map required by the user, such as "a girl sitting and drinking coffee".

[0050] S102, inputting the first natural language text and the first pure noise heat map into a graph diffusion model, where the graph diffusion model is obtained by pre-training a stable diffusion model, and a denoising model in the stable diffusion model is inserted with a graph convolutional neural network.

[0051] The core component of the stable diffusion model (stable diffusion, SD) is the denoising model (UNet), which is based on the convolutional neural network (CNN). The convolutional neural network is suitable for processing neatly arranged grid structure features (such as pixels of RGB images). However, for some data with spatial structure but lacking obvious grid structure (such as protein molecular structure, posture skeleton graph), the convolutional neural network cannot effectively capture the data features of these non-Euclidean spaces. In this regard, the present application inserts the graph convolution neural network (GCNN) into the denoising model. The main difference between the graph convolution neural network and the convolution neural network is that the graph convolution neural network can learn features without grid structure. No matter how the nodes in the graph convolution neural network change, the information exchange between the connected nodes is fixed, that is, the permutation invariance, which enables the denoising model to process both grid structure features and relatively scattered data features with spatial structure. This enables the graph diffusion model to learn the spatial structure of the posture skeleton graph, that is, the distribution of key points and the way bones are connected, so as to generate a stable posture skeleton graph. Specifically, the graph convolutional neural network is inserted between the downsampling component and the upsampling component of the denoising model.

[0052] In addition, since the stable diffusion model was originally designed to generate RGB images, the input and output of the denoising model are 3-channel or 4-channel features. However, in this application, it is necessary to process the posture skeleton graph. Since the k key points in the posture skeleton graph are usually represented as k heat maps, this application stacks the k heat maps to obtain a k-channel feature as the input of the stable diffusion model. At the same time, the output of the stable diffusion model is also a k-channel feature.

[0053] In practical applications, the input and output layers of the stable diffusion model are structurally improved so that it can smoothly predict multi-channel feature maps, not just the original 3-channel RGB image features. Taking a posture skeleton map containing 17 key points as an example, the improved stable diffusion model can output 17-channel features, with each channel corresponding to a key point.

[0054] It should be noted that the posture skeleton diagram in this application refers to a two-dimensional posture diagram of the human body, which is formed by connecting key points (such as head, arms, trunk, legs, etc.) and skeleton lines.

[0055] After inserting the denoising model in the stable diffusion model into the graph convolutional neural network, the stable diffusion model can be trained to obtain the graph diffusion model. Figure 2 , Figure 2 A partial flow chart of a posture controllable generation method provided in an embodiment of the present application. Figure 2 As shown, an embodiment of the present application provides a posture-controllable generation method, wherein the process of pre-training a stable diffusion model to obtain a graph diffusion model may include steps S201 to S204, and these steps are described in detail below.

[0056] S201, obtaining a second natural language text and a second posture skeleton graph, wherein the second posture skeleton graph matches the second natural language text.

[0057] In the embodiment of the present application, a natural language text (ie, a second natural language text) and a matching posture skeleton graph (ie, a second posture skeleton graph) are obtained as samples.

[0058] S202, converting the second posture skeleton image into a second key point heat map.

[0059] In the embodiment of the present application, the second posture skeleton graph is mapped to a corresponding key point heat map, namely, a second key point heat map.

[0060] S203, using the second natural language text and the second key point heat map to perform diffusion training on the stable diffusion model to adjust the model parameters of the denoising model inserted with the graph convolutional neural network.

[0061] In an embodiment of the present application, a second natural language text and a second key point heat map are used as samples to perform diffusion training on a stable diffusion model, and the model parameters of a denoising model with a graph convolutional neural network inserted are adjusted through a diffusion process and an inverse process. Among them, the diffusion process changes the data from an original state to a noisy state, and then the data in the noisy state is gradually restored to the state of the original data through an inverse process. Stable diffusion models are widely used in tasks such as image generation and text generation. In the present invention, this process is used to generate images of human postures from natural language descriptions.

[0062] In the stable diffusion model, the introduction of graph convolutional neural networks is used to capture the spatial relationship between key points in the posture skeleton graph, so as to better generate scientific skeleton postures. Specifically, the graph convolutional neural network can learn the adjacent relationship between key points (such as the left shoulder key point is connected to the left arm key point, the head key point is connected to the chest key point, etc.) and distance (such as the length of the forearm, the length of the thigh, etc.) according to the structural characteristics of the posture. This is crucial to the stability and accuracy of the generated posture skeleton structure.

[0063] See also Figure 3 , Figure 3 This is another partial flow chart of a posture controllable generation method provided in an embodiment of the present application. Figure 3 As shown, an embodiment of the present application provides a posture-controllable generation method, wherein step S203 "using a second natural language text and a second key point heat map to perform diffusion training on a stable diffusion model" may include steps S2031 to S2032, and these steps are described in detail below.

[0064] S2031, entering the forward noisy process, performing multiple noisy steps on the second key point heat map to obtain a second pure noise heat map, and recording the first noisy amount of each noisy step.

[0065] In the embodiment of the present application, the diffusion process is first entered, that is, the forward noise process. In this process, the second key point heat map is noised for multiple noise steps to obtain a corresponding pure noise heat map (that is, the second pure noise heat map). The noise of each noise step can be the same or different. At the same time, the noise amount of each noise step (that is, the first noise amount) is recorded.

[0066] S2032, enter the reverse denoising process, extract the text features of the second natural language text; based on the cross-attention mechanism, fuse the text features with the second pure noise heat map and input them into the stable diffusion model, and trigger the stable diffusion model to output the third key point heat map of this round of training by predicting the second noise amount of each noise step; calculate the loss function value of this round of training according to the first noise amount and the second noise amount; if the loss function value does not meet the corresponding end condition, adjust the network parameters of the denoising model with the graph convolutional neural network inserted according to the loss function value, and return to execute to trigger the stable diffusion model to output the third key point heat map of this round of training by predicting the second noise amount of each noise step; if the loss function value meets the corresponding end condition, end this round of training.

[0067] In the embodiment of the present application, after completing the forward denoising process, the inverse process, i.e., the reverse denoising process, is entered. In this process, the second natural language text is first encoded to obtain the corresponding text features, and then the second key point heat map is used as the posture feature. The text features are fused with the second pure noise heat map based on the cross-attention mechanism and then input into the stable diffusion model, and the stable diffusion model is triggered to perform multiple rounds of training to adjust the network parameters of the denoising model with the graph convolutional neural network inserted multiple times. The following is an illustration of one round of training of the stable diffusion model:

[0068] During this round of training, the stable diffusion model samples Gaussian noise multiple times to predict the amount of noise added for each noise step (i.e., the second amount of noise added), and outputs the key point heat map of this round of training (i.e., the third key point heat map). The third key point heat map is a multi-channel feature map, and each channel corresponds to a key point position in the posture skeleton; the mean square error value of the first amount of noise added and the second amount of noise added is used as the loss function value of this round of training; if the loss function value does not meet the corresponding end condition, the network parameters of the denoising model with the graph convolutional neural network inserted are adjusted according to the loss function value, and the next round of training is entered; if the loss function meets the corresponding end condition, the current round of training is ended, that is, the diffusion training of the stable diffusion model is ended.

[0069] It should be noted that text features can be obtained by encoding the second natural language text by any natural language encoder, such as CLIP (Contrastive Language-Image Pre-Training), BERT (Bidirectional Encoder Representations from Transformers), etc.

[0070] It should also be noted that in order to ensure that the generated skeleton posture conforms to the semantic information described in the input natural language, this application introduces a cross-modal cross-attention mechanism so that the text features can effectively guide the key point positioning in the generation process, and can combine the text features with the posture features to achieve accurate control of the posture skeleton by the natural language model.

[0071] S204, using the stable diffusion model after the diffusion training as the graph diffusion model.

[0072] In an embodiment of the present application, after the diffusion training is completed, the trained stable diffusion model is used as a graph diffusion model. The graph diffusion model introduces spatial topological information between multiple channels based on predicting multi-channel features from natural language descriptions. Experiments have shown that when predicting multi-channel features, the graph diffusion model can integrate the topological information of different channels into the prediction process through a graph convolutional network, so that the feature prediction of each channel not only depends on single-point features, but also combines the position information of adjacent channels, thereby effectively avoiding the problems of skeleton proportion imbalance or posture deformity in common methods, and greatly improving the accuracy and stability of the generated posture.

[0073] Based on this, after the stable diffusion model is trained to obtain the graph diffusion model, the first natural language text and the first pure noise heat map can be input into the graph diffusion model. The graph diffusion model performs model inference, extracts the text features of the first natural speech text, and samples the Gaussian noise multiple times based on the text features to predict the noise amount of each noise adding step, so as to denoise the first pure noise heat map and obtain the predicted multi-channel features, namely the subsequent first key point heat map.

[0074] S103, obtaining a first key point heat map output by the graph diffusion model, where the first key point heat map matches the first natural language text.

[0075] In an embodiment of the present application, a first key point heat map output by a graph diffusion model is obtained. The first key point heat map is generated by the graph diffusion model according to a first natural language text, thereby increasing additional posture information with diversity and correctness that can be used by the posture controllable generation task.

[0076] S104, converting the first key point heat map into a first posture skeleton map.

[0077] In the embodiment of the present application, the first key point heat map is mapped to two-dimensional coordinate points according to rules, and the two-dimensional coordinate points are connected to obtain the corresponding posture skeleton map, that is, the first posture skeleton map.

[0078] The pose skeleton graph will be used as a reference for additional pose information in the pose controllable generation task, thereby enriching the additional pose information that can be utilized in the task, while reducing the cost of manually collecting additional pose samples and dragging the correction skeleton graph.

[0079] On this basis, a gesture controllable generation method provided in an embodiment of the present application further includes the following steps:

[0080] A corresponding actual image is generated according to the first posture skeleton graph.

[0081] See also Figure 4 , Figure 4 This is an example of a posture skeleton diagram provided in an embodiment of the present application. Figure 5 , Figure 5 This is an example of an actual image provided in the embodiment of the present application. The first natural language text is "a girl sitting and drinking coffee", and the corresponding Figure 4 In the posture skeleton diagram shown, each key point is located in the correct part of the human body, and the overall posture is consistent with the natural language description, which reflects the rationality of the skeleton structure. At the same time, the posture skeleton diagram can be directly used in the posture controllable generation task as an additional posture reference and generate Figure 5 Actual image of the high profile compliance shown.

[0082] Therefore, the graph diffusion model in this application realizes efficient interpretation of natural language descriptions and generation of posture skeletons that conform to the human body structure in the posture skeleton generation task, providing new ideas and higher quality and more diverse reference skeletons for the posture controllable generation task.

[0083] Based on the above description, the posture controllable generation method provided in the embodiment of the present application has the following advantages:

[0084] 1) High generation accuracy. This application uses graph convolutional neural networks to model the spatial position information between key points of human posture, so that the generated posture maintains accuracy at each key point position, avoiding posture deformity or inconsistent skeleton proportions.

[0085] 2) High generation diversity. During the training process, the graph diffusion model gradually adds random noise to the posture data and generates it through reverse denoising. This gradual denoising mechanism introduces a certain degree of randomness in each step, allowing the model to generate different samples in different sampling processes. Unlike the common search mechanism of existing methods that matches in a known sample library, the graph diffusion model can search in a continuous latent space. The above mechanism enables the graph diffusion model to generate diverse samples, even including postures that have not appeared in the training data, expanding the diversity of the generated results.

[0086] 3) Low labor cost. This application significantly reduces the need for manual intervention and realizes the automatic generation of posture skeletons that meet the description based on natural language descriptions, without the need to manually drag key points for adjustment, making the posture generation process more efficient and more automated when used for posture controllable generation tasks.

[0087] 4) High adaptability to downstream tasks. When used for posture controllable generation tasks, this application can directly generate a two-dimensional posture skeleton from a natural language description without the need for manual conversion from other forms, which improves the adaptability to downstream tasks and is more convenient and practical.

[0088] In order to verify the effect of this application, a comparative analysis is performed on the implementation scheme of this application (i.e., GUNet in Table 1), the implementation scheme of removing the graph convolutional neural network from GUNet (i.e., GUNet-T2H in Table 1), and the two schemes in the paper "Adversarial Synthesis of HumanPose From Text" (i.e., WGAN-LP and WGAN-LP R in Table 1). Among them, WGAN-LP is implemented based on heat map prediction, and WGAN-LP R is implemented based on coordinate regression.

[0089] Table 1

[0090]

[0091] Specifically, the correctness (full name Mean-Square Error, abbreviated in English MSE) represents the difference between the posture skeleton generated according to the natural language description and the real posture skeleton corresponding to the natural language description. The smaller the correctness, the higher the correctness. It can be seen from Table 1 that GUNet has the highest correctness, which is significantly lower than GUNet-T2H. This shows that due to the introduction of graph convolutional neural networks, the posture skeleton generated by the graph diffusion model from the natural language description is more stable, that is, close to the real posture, and the key point distribution is more stable, which means that the predicted key point position will not deviate much from the real human body, for example, the head key point will not appear in the hands or feet.

[0092] See also Figure 6 , Figure 6 An example of a pose skeleton generated for WGAN-LP R. It can be seen that the key points have different degrees of errors, such as the right arm is too long in the first pose skeleton on the left, the head is too far away from the body in the second pose skeleton in the middle, and the foot key point is generated on the head in the third pose skeleton on the right. These can be understood as insufficient learning of the distribution position and distance of the key points of the body. This application uses a graph convolutional neural network to model the distribution and connection relationship of the key points of the human body as a graph, and then lets the model learn the structure of the human body. This method greatly reduces the probability of the model making the above errors.

[0093] Variety (Variance) refers to the number of posture skeletons predicted based on natural language descriptions. Within a certain range, the larger the value, the higher the diversity. This shows that the posture skeletons generated by the model are not single. From Table 1, we can see that the diversity of GUNet and GUNet-T2H is much higher than that of WGAN-LP and WGAN-LP R, which shows that the posture skeletons generated based on the graph diffusion model are much better than WGAN-LP and WGAN-LP R in diversity.

[0094] See also Figure 7 , Figure 7 Another example of a pose skeleton graph generated by WGAN-LP R. It can be seen that the key points fluctuate within a small range, and the poses presented by the three pose skeleton graphs are highly consistent.

[0095] Experiments show that the present application can generate diverse and correct posture skeleton graphs from natural language descriptions, increasing the amount of additional posture data available for posture controllable generation tasks.

[0096] A posture-controllable generation method provided in an embodiment of the present application is introduced above, and a device for executing the posture-controllable generation method is introduced below.

[0097] See also Figure 8 , Figure 8 This is a schematic diagram of the structure of a gesture controllable generation device provided in an embodiment of the present application. Figure 8 As shown, a gesture controllable generation device provided in an embodiment of the present application includes:

[0098] A model training module 10 is used to pre-train a stable diffusion model to obtain a graph diffusion model, in which a denoising model is inserted into a graph convolutional neural network;

[0099] The posture inference module 20 is used to obtain the first natural language text to be processed and randomly generate a first pure noise heat map; input the first natural language text and the first pure noise heat map into the graph diffusion model; obtain the first key point heat map output by the graph diffusion model, and match the first key point heat map with the first natural language text; convert the first key point heat map into a first posture skeleton map.

[0100] In a possible implementation, the posture reasoning module 20 is further used to:

[0101] A corresponding actual image is generated according to the first posture skeleton graph.

[0102] In a possible implementation, the model training module 10 is specifically used to:

[0103] A second natural language text and a second posture skeleton graph are obtained, and the second posture skeleton graph is matched with the second natural language text; the second posture skeleton graph is converted into a second key point heat map; the second natural language text and the second key point heat map are used to perform diffusion training on the stable diffusion model to adjust the model parameters of the denoising model inserted with the graph convolutional neural network; the stable diffusion model after the diffusion training is completed is used as the graph diffusion model.

[0104] In a possible implementation, the model training module 10 for performing diffusion training on the stable diffusion model using the second natural language text and the second key point heat map is specifically used to:

[0105] Enter the forward noise adding process, add noise to the second key point heat map for multiple noise adding steps to obtain a second pure noise heat map, and record the first noise adding amount of each noise adding step;

[0106] Enter the reverse denoising process to extract the text features of the second natural language text; based on the cross-attention mechanism, fuse the text features with the second pure noise heat map and input them into the stable diffusion model, and trigger the stable diffusion model to output the third key point heat map of this round of training by predicting the second noise amount of each noise step; calculate the loss function value of this round of training according to the first noise amount and the second noise amount; if the loss function value does not meet the corresponding end condition, adjust the network parameters of the denoising model with the graph convolutional neural network inserted according to the loss function value, and return to execute to trigger the stable diffusion model to output the third key point heat map of this round of training by predicting the second noise amount of each noise step; if the loss function value meets the corresponding end condition, end this round of training.

[0107] In one possible implementation, a graph convolutional neural network is inserted between the downsampling component and the upsampling component of the denoising model.

[0108] It should be noted that the detailed functions of each module in the embodiment of the present application can be found in the corresponding public part of the above-mentioned posture controllable generation method embodiment, and will not be repeated here.

[0109] The present application also provides an electronic device in an embodiment. Fig. 9 , Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device in the embodiment of the present application may include but is not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Fig. 9 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0110] like Fig. 9As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 to a random access memory (RAM) 903. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0111] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a memory card, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Fig. 9 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.

[0112] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any of the gesture-controllable generation methods provided in the embodiments of the present application.

[0113] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the posture-controllable generation methods provided in the embodiment of the present application.

[0114] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0115] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0116] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0117] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

Claims

1. A posture controllable generation method, characterized in that: The posture controllable generation method comprises: Obtaining a first natural language text to be processed, and randomly generating a first pure noise heat map; Inputting the first natural language text and the first pure noise heat map into a graph diffusion model, wherein the graph diffusion model is obtained by pre-training a stable diffusion model, and a denoising model in the stable diffusion model is inserted into a graph convolutional neural network; the number of input and output channels of the graph diffusion model is consistent with the number of key points in the posture skeleton graph, and each channel corresponds to a key point; Obtaining a first key point heat map output by the graph diffusion model, where the first key point heat map matches the first natural language text; Converting the first key point heat map into a first posture skeleton map; The process of pre-training the stable diffusion model to obtain the graph diffusion model includes: Acquire a second natural language text and a second posture skeleton graph, wherein the second posture skeleton graph matches the second natural language text; Converting the second posture skeleton image into a second key point heat map; Performing diffusion training on the stable diffusion model using the second natural language text and the second key point heat map to adjust model parameters of the denoising model in which the graph convolutional neural network is inserted; The stable diffusion model after the diffusion training is completed is used as the graph diffusion model.

2. The attitude controllable generation method according to claim 1, characterized in that: The controllable posture generation method further comprises: A corresponding actual image is generated according to the first posture skeleton image.

3. The attitude controllable generation method according to claim 1, characterized in that: The using the second natural language text and the second key point heat map to perform diffusion training on the stable diffusion model includes: Entering the forward noisy process, performing multiple noise-adding steps on the second key point heat map to obtain a second pure noise heat map, and recording the first noise-adding amount of each noise-adding step; Enter the reverse denoising process to extract the text features of the second natural language text; based on the cross-attention mechanism, fuse the text features with the second pure noise heat map and input them into the stable diffusion model, and trigger the stable diffusion model to output the third key point heat map of this round of training by predicting the second noise amount of each noise step; calculate the loss function value of this round of training according to the first noise amount and the second noise amount; if the loss function value does not meet the corresponding end condition, adjust the network parameters of the denoising model inserted with the graph convolutional neural network according to the loss function value, and return to execute the triggering of the stable diffusion model to output the third key point heat map of this round of training by predicting the second noise amount of each noise step; if the loss function value meets the corresponding end condition, end this round of training.

4. The attitude controllable generation method according to claim 1, characterized in that: The graph convolutional neural network is inserted between the downsampling component and the upsampling component of the denoising model.

5. A gesture controllable generation device, characterized in that: The attitude controllable generating device comprises: A model training module is used to pre-train a stable diffusion model to obtain a graph diffusion model, in which a denoising model is inserted into a graph convolutional neural network; the number of input and output channels of the graph diffusion model is consistent with the number of key points in the posture skeleton graph, and each channel corresponds to a key point; A posture inference module is used to obtain a first natural language text to be processed and randomly generate a first pure noise heat map; input the first natural language text and the first pure noise heat map into the graph diffusion model; obtain a first key point heat map output by the graph diffusion model, the first key point heat map matches the first natural language text; convert the first key point heat map into a first posture skeleton map; The model training module is specifically used for: Acquire a second natural language text and a second posture skeleton graph, wherein the second posture skeleton graph matches the second natural language text; convert the second posture skeleton graph into a second key point heat map; perform diffusion training on the stable diffusion model using the second natural language text and the second key point heat map to adjust the model parameters of the denoising model in which the graph convolutional neural network is inserted; and use the stable diffusion model after the diffusion training as the graph diffusion model.

6. A computer program product, characterized in that It comprises computer-readable instructions, and when the computer-readable instructions are executed on an electronic device, the electronic device implements the gesture-controllable generation method as claimed in any one of claims 1 to 4.

7. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the gesture-controllable generation method as described in any one of claims 1 to 4.

8. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the posture-controllable generation method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Figure generation method based on graph alignment large language model

    CN118823159A

  • Attitude estimation method and system based on deep learning

    CN119006598A