General sketch three-dimensional human body posture prediction method for stylized input

By constructing a large-scale sketch and 3D human pose dataset, combining a variational pose encoder and a graphic-text pre-training model, designing a U-shaped network and a hybrid converter, and optimizing sketch feature extraction and pose prediction, we solve the data scarcity and generalization problems in sketch 3D human pose estimation, and achieve efficient and robust pose prediction.

CN120599692APending Publication Date: 2025-09-05NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510639238.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing technologies for 3D human pose estimation from sketches suffer from a lack of large-scale data and insufficient generalization capabilities, resulting in low inference efficiency and poor input diversity.

Method used

By constructing a large-scale sketch and 3D human pose dataset, using a variational pose encoder, a graphic pre-training model and a controllable generative network, combining target detection and pose estimation models, designing a U-shaped network and a hybrid converter, optimizing sketch feature extraction and pose prediction, and adopting an end-to-end learning strategy for 3D human pose prediction.

Benefits of technology

The generalization ability of sketch 3D human pose estimation has been significantly improved, with the inference speed increased by 500 times. The model is more robust under different sketch styles and can accurately predict human body proportions and perspective errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599692A_ABST
    Figure CN120599692A_ABST
Patent Text Reader

Abstract

The invention provides a stylized input-oriented general sketch three-dimensional human body posture prediction method, which comprises the following steps of: 1, sampling SMPL (skin multi-person linear model) posture parameters through a variational posture encoder, and injecting text description and the posture parameters as conditions into a training controllable generation network; 2, training a target detection model network and a visual converter attitude estimation model network; step 3, enabling the sketch human body two-dimensional joint point position to pass through a multi-layer perceptron network to obtain a human body two-dimensional joint point position high-dimensional vector; 4, combining the cross attention map and the feature map in the output layer of the sketch human body posture prediction network to generate a posture feature vector of the sketch; and step 5, outputting three-dimensional human body shape parameters and camera parameters. The method can be widely applied to the fields of game creation, animation film production and the like, and has relatively high practical value and development prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and in particular relates to a general grass Figure 3 dimensional human posture prediction method. Background Art

[0002] Human pose estimation is a cross-modal synthesis technology involving computer vision and graphics, widely used in fields such as computer animation and film production, human-computer interaction, and virtual reality. Among the many data sources for pose estimation, sketches, as a highly practical and versatile form of representation, offer unique advantages. They are widely used in animation and film production due to their ease of access and ability to accurately convey the artist's vision. More broadly, the term "sketch" encompasses a wide range of drawing styles, including charcoal sketches, cartoons, stick figure drawings, children's doodles, oil paintings, and ink wash paintings.

[0003] Sketch-based three-dimensional (3D) human pose estimation can effectively improve the efficiency of animated character generation, optimize animation production pipelines, and provide a more flexible method for acquiring human poses for digital content creation. Therefore, it has widespread application in computer animation and film production. Furthermore, compared to traditional pose estimation methods based on photos or videos, sketches offer greater freedom for creative design and artistic expression, allowing animators and artists to more intuitively manipulate character poses and achieve more creative expressions. However, compared to human pose estimation from real photos, estimating human pose from sketches is a significantly more challenging task. General pose estimation methods perform poorly on this task because their training data is almost entirely derived from real photos, and the network models and supervision methods they use are tailored to people in real photos. In contrast, sketches lack clear character features, conveying less information, and often ignore the normal proportions and geometric perspective of the human body, expressing 3D human pose in a more abstract manner. This increases the complexity of predicting 3D human pose from sketches.

[0004] In recent years, in order to solve this problem, Gesture3D (Mikhail Bessmeltsev, Nicholas Vining, and Alla Sheffer. Gesture3d: Posing 3d characters via gesture drawing. ACM Trans. Graph., 35 (6), 2016.2) reconstructs poses through vector drawing, but it assumes that the input data has minimal noise, precise connections, and no extra strokes. However, these requirements are difficult to meet in sketches drawn freely by users; Brodt et al. proposed Sketch2Pose (Kirill Brodt and Mikhail Bessmeltsev. Sketch2pose: estimating a 3d character pose from a bitmap sketch. ACM Transactions on Graphics (TOG), 41 (4): 1–15, 2022. 2, 3, 7, 8), which first predicts 2D joint positions from the sketch and then aligns the 3D parametric human model to these skeletons through an optimization framework. However, this method runs slowly and is mainly optimized for hand-drawn sketch lines. Therefore, how to find a fast and highly generalized sketch-to-pose estimation solution remains an open problem to be solved.

[0005] In summary, due to the lack of large-scale sketch and 3D human pose paired data, previous methods perform poorly in terms of both inference efficiency and input diversity. Summary of the Invention

[0006] Purpose of the invention: The technical problem to be solved by the present invention is to provide a general grass Figure 3 The method for predicting human posture comprises the following steps:

[0007] Step 1: The variational pose encoder (VPoser) samples the pose parameters of the skinned multi-character linear model (SMPL). A random bone shortening strategy is introduced to adapt the sketch pose features. The pre-trained image-text model BLIP2 is then used to randomly generate text descriptions. The text descriptions and pose parameters are then used as conditions to train a controllable generative network (ControlNet). This method constructs the first large-scale sketch dataset with true 3D pose values ​​through large-scale model generation.

[0008] Step 2: Integrate and unify the sketch and 2D human pose datasets, train the object detection model YOLO-X network to detect the position of the person in the sketch, and train the visual converter pose estimation model ViTPose network to predict the 2D joint position of the human body;

[0009] Step 3: Generate a corresponding joint heat map based on the predicted 2D joint positions of the human body, and combine it with the latent vector encoded by the input sketch as the input of the U-Net. The sketch 2D joint positions of the human body are passed through the Multi-Layer Perceptron (MLP) network to obtain a high-dimensional vector of the 2D joint positions of the human body, which is used as the conditional input of the backbone network.

[0010] Step 4: Use the prior knowledge of the U-Net network in the pre-trained diffusion model Diffusion and copy the network coding layer to inject the 2D pose condition to obtain the sketch human pose prediction network; after training with the generated sketch and 3D human pose dataset, combine the cross attention map and feature map in the output layer of the sketch human pose prediction network to generate the sketch pose feature vector;

[0011] Step 5: Upgrade the sketch’s posture feature vector to three-dimensional space to obtain the sketch’s posture feature in three-dimensional space, align it with the two-dimensional joint position using the hybrid transformer Fusion Transformer network, and finally use the decoder output of the vector quantized variational autoencoder VQVAE to obtain the final sketch. Figure 3 3D human posture parameters, and output 3D human shape parameters and camera parameters through a linear layer.

[0012] In step 1, a variational pose encoder (VPoser) is used to generate random 3D human poses, random bone length deviations are introduced to simulate the random bone proportion features of the sketch, and then projected onto a 2D plane to obtain 3D and 2D joint annotations; then, the sketch and 2D human pose datasets are integrated, and sixteen 2D joint points are used to unify the 2D key point format. The sketch styles are manually divided into six categories (including cartoons, oil paintings, ink paintings, charcoal sketches, stick figure drawings, and children's graffiti), and the image-text pre-training model BLIP2 is used to generate sketch appearance descriptions and human pose description texts, and a controllable generative network ControlNet model is trained; then, edge detection and bounding box screening are used to generate sketch data, and the skinned multi-character linear model (SMPL) mesh occlusion analysis is used to record joint visibility; this method is the first to construct the first large-scale sketch dataset with 3D pose truth values ​​through the idea of ​​large model generation.

[0013] In step 2, the ViTPose network designed based on the visual converter Vison Transformer predicts the positions of the two-dimensional joints of the human body in the input sketch, and further extracts the sketch posture features by using the Gaussian kernel to represent the heat map.

[0014] In step 3, the sketch human feature extraction framework is designed, which is generally expressed as:

[0015]

[0016] where p φ (y|x) represents the process of the sketch human feature extraction framework, x represents the input sketch, y represents the output sketch human mesh parameters, represents the process of extracting 2D human pose conditions from sketches, ∈(x) represents the latent vector encoded by the input sketch, represents the pose condition of the input network (i.e., the 2D joint feature vector extracted from the sketch), Indicates from grass Figure 2 The process of extracting human posture features based on dimensional human posture conditions, represents the output features of the network, Represents the process of regressing human body mesh parameters from network output features.

[0017] In step 3, the input sketch is encoded into the latent space using the encoder of the Vector Quantized Variational Autoencoder (VQVAE). The heat map generated by the two-dimensional joint positions predicted in step 2 is combined with the latent space encoding to obtain unknown guidance in the sketch space. This method can greatly improve the spatial information of the sketch input:

[0018]

[0019] in is the U-Net hidden space input, z0 is the hidden vector encoded by the input sketch, It is the corresponding joint point heat map generated by the two-dimensional joint point position of the human body; Concat represents the connection operation;

[0020] For the conditional injection of the U-Net, the sketch of the two-dimensional joint points of the human body is used instead of the text condition. The two-dimensional joint points of the human body are upgraded to 768 dimensions as the condition vector through the multi-layer perceptron MLP network:

[0021]

[0022] in For the character position condition injection of U-Net, J 2D The 2D joint positions of the human body after encoding the input sketch.

[0023] In step 4, the two-dimensional joint position of the human body is input by copying the U-Net encoding layer, and then the obtained features are input into the U-Net decoding layer through a zero convolution layer to obtain a sketch human posture prediction network that can input the two-dimensional joint position condition information of the human body;

[0024] Extract the cross attention map and feature map in different sampling layers in the U-Net decoding layer of the U-Net, and further extract the output sketch features and the features of the output layer for:

[0025]

[0026] Where θ is the U-Net parameter, θ c is the encoding layer parameter of the replicated U-Net network, θ z is the zero convolution layer parameter, F n (;θ) is the U-Net, z is the input sketch feature of the U-Net, F n (;θ c ) is a copy encoding layer neural network that combines the sketch joint condition input, Z(;θ z ) represents the designed zero convolution layer.

[0027] In step 4, a new Figure 3 This paper presents a training method for dimensional human pose estimation. The supervision of the sketch human pose prediction network uses the designed sketch skeleton parallelism, perspective foreshortening, and the skinned multi-character linear model SMPL pose parameters. The new training objective makes the training process more consistent with the input characteristics of the sketch, thus surpassing all previous methods in performance. Skeletal parallelism is represented by the normal vector of the true skeleton direction vector and the predicted skeleton direction vector:

[0028]

[0029] in represents the skeleton parallelism loss of the sketch, is the i-th bone vector of the real sketch in 2D space, is the length of the i-th bone vector of the real sketch in 2D space, and n is the normal vector of the predicted bone vector;

[0030] The prediction is expressed as the ratio between the actual data and the predicted data:

[0031]

[0032] in represents the perspective foreshortening loss of the sketch, is the length of the i-th bone vector of the real sketch in 3D space, To predict the length of the i-th bone vector of the sketch in 3D space, To predict the length of the i-th bone vector of the sketch in 2D space;

[0033] The skinned multi-character linear model SMPL pose parameters directly perform L1 loss on the predicted results and the real results. Combined with the skeleton parallelism and perspective shortening of the sketch, this is the training goal of the sketch human pose prediction network designed by this method.

[0034] In step 5, the sketch pose features output by the sketch human pose prediction network and the sketch pose features upgraded to three-dimensional space are aligned through the hybrid transformer Fusion Transformer network to generate new three-dimensional pose features; then the decoder of the vector quantized variational autoencoder VQVAE is trained on the motion capture dataset of real images to better extract the output sketch features.

[0035] In step 5, the decoder outputs the SMPL pose parameters of the skinned multi-character linear model, and then the linear layer outputs the SMPL shape parameters and camera translation parameters of the skinned multi-character linear model. After that, the pose parameters, shape parameters and camera translation parameters generated here can be used to reconstruct a sketch-accurate three-dimensional pose.

[0036] The present invention also provides an electronic device, comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the method.

[0037] The present invention also provides a storage medium storing a computer program or instruction, which executes the steps of the method when the computer program or instruction is run on a computer.

[0038] Beneficial effects: This paper proposes an end-to-end sketching algorithm based on learning strategy by generating large-scale, multi-style sketches and 3D human pose datasets. Figure 3 This paper presents a novel 3D human pose prediction method, significantly improving the generalization of pose estimation algorithms across different sketch styles. The paper also develops an efficient feed-forward sketch human pose prediction network, achieving inference speed 500 times faster than the current state-of-the-art method. The paper also meticulously designs the network architecture and loss function, significantly enhancing the model's robustness, enabling accurate pose prediction despite challenges commonly encountered in sketches, such as human proportion distortion and perspective errors. The proposed method has broad application in areas such as game creation and animated film production, demonstrating its high practical value and promising future. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become more apparent.

[0040] Figure 1 Flowchart of the method of the present invention.

[0041] Figure 2 This is a demonstration of creating a data set in an embodiment of the present invention.

[0042] Figure 3 A schematic diagram of creating a sketch and a 3D human posture dataset in an embodiment of the present invention.

[0043] Figure 4 This is a diagram showing the results in an embodiment of the present invention. DETAILED DESCRIPTION

[0044] like Figure 1 As shown, this embodiment provides a general grass Figure 3 dimensional human posture prediction method, the specific process includes:

[0045] Step 1: Propose a new sketch and 3D human pose dataset construction process and the first large-scale sketch dataset SKEP-120k with 3D human pose ground truth to address the problem of scarcity of existing sketch and 3D pose estimation data. The dataset contains 120,000 sketch and 3D human pose pairs, covering six styles: cartoon, oil painting, ink painting, charcoal sketch, stick figure drawing and children's graffiti, with approximately 20,000 images for each style. The dataset construction process includes: (1) using the variational pose encoder VPoser to generate the pose parameters of the diversified skinned multi-character linear model SMPL, and randomly adjust the bone ratio to simulate the perspective deformation characteristics of the sketch; (2) using the image-text pre-training model BLIP2 to generate image description text, and combining it with the controllable generative network ControlNet to train the text-guided image generation model; (3) extracting the sketch contour through Canny detection, and optimizing the human bounding box and joint occlusion annotation;

[0046] Step 2: Integrate the existing sketches and 2D human pose datasets into a unified format of 16 human key points; train the object detection model YOLO-X network to detect the position of the people in the sketches, and train the visual transformer pose estimation model ViTPose network to predict the precise 2D joint positions of the human body;

[0047] Step 3: The accurate 2D joint locations of the human body sketch are obtained through step 2, which in turn generates a Gaussian heat map of the sketch joints. The input sketch is encoded into a latent space using a vector quantized variational autoencoder to obtain the image's latent vector. The latent vector and heat map are then concatenated to provide spatial guidance for the sketch. Simultaneously, the 2D joint locations of the human body are mapped to 768 dimensions using a multi-layer perceptron (MLP) network, replacing the traditional text-based conditional input.

[0048] In step 4, sketch feature extraction primarily extracts high-level knowledge from a pretrained denoising U-Net network. Using a pretrained diffusion model as the image backbone, the sketch's positional features are extracted using a single inference. The trainable encoding layer of the U-Net is then replicated to create a designed sketch pose prediction network. The textual conditions are replaced with sketch pose conditions to enhance the correlation between sketch and pose information. Finally, a method is proposed to extract hierarchical features from the cross-attention map and sketch feature map in the output layer of the sketch pose prediction network for further optimization.

[0049] In step 5, a human mesh regressor is proposed to predict the 3D human pose from the sketch features extracted in step 4. First, the 2D guide features are upgraded to 3D features. The hybrid transformer Fusion Transformer is used to regress the parameters of the skinned multi-character linear model SMPL to align and integrate the 2D and 3D features. The human pose parameters are then generated using the trained vector quantized variational autoencoder.

[0050] In step 1, a dataset of sketches and 3D human poses containing 120,000 data pairs covering a variety of sketch styles is proposed, named SKEP-120k. Figure 2 As shown on the left, the SKEP-120k dataset has 16 joint annotations for the human skeleton in each sketch (0 to 15 in the figure represent the human head joints, neck joints, shoulder joints, elbow joints, hand joints, hip joints, knee joints, ankle joints and foot joints respectively); Figure 2 As shown on the right, the dataset covers six sketch styles, including cartoons, oil paintings, ink paintings, charcoal sketches, stick figure drawings, and children's graffiti, with each style containing approximately 20,000 images. The dataset of an embodiment of the present invention provides a human body bounding box, 16 key points of the human body (two-dimensional / three-dimensional coordinates and their visibility annotations), skinned multi-character linear model SMPL posture parameters, and sketch text information. Due to the differences in the definition of the human skeleton in different datasets, and taking into account the gesture expression characteristics of the sketch, the present invention redefines a new three-dimensional skeleton to represent the human body posture. Specifically, the body parts refer to the MSCOCO database, and additional left and right toe joints are added to describe the leg posture more finely.

[0051] The process of creating a dataset is as follows Figure 3 As shown, first, the variational pose encoder VPoser is used to generate random three-dimensional human poses, and random bone length deviations are introduced to address the common perspective shortening phenomenon in sketches, and then projected onto a two-dimensional plane to obtain three-dimensional and two-dimensional joint annotations. Next, existing sketch and two-dimensional human pose datasets such as Sketch2Pose, HumanArt, and Amateur Drawing are integrated, and the 2D key point format and style classification are unified. In order to enhance data richness, the present invention uses the graphic pre-training model BLIP2 to generate sketch appearance descriptions and human pose description texts, and trains a text-conditional image generation model based on a controllable generation network ControlNet network. Subsequently, Canny edge detection and bounding box screening are combined to generate high-quality, multi-style sketch data, and the skinned multi-character linear model SMPL human mesh occlusion analysis is used to record joint visibility. Finally, SKEP-120k provides a comprehensive and reliable method for sketching through large-scale data synthesis, style diversification, and high-precision annotation. Figure 3 D pose estimation provides a high-quality training benchmark.

[0052] Next, we design a sketch human feature extraction framework, which is represented as follows:

[0053]

[0054] Where x represents the input sketch, y represents the output sketch body mesh parameters, represents the process of extracting 2D human pose conditions from sketches, ∈(x) represents the latent vector encoded by the input sketch, Denotes the posture condition of the input network, Indicates from grass Figure 2 The process of extracting human posture features based on dimensional human posture conditions, represents the output features of the network, Represents the process of regressing human body mesh parameters from network output features.

[0055] In step 2, the existing sketch and 2D human pose dataset are used to train the object detection model YOLO-X network to identify the human body in the sketch and generate the bounding box. The human features are then extracted through the non-hierarchical Vision Transformer in the visual transformer pose estimation model ViTPose, and the 2D joint points are predicted through a lightweight decoder.

[0056] In step 3, the input sketch is first encoded into a latent vector using a vector quantized variational autoencoder (VQVAE). This latent vector is then concatenated with a Gaussian heat map generated from two-dimensional joints to provide spatial guidance for the network input. The traditional text embedding is replaced by the two-dimensional joint positions of the human body. This is then mapped to 768 dimensions using a two-layer MLP to generate the network control condition. The network input and control condition are fed into different channels of the backbone network to guide the understanding of the human body structure in the sketch.

[0057] In step 4, a multi-scale feature extractor is designed based on the pre-trained denoising U-Net. This model fully utilizes the high-level knowledge in the pre-trained diffusion model to extract informative features and uses the learned knowledge to predict the 3D human pose in the sketch. The trainable encoding layer of the U-Net is copied to input the 2D joint control conditions, and the output of the decoding layer is:

[0058]

[0059] Among them, F n (;θ) is the main network layer of the U-Net network, z is the input sketch feature of this layer, F n (;θ c ) is a trainable neural network that combines the sketch joint condition input, is the 2D joint feature vector extracted from the sketch, Z(;θ z ) represents the zero convolution layer in the network. During the training of the network, three main training objectives are designed: skeleton parallelism, perspective foreshortening and skinned multi-character linear model SMPL human mesh parameters. Unlike estimating human pose from real photos, recovering human pose from artificial sketches is more difficult, mainly due to its distorted proportions, perspective effects and foreshortening deformation. Specifically, sketches often depict unrealistic body shapes or exaggerated body proportions, so standard optimization methods that only rely on 2D joint positions may lead to inaccurate or unnatural results. By observing human paintings, three key factors are identified that affect the accuracy of pose recovery: skeleton tangents, perspective foreshortening and self-contact.

[0060] Due to inaccuracies in drawings or artistic manipulation, a character's skeleton often appears visually longer than its actual length. This discrepancy makes it impractical to rely directly on absolute joint positions for optimization. At the same time, artistic research generally emphasizes the importance of accurately describing joint angles. Therefore, to ensure that the reconstructed 3D joint angles have a reasonable projection in the 2D image, the projection of the 3D skeleton must be aligned with the skeleton in the 2D drawing. To achieve this, the principle of bone parallelism can be used, which is represented by the normal vector between the true bone direction vector and the predicted bone direction vector:

[0061]

[0062] in is the i-th bone vector of the real sketch in 2D space, is the length of the i-th bone vector of the real sketch in 2D space, and n is the normal vector of the predicted bone vector;

[0063] Artists typically don't rely on precise mathematical measurements for orthographic or perspective projection when drawing images. Therefore, reconstructing the 3D pose directly from the predicted 2D pose often results in inaccurate angles between the skeleton and the screen. The skeleton is foreshortened to compensate for the incorrect perspective proportions of the character in the sketch, using the ratio between the ground truth and the predicted data as the predicted representation:

[0064]

[0065] in is the length of the i-th bone vector of the real sketch in 2D space, is the length of the i-th bone vector of the real sketch in 3D space, To predict the length of the i-th bone vector of the sketch in 3D space, To predict the length of the i-th bone vector of the sketch in 3D space;

[0066] Self-contact refers to the contact between different parts of the human body. Human observers usually rely on the visual perception of self-contact when judging depth, and map the parts of the body that are in contact with each other to similar depths. Existing methods are mainly optimized based on manually annotated self-contact areas, and physical contact constraints between vertex pairs are achieved by mapping each contact area to roughly aligned Skinned Multi-Character Linear Model SMPL human mesh vertices. In contrast, the dataset of the present invention provides the precise Skinned Multi-Character Linear Model SMPL human pose parameters of the human body in the sketch, and replaces the traditional self-contact loss with the Skinned Multi-Character Linear Model SMPL pose parameter loss, so that the relative depth and joint position of the character skeleton can be accurately obtained.

[0067] Through the above steps, the present invention can perform efficient and accurate 3D human body posture prediction on sketch data of various styles, and generate results such as Figure 4 As shown in the results, it can be seen that the generated human body posture is consistent with the artist's intention and has good rationality.

[0068] The present invention provides a general grass Figure 3There are many methods and approaches to implement this technical solution. The above is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention. Such improvements and modifications should also be considered within the scope of protection of the present invention. Any components not specified in this embodiment may be implemented using existing technologies.

Claims

1. A general sketch 3D human pose prediction method for stylized input, characterized by: The following steps are involved: Step 1: The variational pose encoder (VPoser) is used to sample the pose parameters of the skinned multi-character linear model (SMPL). A random bone shortening strategy is introduced to adapt to the pose characteristics of the sketch. The image-text pre-training model BLIP2 is then used to randomly generate text descriptions. The text descriptions and pose parameters are then injected into the controllable generative network (ControlNet) as conditions for training. Step 2: Integrate and unify the sketch and 2D human pose datasets, train the object detection model YOLO-X network to detect the position of the person in the sketch, and train the visual converter pose estimation model ViTPose network to predict the 2D joint position of the human body; Step 3: Generate a corresponding joint heat map based on the predicted 2D joint positions of the human body, and combine it with the latent vector encoded by the input sketch as the input of the U-Net. The 2D joint positions of the sketch are passed through the multi-layer perceptron (MLP) network to obtain a high-dimensional vector of the 2D joint positions of the human body, which is used as the conditional input of the backbone network. Step 4: Use the prior knowledge of the U-Net network in the pre-trained diffusion model Diffusion and copy the network coding layer to inject the 2D pose condition to obtain the sketch human pose prediction network; after training with the generated sketch and 3D human pose dataset, combine the cross attention map and feature map in the output layer of the sketch human pose prediction network to generate the sketch pose feature vector; In step 5, the sketch's pose feature vector is upgraded to three-dimensional space to obtain the sketch's pose features in three-dimensional space. The hybrid transformer Fusion Transformer network is used to align it with the two-dimensional joint positions. Finally, the decoder output of the vector quantized variational autoencoder (VQVAE) is used to obtain the final sketch's three-dimensional human pose parameters, and the three-dimensional human shape parameters and camera parameters are output through the linear layer.

2. The method according to claim 1, characterized in that In step 1, a variational pose encoder (VPoser) is used to generate random 3D human poses, random bone length deviations are introduced to simulate the random bone proportion features of the sketch, and then projected onto a 2D plane to obtain 3D and 2D joint annotations. Next, the sketch and 2D human pose datasets are integrated, and sixteen 2D joint points are used to unify the 2D key point format. The sketch styles are classified, and the image-text pre-training model BLIP2 is used to generate sketch appearance descriptions and human pose description texts, and the controllable generative network ControlNet model is trained. Subsequently, edge detection and bounding box screening are used to generate sketch data, and the skinned multi-character linear model SMPL mesh occlusion analysis is used to record joint visibility.

3. The method according to claim 2, characterized in that In step 2, the ViTPose network designed based on the visual converter VisonTransformer predicts the positions of the two-dimensional joints of the human body in the input sketch, and further extracts the sketch posture features by using the Gaussian kernel to represent the heat map.

4. The method according to claim 3, characterized in that In step 3, the sketch human feature extraction framework is designed, which is generally expressed as: where p φ (y|x) represents the process of the sketch human feature extraction framework, x represents the input sketch, y represents the output sketch human mesh parameters, represents the process of extracting 2D human pose conditions from sketches, ∈(x) represents the latent vector encoded by the input sketch, represents the posture condition of the input network, It represents the process of extracting human body posture features from the sketch 2D human body posture conditions. represents the output features of the network, Represents the process of regressing human body mesh parameters from network output features.

5. The method according to claim 4, characterized in that In step 3, the input sketch is encoded into the latent space using the encoder of the vector quantized variational autoencoder (VQVAE). The heat map generated by the two-dimensional joint positions predicted in step 2 is combined with the latent space encoding to obtain the unknown guidance in the sketch space: in is the U-Net hidden space input, z0 is the hidden vector encoded by the input sketch, It is the corresponding joint point heat map generated by the two-dimensional joint point position of the human body; Concat represents the connection operation; For the conditional injection of the U-Net, the sketch of the two-dimensional joint points of the human body is used instead of the text condition. The two-dimensional joint points of the human body are upgraded to 768 dimensions as the condition vector through the multi-layer perceptron MLP network: in For the character position condition injection of U-Net, J 2D The 2D joint positions of the human body after encoding the input sketch.

6. The method according to claim 5, characterized in that In step 4, the two-dimensional joint position of the human body is input by copying the U-Net encoding layer, and then the obtained features are input into the U-Net decoding layer through a zero convolution layer to obtain a sketch human posture prediction network that can input the two-dimensional joint position condition information of the human body; Extract the cross attention map and feature map in different sampling layers in the U-Net decoding layer of the U-Net, and further extract the output sketch features and the features of the output layer for: Where θ is the U-Net parameter, θ c is the encoding layer parameter of the replicated U-Net network, θ z is the zero convolution layer parameter, F n (;θ) is the U-Net, z is the input sketch feature of the U-Net, F n (;θ c ) is a copy encoding layer neural network that combines the sketch joint condition input, Z(;θ z ) represents the designed zero convolution layer.

7. The method according to claim 6, characterized in that In step 4, the supervision of training the sketch human pose prediction network uses the designed sketch bone parallelism, perspective foreshortening and skinned multi-character linear model SMPL pose parameters, where bone parallelism is represented by the normal vector of the true bone direction vector and the predicted bone direction vector: in represents the skeleton parallelism loss of the sketch, is the i-th bone vector of the real sketch in 2D space, is the length of the i-th bone vector of the real sketch in 2D space, and n is the normal vector of the predicted bone vector; The prediction is expressed as the ratio between the actual data and the predicted data: in represents the perspective foreshortening loss of the sketch, is the length of the i-th bone vector of the real sketch in 3D space, To predict the length of the i-th bone vector of the sketch in 3D space, To predict the length of the i-th bone vector of the sketch in 2D space; The pose parameters of the skinned multi-character linear model SMPL are directly subjected to L1 loss between the predicted results and the actual results. Combined with the skeleton parallelism and perspective foreshortening of the sketch, this is the training goal of the sketch human pose prediction network.

8. The method according to claim 7, characterized in that In step 5, the sketch pose features output by the sketch human pose prediction network and the sketch pose features upgraded to three-dimensional space are aligned through the hybrid transformer FusionTransformer network to generate new three-dimensional pose features; Then, the decoder of the vector quantized variational autoencoder VQVAE is trained on a motion capture dataset of real images. The decoder outputs the skinned multi-character linear model SMPL pose parameters, and then the linear layer outputs the skinned multi-character linear model SMPL shape parameters and camera translation parameters. The generated pose parameters, shape parameters and camera translation parameters are used to reconstruct a sketch-accurate three-dimensional pose.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor is caused to perform the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is run on a computer, the steps of the method according to any one of claims 1 to 8 are executed.