An aesthetic perception gesture generation method based on visual and text prompts
Patent Information
- Application Number
- CN202610822271.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-06-09
AI Technical Summary
现有姿态生成研究多聚焦于人体与场景的功能交互适配,仅注重生成满足沙发就坐、运动动作等场景功能需求的姿态,完全忽略了姿态本身的美学属性,同时未结合人体个性化特征进行姿态生成,无法满足人像摄影的美学需求
本发明通过预构建姿态原型库,结合多模态特征提取与原型分类器完成目标姿态原型的精准匹配,不再采用直接全局生成姿态的方式,有效降低美学感知姿态生成的整体复杂度,实现了快速定位与场景、人物属性相适配的基础姿态,为姿态的细粒度优化筑牢优质起点,打造出高效生成美学感知姿态的核心思路,大幅提升姿态生成前期的匹配效率与场景适配性。
Smart Images

Figure CN122369124B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and human pose generation technology, specifically to an aesthetically perceptible pose generation method based on visual and textual prompts. Background Technology
[0002] With the widespread use of social media, outdoor travel and everyday scene recording have become common lifestyles, significantly increasing people's demand for aesthetic quality in portrait photography. Capturing aesthetically pleasing portrait photos in tourist attractions and various everyday settings has become a mainstream need. Among the many factors influencing the aesthetic quality of portrait photos, such as lighting conditions, perspective selection, and clothing coordination, human posture is a core key factor. The rationality and aesthetics of the human posture directly determine the shooting effect of the portrait photo. How to generate aesthetically pleasing human postures for specific users in designated scenes has become an open problem that urgently needs to be solved in this field.
[0003] Human pose generation in a scene falls under the research scope of usability-aware pose generation. The core of usability learning is the interaction possibilities provided by the environment. This direction has formed multi-dimensional research branches, covering contextual usability learning, functional understanding, usability classification, usability detection, usability segmentation, etc. Among them, contextual usability learning mainly uses usability relationships extracted from scene context as clues to improve the performance of related tasks. It is also an important research direction in the field of human pose generation. Related research mainly focuses on human pose generation combined with scene context to explore the interaction possibilities between the human body and the physical scene.
[0004] Human pose generation is a research hotspot in the field of computer vision. The development of this technology has gone through a stage from generating poses based on single conditions such as video time information and text descriptions in the early stage, to the rapid development stage of generating poses based on two-dimensional scene context in recent years. Related research mainly focuses on generating human poses that are compatible with scene spatial relationships and functional interactions. Some studies have also explored optimizations for issues such as scene occlusion, three-dimensional pose diversification, and text-driven pose expressiveness.
[0005] Meanwhile, research on image aesthetics has been developing in the field of computer vision for many years, forming several core research directions. Among them, the aesthetic evaluation direction evaluates the visual appeal of an image by identifying and quantifying its aesthetic features; the aesthetic cropping direction optimizes the image composition by selecting the region with the best visual appeal; and the aesthetic perspective selection direction captures the aesthetic features of a scene by identifying and optimizing camera angles or trajectories. All of the above research revolves around the aesthetic evaluation and optimization of the image itself, providing a technical foundation for computer vision tasks related to aesthetic perception.
[0006] In recent years, prototype-based generation methods have been gradually applied in the field of human pose generation. The Transformer model, due to its excellent performance in capturing contextual dependencies, has also become an important technical means in the field of computer vision, providing a new direction for the optimization of human pose generation technology.
[0007] Although technologies such as human pose generation, usability learning, and image aesthetic evaluation have made some progress, existing technologies still have many insurmountable shortcomings in the aesthetic perception pose generation scenario for scene-based portrait photography. Specifically, these shortcomings include: Existing research on pose generation mostly focuses on the functional interaction adaptation between the human body and the scene, only emphasizing the generation of poses that meet the functional requirements of scenes such as sitting on a sofa and moving, completely ignoring the aesthetic attributes of the pose itself. At the same time, it does not combine the personalized characteristics of the human body for pose generation, and thus cannot meet the aesthetic requirements of portrait photography.
[0008] Traditional pose generation methods all adopt a global pose generation strategy, generating key points of human pose as a whole. This makes it difficult to capture the fine-grained interaction relationships between key points, resulting in stiff and uncoordinated poses that do not conform to the laws of human kinematics.
[0009] Existing multimodal pose generation technologies lack the ability to effectively integrate visual and textual prompts. They only fuse visual and textual features through simple splicing, failing to achieve deep semantic alignment between the two modalities. They cannot fully utilize human attributes and scene features in the text description, and the generated poses often fail to meet both textual semantic requirements and scene spatial adaptation requirements. This often results in problems such as inconsistencies between the pose and the text description, and conflicts between the pose and the scene environment. The generated results have extremely low adaptability to the user's actual needs.
[0010] Currently, there is a lack of large-scale aesthetic pose datasets in this field to support the task of generating aesthetically perceived poses. Manually labeling aesthetic poses for scenes is difficult and produces ambiguous results. Existing dataset construction methods also suffer from limited scene coverage and high construction costs. Furthermore, the method of automatically expanding and constructing datasets through existing models is prone to generating sample noise, making it difficult to form high-quality datasets to support the task.
[0011] Existing pose generation models are not robust to data noise and have limited generalization ability in real-world scenarios. Among them, pose generation schemes based on generative adversarial networks suffer from pattern collapse, lack diversity in pose generation, and perform poorly in generating complex poses and unconventional actions. Pose generation schemes based on diffusion models, on the other hand, have extremely high requirements for the quantity and quality of training data, have weak generalization ability in small sample scenarios, and have low generation accuracy in occluded scenes, complex perspective scenes, and niche action poses.
[0012] Existing pose generation models suffer from slow convergence and unstable training states during training. They are unable to perform precise optimization learning for complex poses and difficult movements, and lack scientific progressive training strategies. The models are not good at learning pose detail features, and the final generated poses fail to meet the professional application standards of portrait photography in terms of both accuracy and aesthetics.
[0013] There is a significant gap in the application of emerging technologies in aesthetic perception pose generation scenarios. Existing prototype-based pose generation methods only focus on simple pose template matching without combining aesthetic constraints for prototype design and matching. The Transformer model has not yet been effectively applied to fine-grained interactive modeling of key points of human pose, and cannot fully leverage its technical advantages in capturing contextual dependencies.
[0014] In summary, existing technologies cannot generate aesthetically pleasing human poses at specified locations within a scene based on scene images and character descriptions. There is an urgent need in this field for a pose generation method that integrates multimodal cueing information, aesthetic constraints, prototype matching, and fine-grained interactive modeling of key points. This method would address core issues in existing technologies, such as the lack of aesthetic attributes in poses, poor multimodal information fusion, and weak model robustness and generalization ability, in order to meet the aesthetic perception pose generation requirements of scene-based portrait photography. Summary of the Invention
[0015] This invention provides an aesthetically perceptive pose generation method based on visual and textual prompts. It aims to achieve aesthetically perceptive human pose generation that conforms to portrait photography aesthetic standards and accurately matches the input scene and human attributes by constructing a pose prototype library, designing an efficient multimodal feature fusion strategy, building a partial-level Transformer variational autoencoder architecture, and combining it with a course learning strategy based on prototype complexity. Simultaneously, it improves the efficiency and accuracy of pose generation, enhances the model's robustness and scene adaptability, and meets the practical application needs of aesthetically perceptive pose generation in scene-based portrait photography and other related scenarios.
[0016] To solve the above-mentioned technical problems, the technical solution provided by the present invention is as follows: An aesthetically perceptual gesture generation method based on visual and textual cues includes the following steps: S1. Obtain multimodal input data, which includes images of scenes without people, specified locations of people in the scene, text descriptions of people's attributes, and text descriptions of scene features. S2. Multimodal feature extraction is performed on the multimodal input data to obtain the corresponding visual features, visual feature maps and text features that are homologous to the visual features; based on the visual features and text features, a suitable target pose prototype is obtained by matching through a pre-trained multimodal prototype classifier. The target pose prototype is selected from a pre-built pose prototype library, which is constructed by normalizing the true values of human poses in the training set used for model training and then using the K-means clustering algorithm. S3. Input the target pose prototype, visual feature map and text features into the pre-trained partial-level Transformer variational autoencoder, and output the predicted deformation parameters, including the scale factor and coordinate and visibility offset. The partial-level Transformer variational autoencoder uses a single human keypoint as the smallest processing unit, models the spatial adaptation relationship between human keypoints and scene environment through cross-attention mechanism, models the limb coordination dependency relationship between human keypoints through self-attention mechanism, and independently encodes the corresponding normal distribution for each human keypoint. S4. The predicted deformation parameters are applied inversely to the target pose prototype to restore the aesthetically perceived human pose that conforms to the aesthetic standards of portrait photography and matches the input scene and the attributes of the person. The human pose consists of 17 human key points, and each human key point contains coordinate information and visibility score information.
[0017] Furthermore, the method for constructing the pose prototype library specifically includes the following steps: S211. Use the true values of each human pose in the training set to train the model. Normalization is performed: the width of the bounding box of the pose keypoints is calculated. and height Divide the x-coordinates of all key points by y coordinate divided by The pose bounding boxes are uniformly adjusted to unit boxes to eliminate scale differences and obtain normalized poses. ; S212. Using the K-means clustering algorithm, with Euclidean distance as the metric and ignoring invisible keypoints, for all normalized poses... Perform clustering until convergence, then select... Cluster centers serve as pose prototypes A pose prototype library was constructed. This represents the total number of posture prototypes in the posture prototype library. Each posture prototype contains the same number of coordinate and visibility information of key human body points as the aesthetically perceived human posture, covering aesthetic posture types such as standing, sitting, half-squatting, and side-facing.
[0018] Furthermore, in step S2, multimodal feature extraction and prototype matching specifically include the following steps: S221. Visual Feature Extraction: Using a specified location of a person Centered on a single image, two square image patches of different sizes are cropped. These two patches, along with an original scene image scaled to the same preset size, are then input into three independent ResNet-50 backbone networks, resulting in three sets of backbone feature vectors. These three sets of backbone feature vectors are then concatenated and compressed into a fixed-dimensional array. visual features Simultaneously, the unpooled features from the output of the last convolutional layer of the ResNet-50 backbone network are extracted as visual feature maps. ; S222, Text Feature Extraction: Using a pre-trained CLIP-Text encoder with frozen parameters, text descriptions of character attributes are extracted. Scene feature text description The text features are concatenated and compressed into a fixed dimension. Text features ; S223, Prototype Classification and Matching: Classifying and Matching Visual Features Text features pass After the operation is concatenated, the input is a linear classifier, and the output is... The classification log probability of a pose prototype The calculation formula is: ; in, For the categorical log probability, For linear classifiers, As a visual feature, For text features, This is a vector concatenation operation. This represents the total number of pose prototypes in the pose prototype library. pass The operation selects the pose prototype with the highest probability as the target pose prototype. .
[0019] Furthermore, a partial-level Transformer variational autoencoder includes an encoder and a decoder. The encoder's processing specifically includes the following steps: S311. Initialize learnable keypoint embeddings. The number of embeddings is equal to the number of human keypoints. Each embedding is used to store the inherent attribute features of the corresponding human joint. S312. Calculate the true offset between the normalized pose and the target pose prototype in the training samples. Combined with the width and height scale factors in the attitude normalization process Construct a real deformation parameter matrix and extract deformation features; S313. Concatenate the keypoint embedding, deformation features, and repetition extensions to text features that match the number of keypoints, and obtain the query embedding through a compression function. The calculation formula is as follows: ; in, For query embedding, For the characteristic compression function, Embedding of learnable key points Deformation characteristics, Text feature vectors repeated for dimension alignment This is a vector concatenation operation; S314. Embed the query into the input cross-attention mechanism, interact with the visual feature map, model the spatial adaptation relationship between human key points and the scene environment, and output the cross-attention features. The calculation formula is as follows: ; in, For cross-attention features, This is the computation function corresponding to the cross-attention mechanism. For query embedding, Visual feature map; S315. Input the cross-attention features into the self-attention mechanism to model the limb coordination dependency relationship between human key points, and output the key point interaction features. The calculation formula is as follows: ; in, For key point interaction features, This is the computation function corresponding to the self-attention mechanism. Cross-attention features; S316. Input the keypoint interaction features into the fully connected layer, and output the mean and variance for each human keypoint to construct an independent normal distribution for each keypoint. The calculation formula is as follows: ; in, The mean of a normal distribution is . The variance of the normal distribution is... For the fully connected layer encoding function, For key point interaction features, This is a vector concatenation operation.
[0020] Furthermore, the decoder's processing procedure specifically includes the following steps: S321. From the normal distribution corresponding to each key point Sample latent variables separately To form a latent variable matrix ,Will Input fully connected layers extract latent features The calculation formula is: ; in, Indicates a normal distribution. For the first Each key point corresponds to the mean of a normal distribution. For the first Each key point corresponds to the variance of a normal distribution. For from the first Each key point corresponds to a latent variable sampled from a normal distribution. The latent variable matrix consists of the latent variables of all key points. It is a linear rectified activation function. For the weights of the fully connected layer, For bias terms of fully connected layers; S322, Embed key points Potential characteristics Repeated expansion to text features matching the number of key points Concatenate and compress to obtain the decoded query embedding ; S323, Embed the decoding query After being processed by the cross-attention computation unit and the self-attention computation unit in sequence, the decoded features are output. ; S324, Decode the features The input is a fully connected layer, and the predicted deformation parameters are obtained by decoding. The calculation formula is as follows: ; in, The predicted scaling factor matrix, For the predicted coordinates and visibility offset matrix, Decoding function for fully connected layer For the weights of the fully connected layer, For bias terms of fully connected layers, This indicates a matrix concatenation operation.
[0021] Furthermore, the total loss function used for pre-training the partial-level Transformer variational autoencoder is: ; in, The total loss function for model training. For KL divergence loss, To balance hyperparameters, For attitude loss; KL divergence loss The calculation formula is: ; in, An index of key points on the human body. For the first Each key point corresponds to the variance of a normal distribution. For the first Each key point corresponds to the mean of a normal distribution. For the first Each key point corresponds to the square of the L2 norm of the mean. For the first The logarithm of the determinant of the variance corresponding to each key point; Attitude loss The calculation formula is: ; in, An index of key points on the human body. For the first The true visibility score of each key point For the first Predicted coordinate offset of each key point For the first The true coordinate offset of each key point The cross-entropy loss function is used for binary classification. For the first Predictive visibility score for each key point The set of true visibility scores for all key points; Predictive visibility score of key points The calculation formula is: ; in, The predicted visibility score for key points. It is the Sigmoid activation function. It is the inverse of the Sigmoid function. The visibility score of the target pose prototype. This represents the predicted visibility offset.
[0022] Furthermore, the partial-level Transformer variational autoencoder is trained using a prototype complexity-based course learning strategy, specifically including the following steps: S331. Calculate the average Euclidean distance between each posture prototype in the posture prototype library and all other prototypes, and use it as a measure of the complexity of the corresponding prototype. S332. Sort all pose prototypes in ascending order of complexity metric, divide them into three groups, corresponding to three training courses from easy to difficult. S333. Train in stages according to the order of course difficulty from low to high. After completing the preset training rounds for each course, move on to the next course until all courses are completed.
[0023] Furthermore, in step S4, the predicted deformation parameters are applied inversely to the target posture prototype to reconstruct the aesthetically perceived human posture, including the following steps: S41. The attitude scale of the target attitude prototype is restored by predicting the scale factor in the deformation parameters. The calculation formula is as follows: ; Where S is the predicted scale factor matrix, and Δ is the predicted coordinate and visibility offset matrix. The x-axis original coordinates of the key points in the target pose prototype. The original y-axis coordinates of key points in the target pose prototype. These are the x-axis coordinates of the key points after scaling. These are the y-axis coordinates of the key points after scaling. The 0th column of the scale factor matrix corresponds to the x-axis scale coefficient. This is the first column of the scale factor matrix, corresponding to the y-axis scale coefficient; S42. Adjust the coordinate position of key points by predicting the coordinate offset in the deformation parameters. The calculation formula is as follows: ; in, The final output contains the x-axis coordinates of each key point. The final output contains the y-axis coordinates of each key point. The 0th column of the offset matrix corresponds to the x-axis coordinate offset. The first column of the offset matrix corresponds to the y-axis coordinate offset. S43, via offset Adjust the keypoint visibility score and filter keypoints with a predicted visibility score V < 0.2 to obtain the final aesthetically perceived human posture.
[0024] The technical effects of this invention are as follows: This invention achieves accurate matching of target pose prototypes by pre-constructing a pose prototype library and combining multimodal feature extraction with a prototype classifier. It no longer adopts the method of directly generating poses globally, effectively reducing the overall complexity of aesthetically perceptible pose generation. It realizes the rapid positioning of basic poses that are compatible with scene and character attributes, laying a solid foundation for fine-grained pose optimization, creating a core idea for efficiently generating aesthetically perceptible poses, and greatly improving the matching efficiency and scene adaptability in the early stage of pose generation.
[0025] This invention designs a proprietary multimodal feature extraction and deep fusion strategy to extract scene visual features and text description features separately and then complete feature splicing and fusion. It deeply explores the cross-modal dependency between visual and text information, realizes explicit promotion and enhancement of the complementarity of multi-source information, can effectively explore potential intrinsic pose adaptation cues from multimodal inputs, improves the insufficient adaptability of simple multimodal information fusion, and strengthens the accurate guidance role of multimodal information in pose generation.
[0026] This invention constructs a partial-level Transformer variational autoencoder architecture, using a single human keypoint as the smallest processing unit. It combines self-attention and cross-attention mechanisms to capture fine-grained interactions between keypoints and between keypoints and the scene. At the same time, it introduces probabilistic modeling and latent variable sampling to deeply explore the potential laws of pose deformation. This allows the generated poses to get rid of the stiffness and incoordination caused by global generation, and output more natural, coordinated human poses that conform to the aesthetic standards of portrait photography.
[0027] This invention guides model pre-training by employing a course learning strategy based on prototype complexity. The pose prototypes are graded according to complexity and trained in stages from easy to difficult, allowing the model to gradually adapt to the learning difficulty. This enables the model to first master the basic rules from simple samples and then gradually adapt to complex samples and data noise, greatly improving the robustness of the model and giving it the superior ability to stably output high-quality aesthetically perceptual poses in real-world scenarios.
[0028] This invention reconstructs the implementation path of pose generation and adopts a two-stage generation mode of prototype matching and fine-grained optimization, endowing the model with ideal modular design properties. It breaks the limitations of a single generation mode and makes the model more flexible in adjustment and expansion. The pose prototype library can be changed for different application scenarios, and the optimization target of the model can be adjusted according to different aesthetic requirements, effectively improving the model's adaptability to diverse scenarios and aesthetic needs.
[0029] This invention organically integrates multimodal fusion, probabilistic modeling, and attention mechanism technologies to construct a complete framework for aesthetically perceptible pose generation, achieving efficient collaboration and maximizing strengths while minimizing weaknesses. This effectively promotes a comprehensive improvement in efficiency, accuracy, and aesthetics in pose generation tasks, fully meeting the core needs of scene-based portrait photography for aesthetically perceptible pose generation.
[0030] This invention organically combines course learning strategies with fine-grained interactive modeling of the model, allowing the course learning strategies to run through the entire model pre-training process. This enables the model to accurately filter effective learning information from noisy data, while reasonably measuring the learning value of samples with different complexities. This significantly improves the model's generalization ability in an orderly learning process, generating more scene-adaptable aesthetic perception postures that are accurately adapted to diverse real-world application scenarios. Attached Figure Description
[0031] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application.
[0032] In the attached diagram: Figure 1 This is a flowchart of a method for generating aesthetically perceptual gestures based on visual and textual cues.
[0033] Figure 2 This is a schematic diagram of the structure of the partial-level Transformer variational autoencoder in this invention.
[0034] Figure 3 This is a schematic diagram of the posture prototype library construction process in this invention.
[0035] Figure 4 This is a schematic diagram of the multimodal feature extraction and prototype matching process in this invention.
[0036] Figure 5 This is a schematic diagram of the process of restoring human posture through aesthetic perception in this invention.
[0037] Figure 6 This is a schematic diagram of the training process for the course learning strategy based on prototype complexity in this invention.
[0038] Figure 7 This is a qualitative comparison chart between the present invention and competing baselines. Detailed Implementation
[0039] The detailed description of the following embodiments is used to illustrate the principles of this application, but should not be used to limit the scope of this application. That is, the method and system for generating multilingual code based on domain rules of a large language model in this application are not limited to the described embodiments.
[0040] The present invention will be further described below with reference to embodiments.
[0041] Example 1
[0042] like Figure 1 As shown, an aesthetically perceptual gesture generation method based on visual and textual cues includes the following steps: S1. Obtain multimodal input data, which includes images of scenes without people, specified locations of people in the scene, text descriptions of people's attributes, and text descriptions of scene features. S2. Multimodal feature extraction is performed on the multimodal input data to obtain the corresponding visual features, visual feature maps and text features that are homologous to the visual features; based on the visual features and text features, a suitable target pose prototype is obtained by matching through a pre-trained multimodal prototype classifier. The target pose prototype is selected from a pre-built pose prototype library, which is constructed by normalizing the true values of human poses in the training set used for model training and then using the K-means clustering algorithm. S3. Input the target pose prototype, visual feature map and text features into the pre-trained partial-level Transformer variational autoencoder, and output the predicted deformation parameters, including the scale factor and coordinate and visibility offset. The partial-level Transformer variational autoencoder uses a single human keypoint as the smallest processing unit, models the spatial adaptation relationship between human keypoints and scene environment through cross-attention mechanism, models the limb coordination dependency relationship between human keypoints through self-attention mechanism, and independently encodes the corresponding normal distribution for each human keypoint. S4. The predicted deformation parameters are applied inversely to the target pose prototype to restore the aesthetically perceived human pose that conforms to the aesthetic standards of portrait photography and matches the input scene and the attributes of the person. The human pose consists of 17 human key points, and each human key point contains coordinate information and visibility score information.
[0043] In a specific embodiment, four types of multimodal input data are first acquired: a scene image without human figures, a designated location of a person within the scene, textual descriptions of person attributes, and textual descriptions of scene features. The scene image is uniformly adjusted to a size of 256×256 using bilinear interpolation. The designated location of the human figure is the coordinate of the center point of the pose bounding box. The textual descriptions of person attributes include attributes related to human appearance, clothing, and style. The textual descriptions of scene features include features related to scene terrain, vegetation, lighting, and atmosphere. Multimodal feature extraction is performed on the multimodal input data to obtain corresponding visual features, visual feature maps derived from the visual features, and textual features. Based on the visual and textual features, a suitable target pose prototype is obtained by matching a pre-trained multimodal prototype classifier. The target pose prototype is selected from a pre-built pose prototype library, which is constructed by normalizing the ground truth human pose values in the training set used for model training and then using the K-means clustering algorithm. The target pose prototype, visual feature map, and textual features are then input into a pre-trained partial-level Transformer. The variational autoencoder outputs predicted deformation parameters, including scale factors and coordinate and visibility offsets. The partial-level Transformer variational autoencoder uses a single human keypoint as the smallest processing unit. It models the spatial adaptation relationship between human keypoints and the scene environment through a cross-attention mechanism and the limb coordination dependency relationship between human keypoints through a self-attention mechanism. It independently encodes the corresponding normal distribution for each human keypoint and applies the predicted deformation parameters inversely to the target pose prototype, restoring an aesthetically perceptible human pose that conforms to portrait photography aesthetic standards and matches the input scene and human attributes. The human pose consists of 17 human keypoints, each containing coordinate information and visibility score information. Finally, keypoints with a predicted visibility score below 0.2 are filtered out to ensure the integrity and rationality of the output pose.
[0044] Furthermore, such as Figure 3 As shown, the method for building the pose prototype library specifically includes the following steps: S211. Use the true values of each human pose in the training set to train the model. Normalization is performed: the width of the bounding box of the pose keypoints is calculated. and height Divide the x-coordinates of all key points by y coordinate divided by The pose bounding boxes are uniformly adjusted to unit boxes to eliminate scale differences and obtain normalized poses. ; S212. Using the K-means clustering algorithm, with Euclidean distance as the metric and ignoring invisible keypoints, for all normalized poses... Perform clustering until convergence, then select... Cluster centers serve as pose prototypes A pose prototype library was constructed. This represents the total number of posture prototypes in the posture prototype library. Each posture prototype contains the same number of coordinate and visibility information of key human body points as the aesthetically perceived human posture, covering aesthetic posture types such as standing, sitting, half-squatting, and side-facing.
[0045] In a specific embodiment, the ground truth values of each human pose in the training set used for model training are normalized. The width and height of the bounding box of the pose keypoints are calculated. The x-coordinate of all keypoints is divided by the width of the bounding box, and the y-coordinate is divided by the height of the bounding box. The pose bounding boxes are uniformly adjusted to unit boxes to eliminate the scale differences between different poses and obtain normalized poses. The K-means clustering algorithm is used, with Euclidean distance as the metric and invisible keypoints are ignored. All normalized poses are clustered until convergence. The number of iterations is set to 100. 30 cluster centers are selected as pose prototypes to build a pose prototype library. Each pose prototype contains the same number of coordinate and visibility information of human keypoints as the aesthetically perceived human pose, covering common aesthetic pose types such as standing, sitting, half-squatting, and side-facing.
[0046] Furthermore, such as Figure 4 As shown, in step S2, multimodal feature extraction and prototype matching specifically include the following steps: S221. Visual Feature Extraction: Using a specified location of a person Centered on a single image, two square image patches of different sizes are cropped. These two patches, along with an original scene image scaled to the same preset size, are then input into three independent ResNet-50 backbone networks, resulting in three sets of backbone feature vectors. These three sets of backbone feature vectors are then concatenated and compressed into a fixed-dimensional array. visual features Simultaneously, the unpooled features from the output of the last convolutional layer of the ResNet-50 backbone network are extracted as visual feature maps. ; S222, Text Feature Extraction: Using a pre-trained CLIP-Text encoder with frozen parameters, text descriptions of character attributes are extracted. Scene feature text description The text features are concatenated and compressed into a fixed dimension. Text features ; S223, Prototype Classification and Matching: Classifying and Matching Visual Features Text features pass After the operation is concatenated, the input is a linear classifier, and the output is... The classification log probability of a pose prototype The calculation formula is: ; in, For the categorical log probability, For linear classifiers, As a visual feature, For text features, This is a vector concatenation operation. This represents the total number of pose prototypes in the pose prototype library. pass The operation selects the pose prototype with the highest probability as the target pose prototype. .
[0047] In a specific embodiment, during visual feature extraction, two square image patches of sizes, 256×256 and 128×128, are cropped centered on a designated location on the person. These two image patches, along with the original scene image scaled to 256×256, are input into three independent ResNet-50 backbone networks. The pre-trained weights are the training results from the ImageNet dataset. Each ResNet-50 network outputs a 2048-dimensional feature vector. The output features of the three networks are concatenated to obtain a 6144-dimensional vector, which is then compressed into a 512-dimensional visual feature map through a fully connected layer with a hidden layer dimension of 1024 and an activation function of ReLU. Simultaneously, the unpooled features output from the last convolutional layer of the ResNet-50 backbone network are extracted as the visual feature map. During text feature extraction, a pre-trained CLIP-Text encoder with frozen parameters is used to extract 512-dimensional feature vectors for the text descriptions of the person's attributes and the scene's features. These 512-dimensional feature vectors are concatenated to obtain a 1024-dimensional vector, which is then compressed into a 512-dimensional vector through a fully connected layer with a hidden layer dimension of 512 and an activation function of ReLU. The fully connected layer is compressed into 512-dimensional text features. The visual features and text features are concatenated into a 1024-dimensional feature vector. The input is a linear classifier with two hidden layers. The first layer has a dimension of 1024 and the second layer has a dimension of 512. The activation function for both layers is ReLU. The output is 30-dimensional classification logits. The pose prototype with the highest probability is selected as the target pose prototype through the argmax operation.
[0048] Furthermore, such as Figure 2 As shown, a partial-level Transformer variational autoencoder includes an encoder and a decoder. The encoder's processing steps specifically include the following: S311. Initialize learnable keypoint embeddings. The number of embeddings is equal to the number of human keypoints. Each embedding is used to store the inherent attribute features of the corresponding human joint. S312. Calculate the true offset between the normalized pose and the target pose prototype in the training samples. Combined with the width and height scale factors in the attitude normalization process Construct a real deformation parameter matrix and extract deformation features; S313. Concatenate the keypoint embedding, deformation features, and repetition extensions to text features that match the number of keypoints, and obtain the query embedding through a compression function. The calculation formula is as follows: ; in, For query embedding, For the characteristic compression function, Embedding of learnable key points Deformation characteristics, Text feature vectors repeated for dimension alignment This is a vector concatenation operation; S314. Embed the query into the input cross-attention mechanism, interact with the visual feature map, model the spatial adaptation relationship between human key points and the scene environment, and output the cross-attention features. The calculation formula is as follows: ; in, For cross-attention features, This is the computation function corresponding to the cross-attention mechanism. For query embedding, Visual feature map; S315. Input the cross-attention features into the self-attention mechanism to model the limb coordination dependency relationship between human key points, and output the key point interaction features. The calculation formula is as follows: ; in, For key point interaction features, This is the computation function corresponding to the self-attention mechanism. Cross-attention features; S316. Input the keypoint interaction features into the fully connected layer, and output the mean and variance for each human keypoint to construct an independent normal distribution for each keypoint. The calculation formula is as follows: ; in, The mean of a normal distribution is . The variance of the normal distribution is... For the fully connected layer encoding function, For key point interaction features, This is a vector concatenation operation.
[0049] In a specific embodiment, firstly, 17 learnable keypoint embeddings of 512 dimensions are initialized, with the number of embeddings equal to the number of human keypoints. Each embedding corresponds to a pose keypoint, used to store the joint range of motion, association distance with other keypoints, and other inherent attribute features. The true offset between the normalized pose and the target pose prototype in the training samples is calculated. Combining the width and height scale factors in the pose normalization process, a 17×5 dimension true deformation parameter matrix is constructed. The deformation parameter matrix is input into two fully connected layers, the first layer with a dimension of 1024 and the second layer with a dimension of 512, both using ReLU activation functions to extract 512-dimensional deformation features. The keypoint embeddings, deformation features, and text features that are repeatedly extended to match the dimension of the number of keypoints are concatenated into a 17×1536-dimensional matrix. A compression function containing fully connected layers, ReLU activation, and layer normalization is used to obtain a 17×512 matrix. The query embedding is input into the cross-attention mechanism, which interacts with the visual feature map to model the spatial adaptation relationship between human keypoints and the scene environment, and outputs cross-attention features. The cross-attention module adopts the MaskFormer architecture with 8 multi-head attention heads and a head dimension of 64. The cross-attention features are input into the self-attention mechanism to model the limb coordination dependency relationship between human keypoints and output 17×512-dimensional keypoint interaction features. The self-attention module also has 8 attention heads and a head dimension of 64. The keypoint interaction features are input into the fully connected layer with a weight dimension of 512×128 and a bias term of 128 dimensions, outputting the corresponding mean and variance for each human keypoint and constructing an independent normal distribution for each keypoint.
[0050] Furthermore, the decoder's processing procedure specifically includes the following steps: S321. From the normal distribution corresponding to each key point Sample latent variables separately To form a latent variable matrix ,Will Input fully connected layers extract latent features The calculation formula is: ; in, Indicates a normal distribution. For the first Each key point corresponds to the mean of a normal distribution. For the first Each key point corresponds to the variance of a normal distribution. For from the first Each key point corresponds to a latent variable sampled from a normal distribution. The latent variable matrix consists of the latent variables of all key points. It is a linear rectified activation function. For the weights of the fully connected layer, For bias terms of fully connected layers; S322, Embed key points Potential characteristics Repeated expansion to text features matching the number of key points Concatenate and compress to obtain the decoded query embedding ; S323, Embed the decoding query After being processed by the cross-attention computation unit and the self-attention computation unit in sequence, the decoded features are output. ; S324, Decode the features The input is a fully connected layer, and the predicted deformation parameters are obtained by decoding. The calculation formula is as follows: ; in, The predicted scaling factor matrix, For the predicted coordinates and visibility offset matrix, Decoding function for fully connected layer For the weights of the fully connected layer, For bias terms of fully connected layers, This indicates a matrix concatenation operation.
[0051] In a specific embodiment, 64-dimensional latent variables are sampled from the normal distribution corresponding to each keypoint to form a 17×64-dimensional latent variable matrix. The latent variable matrix is input into a fully connected layer with a weight dimension of 64×512, a bias term of 512, and a ReLU activation function to extract 17×512-dimensional latent features. The embedding fusion logic of the repeat encoder concatenates and compresses the keypoint embedding, latent features, and text features extended to match the number of keypoints to obtain the decoded query embedding. The decoded query embedding is then processed by the cross-attention calculation unit and the self-attention calculation unit to output 17×512-dimensional decoded features. The decoded features are then input into a fully connected layer with a weight dimension of 512×5 and a bias term of 5 to decode 17×5-dimensional predicted deformation parameters. The predicted deformation parameters include a 17×2-dimensional scale factor matrix and a 17×3-dimensional coordinate and visibility offset matrix.
[0052] Furthermore, the total loss function used for pre-training the partial-level Transformer variational autoencoder is: ; in, The total loss function for model training. For KL divergence loss, To balance hyperparameters, For attitude loss; KL divergence loss The calculation formula is: ; in, An index of key points on the human body. For the first Each key point corresponds to the variance of a normal distribution. For the first Each key point corresponds to the mean of a normal distribution. For the first Each key point corresponds to the square of the L2 norm of the mean. For the first The logarithm of the determinant of the variance corresponding to each key point; Attitude loss The calculation formula is: ; in, An index of key points on the human body. For the first The true visibility score of each key point For the first Predicted coordinate offsets of key points For the first The true coordinate offset of each key point The cross-entropy loss function is used for binary classification. For the first Predictive visibility score for each key point The set of true visibility scores for all key points; Predictive visibility score of key points The calculation formula is: ; in, The predicted visibility score for key points. It is the Sigmoid activation function. It is the inverse of the Sigmoid function. The visibility score of the target pose prototype. This represents the predicted visibility offset.
[0053] In a specific embodiment, the total loss function used for pre-training the partial-level Transformer variational autoencoder consists of two parts: KL divergence loss and pose loss. The balancing hyperparameter is set to 0.1. The KL divergence loss is used to minimize the difference between the predicted distribution and the standard normal distribution, thereby improving generation stability. The KL divergence of the corresponding normal distribution is calculated for each of the 17 keypoints and then averaged. The pose loss includes the coordinate regression loss of the visible keypoints and the binary cross-entropy loss for visibility classification. The coordinate regression loss uses L2 loss to calculate the difference between the predicted coordinate offset and the true coordinate offset of the visible keypoints. The visibility classification loss uses the Sigmoid activation function to calculate the binary cross-entropy between the predicted visibility score and the true visibility score.
[0054] Furthermore, such as Figure 6 As shown, the partial-level Transformer variational autoencoder is trained using a prototype complexity-based course learning strategy, specifically including the following steps: S331. Calculate the average Euclidean distance between each posture prototype in the posture prototype library and all other prototypes, and use it as a measure of the complexity of the corresponding prototype. S332. Sort all pose prototypes in ascending order of complexity metric, divide them into three groups, corresponding to three training courses from easy to difficult. S333. Train in stages according to the order of course difficulty from low to high. After completing the preset training rounds for each course, move on to the next course until all courses are completed.
[0055] In a specific embodiment, the partial-level Transformer variational autoencoder is trained using a course learning strategy based on prototype complexity. The average Euclidean distance between each pose prototype and the other 29 prototypes in the pose prototype library is calculated as a complexity metric for the corresponding prototype. The larger the distance, the more complex the prototype. The 30 pose prototypes are sorted in ascending order of complexity metric and divided into three groups, corresponding to three training courses from easy to difficult. Each group contains 10 prototypes. The training is carried out in stages according to the order of course difficulty from low to high. First, the first course is trained using only samples corresponding to simple prototypes for 30 rounds. Then, the second course is trained using samples corresponding to the first 20 prototypes for 30 rounds. Finally, the third course is trained using samples corresponding to all 30 prototypes for 30 rounds. The learning rate of each course is kept constant at 1e-4, and the batch size is 8. By gradually increasing the learning difficulty, the interference of noisy data is reduced, and the robustness of the model is improved.
[0056] Furthermore, such as Figure 5As shown, in step S4, the predicted deformation parameters are applied inversely to the target posture prototype to restore the aesthetically perceived human posture. This process includes the following steps: S41. The attitude scale of the target attitude prototype is restored by predicting the scale factor in the deformation parameters. The calculation formula is as follows: ; Where S is the predicted scale factor matrix, and Δ is the predicted coordinate and visibility offset matrix. The x-axis original coordinates of the key points in the target pose prototype. The original y-axis coordinates of key points in the target pose prototype. These are the x-axis coordinates of the key points after scaling. These are the y-axis coordinates of the key points after scaling. The 0th column of the scale factor matrix corresponds to the x-axis scale coefficient. This is the first column of the scale factor matrix, corresponding to the y-axis scale coefficient; S42. Adjust the coordinate position of key points by predicting the coordinate offset in the deformation parameters. The calculation formula is as follows: ; in, The final output contains the x-axis coordinates of each key point. The final output contains the y-axis coordinates of each key point. The 0th column of the offset matrix corresponds to the x-axis coordinate offset. The first column of the offset matrix corresponds to the y-axis coordinate offset. S43, via offset Adjust the keypoint visibility score and filter keypoints with a predicted visibility score V < 0.2 to obtain the final aesthetically perceived human posture.
[0057] In a specific embodiment, when the predicted deformation parameters are applied inversely to the target posture prototype to restore the aesthetically perceived human posture, the posture scale of the target posture prototype is first restored by the scale factor in the predicted deformation parameters. The original x-axis coordinates of the key points in the target posture prototype are multiplied by the value of the 0th column of the scale factor matrix, and the original y-axis coordinates are multiplied by the value of the 1st column of the scale factor matrix to complete the scale restoration. Then, the coordinate positions of the key points are adjusted by the coordinate offset in the predicted deformation parameters. The scale-restored x-axis coordinates are added to the value of the 0th column of the offset matrix, and the scale-restored y-axis coordinates are added to the value of the 1st column of the offset matrix to obtain the final output coordinates of each key point. The visibility score of the key points is adjusted by the value of the 2nd column of the offset matrix, and key points with a predicted visibility score lower than 0.2 are filtered out to obtain the final aesthetically perceived human posture. The generated posture is natural and coordinated, the limb proportions are reasonable, and it is highly adapted to the scene features and human attribute descriptions, which conforms to the principles of photographic aesthetic composition.
[0058] Example 2
[0059] This embodiment is implemented on a Python 3.10, PyTorch 2.9, and Ubuntu 18.04 compatible operating system, with an AMD EPYC 8324P CPU and two NVIDIA RTX 4090 GPUs. To ensure experimental reproducibility, unless otherwise specified, the random seed for all experiments is fixed at 0. For data augmentation, random horizontal flipping, random vertical flipping, random rotation within ±10 degrees, and color jitter are used as preprocessing operations to improve the model's generalization ability. After augmentation, the input image is uniformly adjusted to 256×256 pixels, and bilinear interpolation is used for scaling to maintain spatial consistency.
[0060] The dataset used in this embodiment was collected from four platforms: Depositphotos, Getty Images, Xiaohongshu, and Visual China (VCG). After initial image crawling, images containing a single prominent human figure were selected and retained, while images with multiple people were excluded to reduce task complexity. Ultimately, 25,000 valid images were retained from each platform, totaling 100,000 samples to form the complete experimental dataset. To verify the aesthetic value of the collected images, the HumanAesExpert model was used to evaluate the aesthetic quality of the samples from each platform. The scene aesthetic scores for each platform were 0.6194 ( Depositphotos), 0.5992 (Xiaohongshu), 0.6212 (Getty Images), and 0.6549 (Visual China), respectively, while the human figure aesthetic scores were 0.6644 ( Depositphotos), 0.6275 (Xiaohongshu), 0.6433 (Getty Images), and 0.6542 (Visual China), respectively. The overall aesthetic scores were higher than the reference values (scene aesthetics 0.5708, human figure aesthetics 0.5632), verifying that the aesthetic attributes of the samples met the experimental requirements.For each sample in the dataset, the VitPose model was used to estimate human pose, which includes 17 keypoints (nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, and left and right ankles). A confidence threshold of 0.2 was set, retaining visible or high-confidence keypoints and filtering low-confidence noise keypoints. The PCKh indexes for pose estimation on various platforms were 0.915 ( Depositphotos), 0.753 (Xiaohongshu), 0.840 (Getty Images), and 0.750 (Visual China), ensuring the validity of the ground truth pose values. The center coordinates of the pose bounding box were calculated as the specified human position in the sample. Two sets of text descriptions were generated using the Google Gemini model. The scene text description focused on terrain, vegetation, lighting, atmosphere, and spatial composition, strictly excluding any content related to the human body. The human body text description focused on the person's appearance, clothing, style, and overall impression, without involving any pose or action descriptions, based on CLIP. The model calculated the text-image matching score. The text matching scores for people across platforms were 0.1974 ( Depositphotos), 0.1958 (Xiaohongshu), 0.2017 (Getty Images), and 0.1932 (Visual China), respectively. The text matching scores for scenes were 0.2367 ( Depositphotos), 0.2059 (Xiaohongshu), 0.2193 (Getty Images), and 0.1903 (Visual China), respectively. These scores were significantly higher than the reference matching scores of 0.1412 (people) and 0.1563 (scenes), validating the correlation between text descriptions and images. The Wan 2.2 model was used to repair human body regions in the images. This model was trained on billions of high-quality images and videos to obtain complete scene images without human figures. The image repair quality was evaluated using the QualiCLIP model. The repair quality scores across platforms were 0.3683 ( Depositphotos), 0.6477 (Xiaohongshu), 0.3727 (Getty Images), and 0.4229 (Visual China), respectively, reaching the reference quality threshold of 0.4 overall. The dataset was constructed using the established standards. Simultaneously, 1000 samples were randomly selected from the dataset, and human poses were manually corrected to create an independent validation set. This validation set, with no overlap with the training set, was used for quantitative evaluation of the model's performance. Additionally, 100 real-world scene images without human figures were collected for subjective effect verification in real-world scenarios.
[0061] like Figure 7 As shown, this embodiment selects three mainstream existing technical methods in the industry as comparison baselines, namely: The first category of traditional pose estimators includes Heatmap, Regression, UniPose (B. Artacho and ASavakis. UniPose: Unified Human Pose Estimation in Single Images and Videos. IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2020, pp. 7033-7042, 2020. 7), and PRTR (Ke Li, Shijie Wang, Xiang Zhang, Yifan Xu, Weijian Xu, and Zhuowen Tu. Pose Recognition Based on Cascaded Transformer. IEEE / CVF Conference on Computer Vision and Pattern Recognition Proceedings, pp. 1944-1953, 2021. 7); The second category of target placement algorithms includes PlaceNet (Lingzhi Zhang, Tarmily Wen, Jie Min, Jiancong Wang, David Han, and Jianbo Shi. Object Placement Learning Based on Repair for Synthetic Data Augmentation. European Conference on Computer Vision Proceedings, pp. 566-581. Springer, 2020. 7), and GracoNet (Siyuan Zhou, Liu Liu, LiNiu, and Liqing Zhang. Object placement learning based on dual-path graph completion. Proceedings of the European Conference on Computer Vision, pp. 373-389. Springer, 2022. 7); The third category of existing pose generation methods: Wang et al. (Xiaolong Wang, Rohit Girdhar, and Abhinav Gupta. binge watching: scaling usability learning from comedy sketches. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2596-2605, 2017.), Zhang et al. (Lingzhi Zhang, Weiyu Du, Shenghao Zhou, Jiancong Wang, and Jianbo Shi. Inpaint2learn: a self-supervised framework for usability learning. Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, pp. 2665-2674, 2022. 1, 2, 3, 7, 8), Yao et al. (Jieteng Yao, Junjie Chen, Yixin Chen, and Liang Wang. Human pose generation in the scene based on prototype guidance.Proceedings of the International Conference on Computer Vision, pp. 14567-14577, 2023. 1, 2, 4, 7, 8; MAMC (PrasunRoy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal, and Michael Blumenstein. Exploring Interactive Modal Attention for Context-Aware Human Usability Generation. IEEE Transactions on Artificial Intelligence, 2025. 5, 6, 7, 8).
[0062] The comparative experiments were conducted on a manually calibrated validation set of 1000 samples, using seven industry-standard core evaluation metrics: PCK accuracy, PCKh accuracy, AKD (Anatomical Angle Difference), mean absolute error (MAE), mean squared error (MSE), pose similarity (SIM), and pose region intersection-union ratio (IOU).
[0063] To verify the contribution of each core module of this invention, this embodiment conducts ablation experiments by disabling different modules. The experiments are performed on the same validation set, and PCK, PCKh, AKD, MAE, and MSE are used as core evaluation indicators, as shown in the table below: In addition, this embodiment adopts the subjective user research scheme of the MAMC method, inviting 30 volunteers to participate in the test, and comparing the pose generation effect of the method of this invention with that of three mainstream existing technologies: Zhang et al., Yao et al., and MAMC. The mean opinion score (MOS) is used as the core evaluation index in this test, as shown in the table below: Therefore, the MOS of this method is significantly better than the baseline, and 0.4 higher than MAMC, demonstrating its advantage in generating aesthetically perceptible poses.
[0064] It should be noted that the combination of the technical features in this case is not limited to the combination methods described in the claims of this case or the combination methods described in the specific embodiments. All technical features described in this case can be freely combined or combined in any way, unless they contradict each other.
[0065] It should also be noted that the embodiments listed above are merely specific embodiments of the present invention. Obviously, the present invention is not limited to the above embodiments, and similar changes or modifications made thereto are those that can be directly derived or easily conceived by those skilled in the art from the content disclosed in the present invention, and should all fall within the protection scope of the present invention.
Claims
1. A method for generating aesthetically perceptual postures based on visual and textual cues, characterized in that, Includes the following steps: S1. Obtain multimodal input data, which includes a scene image without people, a specified location of a person in the scene, a text description of the person's attributes, and a text description of the scene features. S2. Multimodal feature extraction is performed on the multimodal input data to obtain the corresponding visual features, visual feature maps and text features that are homologous to the visual features; based on the visual features and text features, a suitable target pose prototype is obtained by matching through a pre-trained multimodal prototype classifier. The target pose prototype is selected from a pre-built pose prototype library, which is constructed by normalizing the true values of human poses in the training set used for model training and then using the K-means clustering algorithm. S3. Input the target pose prototype, visual feature map, and text features into a pre-trained partial-level Transformer variational autoencoder, and output predicted deformation parameters, including scale factors and coordinate and visibility offsets. The partial-level Transformer variational autoencoder uses a single human keypoint as the smallest processing unit, models the spatial adaptation relationship between human keypoints and the scene environment through a cross-attention mechanism, models the limb coordination dependency relationship between human keypoints through a self-attention mechanism, and independently encodes the corresponding normal distribution for each human keypoint. The partial-level Transformer variational autoencoder includes an encoder and a decoder. The encoder processing specifically includes the following steps: S311. Initialize learnable keypoint embeddings. The number of embeddings is equal to the number of human keypoints. Each embedding is used to store the inherent attribute features of the corresponding human joint. S312. Calculate the true offset between the normalized pose and the target pose prototype in the training samples. Combined with the width and height scale factors in the attitude normalization process Construct a real deformation parameter matrix and extract deformation features; S313. The keypoint embedding, deformation features, and repetition extensions are concatenated with the text features that match the number of keypoints. The query embedding is obtained through a compression function, and the calculation formula is as follows: ; in, For query embedding, For the characteristic compression function, Embedding of learnable key points Deformation characteristics, Text feature vectors repeated for dimension alignment This is a vector concatenation operation; S314. The query is embedded into the input cross-attention mechanism, interacts with the visual feature map, models the spatial adaptation relationship between human key points and scene environment, and outputs cross-attention features. The calculation formula is as follows: ; in, For cross-attention features, This is the computation function corresponding to the cross-attention mechanism. For query embedding, Visual feature map; S315. Input the cross-attention features into the self-attention mechanism to model the limb coordination dependency relationship between human key points, and output the key point interaction features. The calculation formula is as follows: ; in, For key point interaction features, This is the computation function corresponding to the self-attention mechanism. Cross-attention features; S316. Input the keypoint interaction features into the fully connected layer, and output the mean and variance for each human keypoint to construct an independent normal distribution for each keypoint. The calculation formula is as follows: ; in, The mean of a normal distribution is . The variance of the normal distribution is... For the fully connected layer encoding function, For key point interaction features, This is a vector concatenation operation; S4. The predicted deformation parameters are applied in reverse to the target pose prototype to restore the aesthetically perceived human pose that conforms to the aesthetic standards of portrait photography and matches the input scene and the attributes of the person. The human pose consists of 17 human key points, and each human key point contains coordinate information and visibility score information.
2. The method for generating aesthetically perceptual postures based on visual and textual cues according to claim 1, characterized in that: The method for constructing the pose prototype library specifically includes the following steps: S211. Use the true values of each human pose in the training set to train the model. Normalization is performed: the width of the bounding box of the pose keypoints is calculated. and height Divide the x-coordinates of all key points by y coordinate divided by By uniformly adjusting the pose bounding boxes to unit boxes to eliminate scale differences, normalized poses are obtained. ; S212. Using the K-means clustering algorithm, with Euclidean distance as the metric and ignoring invisible keypoints, for all normalized poses... Perform clustering until convergence, then select... Cluster centers serve as pose prototypes A pose prototype library was constructed. The total number of posture prototypes in the posture prototype library. Each posture prototype contains the same number of coordinate and visibility information of key human body points as the aesthetically perceived human posture, covering aesthetic posture types such as standing, sitting, half-squatting, and side-facing.
3. The method for generating aesthetically perceptual postures based on visual and textual cues according to claim 1, characterized in that, In step S2, multimodal feature extraction and prototype matching specifically include the following steps: S221. Visual feature extraction: using the specified location of the person... Centered on a single image, two square image patches of different sizes are cropped. These two patches, along with an original scene image scaled to the same preset size, are then input into three independent ResNet-50 backbone networks, resulting in three sets of backbone feature vectors. These three sets of backbone feature vectors are then concatenated and compressed into a fixed-dimensional array. visual features Simultaneously, the unpooled features from the output of the last convolutional layer of the ResNet-50 backbone network are extracted as visual feature maps. ; S222, Text Feature Extraction: Using a pre-trained CLIP-Text encoder with frozen parameters, extract the text descriptions of the character attributes respectively. The scene feature text description The text features are concatenated and compressed into a fixed dimension. Text features ; S223, Prototype Classification and Matching: Classifying and matching the aforementioned visual features... With the text features pass After the operation is concatenated, the input is a linear classifier, and the output is... The classification log probability of a pose prototype The calculation formula is: ; in, For the categorical log probability, For linear classifiers, As a visual feature, For text features, This is a vector concatenation operation. This represents the total number of pose prototypes in the pose prototype library. pass The operation selects the pose prototype with the highest probability as the target pose prototype. .
4. The method for generating aesthetically perceptual postures based on visual and textual cues according to claim 1, characterized in that: The decoding process specifically includes the following steps: S321. From the normal distribution corresponding to each key point Sample latent variables separately To form a latent variable matrix ,Will Input fully connected layers extract latent features The calculation formula is: ; in, Indicates a normal distribution. For the first Each key point corresponds to the mean of a normal distribution. For the first Each key point corresponds to the variance of a normal distribution. For from the first Each key point corresponds to a latent variable sampled from a normal distribution. The latent variable matrix consists of the latent variables of all key points. It is a linear rectified activation function. For the weights of the fully connected layer, For bias terms of fully connected layers; S322, Embed the key points Potential characteristics Repeatedly expand the text features to match the number of key points. Concatenate and compress to obtain the decoded query embedding ; S323, Embed the decoding query After being processed by the cross-attention computation unit and the self-attention computation unit in sequence, the decoded features are output. ; S324, Decode the features The input is a fully connected layer, and the predicted deformation parameters are obtained by decoding. The calculation formula is as follows: ; in, The predicted scaling factor matrix, For the predicted coordinates and visibility offset matrix, Decoding function for fully connected layer For the weights of the fully connected layer, For bias terms of fully connected layers, This indicates a matrix concatenation operation.
5. The method for generating aesthetically perceptual postures based on visual and textual cues according to claim 1, characterized in that: The total loss function used for the pre-training of the partial-level Transformer variational autoencoder is: ; in, The total loss function for model training. For KL divergence loss, To balance hyperparameters, For attitude loss; The KL divergence loss The calculation formula is: ; in, An index of key points on the human body. For the first Each key point corresponds to the variance of a normal distribution. For the first Each key point corresponds to the mean of a normal distribution. For the first Each key point corresponds to the square of the L2 norm of the mean. For the first The logarithm of the determinant of the variance corresponding to each key point; The attitude loss The calculation formula is: ; in, An index of key points on the human body. For the first The true visibility score of each key point For the first Predicted coordinate offsets of key points For the first The true coordinate offset of each key point The cross-entropy loss function is used for binary classification. For the first Predictive visibility score for each key point The set of true visibility scores for all key points; The predicted visibility score of the key points The calculation formula is: ; in, The predicted visibility score for key points. It is the Sigmoid activation function. It is the inverse of the Sigmoid function. The visibility score of the target pose prototype. This represents the predicted visibility offset.
6. The method for generating aesthetically perceptual postures based on visual and textual cues according to claim 5, characterized in that: The partial-level Transformer variational autoencoder is trained using a prototype complexity-based course learning strategy, specifically including the following steps: S331. Calculate the average Euclidean distance between each posture prototype in the posture prototype library and all other prototypes, and use it as the complexity metric of the corresponding prototype. S332. Sort all pose prototypes in ascending order of complexity metric, and divide them into three groups on average, corresponding to three training courses from easy to difficult. S333. Train in stages according to the order of course difficulty from low to high. After completing the preset training rounds for each course, move on to the next course until all courses are completed.
7. The method for generating aesthetically perceptual postures based on visual and textual cues according to claim 1, characterized in that: In step S4, the predicted deformation parameters are applied inversely to the target posture prototype to restore the aesthetically perceived human posture, which includes the following steps: S41. The attitude scale of the target attitude prototype is restored by predicting the scale factor in the deformation parameters. The calculation formula is as follows: ; Where S is the predicted scale factor matrix, and Δ is the predicted coordinate and visibility offset matrix. Key points in the target pose prototype The original coordinates of the axis Key points in the target pose prototype The original coordinates of the axis Keypoints after scale reduction Axis coordinates Keypoints after scale reduction Axis coordinates The 0th column of the scale factor matrix corresponds to Axis dimension factor, The first column of the scale factor matrix corresponds to Axis dimension factor; S42. Adjust the coordinate position of key points by predicting the coordinate offset in the deformation parameters. The calculation formula is as follows: ; in, The final output contains the x-axis coordinates of each key point. The final output contains the y-axis coordinates of each key point. The 0th column of the offset matrix corresponds to the x-axis coordinate offset. The first column of the offset matrix corresponds to the y-axis coordinate offset. S43, via offset Adjust the keypoint visibility score and filter keypoints with a predicted visibility score V < 0.2 to obtain the final aesthetically perceived human posture.
Citation Information
Patent Citations
Character image generation method guided by text based on generative adversarial network
CN110021051A
Positional encoding for neural network attention
WO2024207020A1