Method and device for processing character object in image, storage medium and program product

By extracting feature vectors from the character image and processing them using a self-attention mechanism to generate 2D and 3D Mesh point coordinates, the problem of insufficient three-dimensional spatial information capture in traditional methods is solved, and the body beauty effect is improved.

CN120339799APending Publication Date: 2025-07-18XIAMEN MEITUZHIJIA TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478796.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional human posture estimation methods are difficult to accurately capture three-dimensional spatial information, resulting in poor body beauty effects of characters and objects in complex postures and obstructions.

Method used

By extracting feature vectors from the character image, integrating position coded vectors, and coding with self-attention mechanism, 2D Mesh point coordinates, 3D Mesh point coordinates and stretching directions are generated, and context relationships are captured using the Transformer model.

Benefits of technology

It realizes accurate prediction of 2D Mesh point coordinates, 3D Mesh point coordinates and stretching directions under complex postures and occlusions, improving the body beauty effect of the character object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339799A_ABST
    Figure CN120339799A_ABST
Patent Text Reader

Abstract

The invention relates to a character object processing method and device, a storage medium and a program product. Extracting a feature vector from the character image containing the character to be processed; fusing the position coding vector into the feature vector to obtain a target feature vector; encoding the target feature vector to obtain an encoded feature vector with a context relationship; the step of coding the target feature vector comprises the following steps: in a coding process, carrying out self-attention processing on the target feature vector by adopting a self-attention mechanism so as to capture the context relationship; decoding the coding feature vector to obtain a 2D Mesh point coordinate, a 3D Mesh point coordinate and a stretching direction; and based on the 2D Mesh point coordinates, the 3D Mesh point coordinates and the stretching direction, generating a figure beautifying figure corresponding to the figure to be processed. By adopting the method, the figure beautifying effect of the figure object can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technologies, and particularly to a method, apparatus, storage medium, and program product for processing human objects in images. Background Art

[0002] With the rapid development of computer vision technologies, human pose estimation, as an important branch thereof, shows broad application prospects in body shaping. Traditional human pose estimation methods are mainly divided into methods based on 2D key points. Such methods define a small number of key points on the two-dimensional image plane to represent the shape and structure of the human body, and use these key points for pose estimation. Due to their advantages such as high computational efficiency and simple implementation, they have been widely applied in the past period of time.

[0003] However, with the increasing complexity of application scenarios, the methods based on 2D key points gradually expose their limitations. Due to the lack of effective modeling of depth information, such methods are difficult to accurately capture the three-dimensional spatial information of the human body, resulting in poor performance in processing complex poses, occlusion and other scenarios, and thus leading to poor body shaping effects of human objects. Summary of the Invention

[0004] Based on this, the present application provides a method, apparatus, storage medium, and program product for processing human objects in images, which can improve the body shaping effect of human objects.

[0005] On the one hand, the present application provides a method for processing human objects in images, including:

[0006] Extracting a feature vector from a human image including a human object to be processed;

[0007] Incorporating a position encoding vector into the feature vector to obtain a target feature vector;

[0008] Encoding the target feature vector to obtain an encoded feature vector with context relationships; the encoding of the target feature vector includes, during the encoding process, performing self-attention processing on the target feature vector by using a self-attention mechanism to capture the context relationships;

[0009] Decoding the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and a stretching direction;

[0010] Generating a body-shaped human corresponding to the human object to be processed based on the 2D Mesh point coordinates, the 3D Mesh point coordinates, and the stretching direction.

[0011] In one embodiment, the extracting a feature vector from a human image including a human object to be processed includes:

[0012] Extract a first feature map from a person image including a person to be processed;

[0013] Process the first feature map through a ResNet network to obtain a second feature map;

[0014] Flatten the second feature map into a feature vector.

[0015] In one embodiment, the extracting a first feature map from a person image including a person to be processed includes:

[0016] Extract features from a person image including a person to be processed through a feature extraction network, and during the feature extraction process, perform multiple information exchanges between parallel sub-networks in the feature extraction network to obtain a first feature map.

[0017] In one embodiment, the extracting features from a person image including a person to be processed through a feature extraction network, and during the feature extraction process, performing multiple information exchanges between parallel sub-networks in the feature extraction network to obtain a first feature map includes:

[0018] Input a person image including a person to be processed into the feature extraction network; the feature extraction network includes a first sub-network, a second sub-network, and a third sub-network in parallel;

[0019] Extract features from the person image through the feature extraction network, and during the feature extraction process, the first sub-network, the second sub-network, and the third sub-network respectively process feature maps with different resolutions, and during the processing, perform cross-resolution information exchanges between different sub-networks to obtain a first feature map.

[0020] In one embodiment, the target feature vector is obtained by encoding through an encoder of a Transformer model, and the encoder includes a multi-head self-attention layer, a feed-forward neural network layer, as well as residual connections and layer normalization;

[0021] The encoding the target feature vector to obtain an encoded feature vector with context relationship includes:

[0022] Input the target feature vector into the encoder of the Transformer model;

[0023] For each head in the multi-head self-attention layer, the self-attention mechanism is respectively used to linearly process the target feature vector to obtain a query vector, a key vector, and a value vector, calculate the attention score according to the query vector and the key vector, perform normalization processing on the attention score, and obtain context information based on the normalized attention score and the value vector;

[0024] The context information is non-linearly transformed through the feed-forward neural network layer to obtain a non-linearly transformed feature;

[0025] Through the residual connection and layer normalization, the non-linearly transformed feature is fused with the target feature vector, and the obtained fused feature is normalized to obtain an encoded feature vector with context relationship.

[0026] In one embodiment, the decoding of the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching direction includes:

[0027] Input the encoded feature vector into the decoder of the Transformer model;

[0028] Through the decoder, cross-attention processing is performed on the encoded feature vector and the target query vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching direction.

[0029] In one embodiment, the method further includes:

[0030] Extract a training feature vector from a person image sample containing a target person;

[0031] Integrate the position encoding vector into the training feature vector to obtain a training target feature vector;

[0032] Encode the training target feature vector through the encoder of the initial Transformer model to obtain a training encoded feature vector with context relationship; the encoding of the training target feature vector includes, during the encoding process, performing self-attention processing on the training feature vector by using the self-attention mechanism to capture the context relationship;

[0033] Decode the training encoded feature vector through the decoder of the initial Transformer model to obtain predicted 2D Mesh point coordinates, predicted 3D Mesh point coordinates, and predicted stretching direction;

[0034] Determine a loss value based on the predicted 2D Mesh point coordinates, the predicted 3D Mesh point coordinates, the predicted stretching direction, and the corresponding labels; wherein, the labels are 2D Mesh point labels, 3D Mesh point labels, and stretching direction labels obtained by processing based on the SMPL human model;

[0035] Optimize the initial Transformer model according to the loss value until the model converges.

[0036] On the one hand, the present application also provides a device for processing a human object in an image, including:

[0037] An extraction module, configured to extract a feature vector from a human image including a person to be processed;

[0038] A fusion module, configured to integrate a position encoding vector into the feature vector to obtain a target feature vector;

[0039] An encoding module, configured to encode the target feature vector to obtain an encoded feature vector with context relationships; the encoding of the target feature vector includes, during the encoding process, performing self-attention processing on the target feature vector by using a self-attention mechanism to capture the context relationships;

[0040] A decoding module, configured to decode the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and a stretching direction;

[0041] A generation module, configured to generate a figure with a beautiful figure corresponding to the person to be processed based on the 2D Mesh point coordinates, the 3D Mesh point coordinates, and the stretching direction.

[0042] In one embodiment, the extraction module is further configured to extract a first feature map from a human image including a person to be processed; process the first feature map through a ResNet network to obtain a second feature map; and flatten the second feature map into a feature vector.

[0043] In one embodiment, the extraction module is further configured to perform feature extraction on a human image including a person to be processed through a feature extraction network, and during the feature extraction process, perform multiple information exchanges between parallel sub-networks in the feature extraction network to obtain a first feature map.

[0044] In one embodiment, the extraction module is further configured to input a person image including a person to be processed into a feature extraction network; the feature extraction network includes a first sub-network, a second sub-network, and a third sub-network in parallel; the feature extraction network performs feature extraction on the person image, and during the feature extraction process, the first sub-network, the second sub-network, and the third sub-network respectively process feature maps of different resolutions, and during the processing, cross-resolution information exchange is performed between different sub-networks to obtain a first feature map.

[0045] In one embodiment, the target feature vector is obtained by encoding through the encoder of the Transformer model, and the encoder includes a multi-head self-attention layer, a feed-forward neural network layer, as well as residual connections and layer normalization;

[0046] The encoding module is further configured to input the target feature vector into the encoder of the Transformer model; for each head in the multi-head self-attention layer, the self-attention mechanism is respectively used to linearly process the target feature vector to obtain a query vector, a key vector, and a value vector, calculate an attention score according to the query vector and the key vector, perform normalization processing on the attention score, and obtain context information based on the normalized attention score and the value vector; perform a non-linear transformation on the context information through the feed-forward neural network layer to obtain a non-linearly transformed feature; through the residual connection and layer normalization, fuse the non-linearly transformed feature with the target feature vector, and perform normalization processing on the obtained fused feature to obtain an encoded feature vector with context relationship.

[0047] In one embodiment, the decoding module is further configured to input the encoded feature vector into the decoder of the Transformer model;

[0048] Through the decoder, cross-attention processing is performed on the encoded feature vector and the target query vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and a stretching direction.

[0049] In one embodiment, the device further includes:

[0050] A training module, configured to extract training feature vectors from a person image sample containing a target person; integrate a position encoding vector into the training feature vectors to obtain training target feature vectors; encode the training target feature vectors through an encoder of an initial Transformer model to obtain training encoded feature vectors with context relationships; the encoding of the training target feature vectors includes, during the encoding process, performing self-attention processing on the training feature vectors by using a self-attention mechanism to capture the context relationships; decode the training encoded feature vectors through a decoder of the initial Transformer model to obtain predicted 2D Mesh point coordinates, predicted 3D Mesh point coordinates, and a predicted stretching direction; determine a loss value according to the predicted 2D Mesh point coordinates, the predicted 3D Mesh point coordinates, the predicted stretching direction, and corresponding labels; wherein the labels are 2D Mesh point labels, 3D Mesh point labels, and stretching direction labels obtained by processing according to the SMPL human model; optimize the initial Transformer model according to the loss value until the model converges.

[0051] On the one hand, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method for processing a person object in the above image are implemented.

[0052] On the one hand, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for processing a person object in the above image are implemented.

[0053] On the one hand, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method for processing a person object in the above image are implemented.

[0054] The method, device, storage medium, and program product for processing a human object in an image extract a feature vector from a human image including a human object to be processed; integrate a position encoding vector into the feature vector to obtain a target feature vector; encode the target feature vector to obtain an encoded feature vector with context relationships. During the encoding process, a self-attention mechanism is used to perform self-attention processing on the target feature vector to capture context relationships; decode the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and a stretching direction, thereby introducing three-dimensional technology into the body shaping of the human object, accurately capturing the three-dimensional spatial information of the human body. Therefore, even in the case of complex postures and occlusions, the 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching direction can be accurately predicted, and based on the 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching direction, a body-shaped human corresponding to the human object to be processed is generated, thereby effectively improving the body shaping effect of the human object. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the following-described drawings are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 It is an application environment diagram of the method for processing a human object in an image in an embodiment;

[0057] Figure 2 It is a flowchart of the method for processing a human object in an image in an embodiment;

[0058] Figure 3 It is a schematic structural diagram of a body shaping model in an embodiment;

[0059] Figure 4 It is a schematic structural diagram of a feature extraction model in an embodiment;

[0060] Figure 5 It is a schematic diagram of Mesh points in an embodiment;

[0061] Figure 6 It is a schematic diagram of Mesh points in another embodiment;

[0062] Figure 7 It is a flowchart of model training in an embodiment;

[0063] Figure 8 It is a block diagram of the structure of the device for processing a human object in an image in an embodiment;

[0064] Figure 9 is a structural block diagram of a human object processing device in an image in another embodiment;

[0065] Figure 10 is an internal structure diagram of a computer device in one embodiment. Detailed implementation manners

[0066] In order to make the objectives, technical solutions, and beneficial effects of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0067] The method for processing a human object in an image provided by an embodiment of the present application can be applied to, for example, Figure 1 the application environment shown. The terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or on other network servers.

[0068] It should be noted that Figure 1 this is only one application environment of the method for processing a human object, and it can also be applied to an application environment that only includes the terminal 102 or the server 104.

[0069] The terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 can be an independent physical server, or can be a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server that provides cloud computing services.

[0070] In an exemplary embodiment, as Figure 2 shown, a method for processing a human object in an image is provided. Taking the method applied to the Figure 1 terminal in it as an example for description, it includes the following steps 202 to step 210, where:

[0071] Step 202, extract a feature vector from a human image containing a human object to be processed.

[0072] Among them, the person to be processed can be the person in the person image who needs to be processed as a person object, such as the person in the person image who needs to be processed for body shaping. The person to be processed in the person image can present all body parts, or only the upper body or the body part below the head.

[0073] The person image can be a photo taken in real time using a camera application. During the shooting process, the body shaping method of the person object processing method of the present application is used to finally obtain a target person image after body shaping (that is, the photo that the user finally sees when taking a photo). The person in the target person image is a person after body shaping (abbreviated as a body-shaped person).

[0074] In addition, the person image can also be a photo stored in the album (including photos sent by friends, photos obtained by the user from the network, or photos taken by the user himself), or a photo not stored in the album. By triggering the body shaping tool provided in the album application, camera application, or image processing application, the person object processing method of the present application can be called to perform body shaping, and finally a target person image after shaping can be obtained.

[0075] In one embodiment, the terminal can display a person image including the person to be processed through the display page of the album application, camera application, or image processing application. The display page includes a control of the body shaping tool. In response to the triggering operation on the control, the person object processing method corresponding to the body shaping tool is called, so that the person to be processed in the person image can be body-shaped.

[0076] For example, the user can use an image processing application that provides a body shaping tool. After selecting one of the person images, the body shaping tool is triggered, so that the person object processing method of the present application can be used to body-shape the person to be processed in the person image.

[0077] It should be noted that the person object processing method of the present application can be integrated not only in the album application, camera application, or image processing application, but also in other application programs, such as social applications, video applications, takeaway applications, sharing applications, shopping applications, map applications, or taxi applications. When obtaining a person image using the chat function in these application programs, or when viewing or downloading a person image using these application programs, by triggering the body shaping tool, the person object processing method can be executed on the obtained person image, so as to realize body shaping of the person to be processed in the person image.

[0078] In one embodiment, the terminal can first preprocess the person image including the person to be processed. The preprocessing can include at least one of denoising processing, person beautification processing, or effective area cropping processing.

[0079] Among them, the portrait beauty processing can include adding filters, whitening processing, skin smoothing, etc.

[0080] In one embodiment, the terminal can extract a first feature map from a portrait image including a person to be processed; process the first feature map through a Resnet network to obtain a second feature map; and flatten the second feature map into a feature vector.

[0081] Among them, the Resnet network can extract a second feature map containing rich semantic information. The Resnet network outputs a feature map of a fixed size, and after flattening this feature map, the result of adding it to the position encoding vector is used as the input to the Transformer model.

[0082] The first feature map can be obtained by extraction through a feature extraction network. The feature extraction network can be an HRNet network, such as the HRNet32 network. For the HRNet network, it mainly consists of the following parts:

[0083] Stem module: The HRNet network uses a Stem module to reduce the resolution of the input image and increase the number of channels. It usually contains one or two convolutional layers, and the output resolution is 1 / 4 of the original input image;

[0084] Stage module: Each Stage module contains multiple parallel branches. These branches process feature maps of different resolutions. The network starts from a high resolution and continuously introduces new branches in each Stage. Each branch is responsible for processing feature maps with different downsampling rates. The feature maps perform convolutional operations within each branch, and information exchange is carried out between branches through bilinear interpolation or convolutional fusion;

[0085] Feature fusion: At the end of each Stage module, the features of all branches are integrated through cross-branch feature fusion. This fusion method ensures that the high-resolution feature map can obtain more context information from the low-resolution branches while maintaining the resolution;

[0086] Model output: The output of the HRNet network is the fused multi-scale feature map, which is then restored to the resolution size of the original input image through an upsampling operation.

[0087] In one embodiment, the terminal performs feature extraction on a portrait image including a person to be processed through a feature extraction network, and during the feature extraction process, multiple information exchanges are carried out between the parallel sub-networks in the feature extraction network to obtain the first feature map.

[0088] Among them, the sub-networks in the feature extraction network can be parallel network sub-networks in the feature extraction network (i.e., parallel sub-networks). Different sub-networks can process feature maps of different resolutions. The feature extraction network can be a neural network for feature extraction in the body beautification model. For example, it can be Figure 3 the HRNet network in

[0089] In the process of feature extraction, starting from the high-resolution sub-network as the first stage, gradually adding high-resolution to low-resolution sub-networks to form multiple stages, and connecting these multi-resolution sub-networks in parallel. Throughout the process, information is repeatedly exchanged between the parallel multi-resolution sub-networks to achieve multi-scale repeated fusion.

[0090] Specifically, the terminal inputs a person image containing the person to be processed into the feature extraction network; the feature extraction network includes parallel first, second, and third sub-networks; the feature extraction network extracts features from the person image, and during the feature extraction process, the first, second, and third sub-networks respectively process feature maps of different resolutions, and during the processing, information exchange across resolutions occurs between different sub-networks to obtain a first feature map.

[0091] Among them, the first sub-network can be a high-resolution sub-network of the feature extraction network, that is, a sub-network for processing high-resolution feature maps; the second sub-network can be a low-resolution (which can be called the first low-resolution) sub-network of the feature extraction network, that is, a sub-network for processing low-resolution feature maps; the third sub-network can be a lower-resolution (which can be called the second low-resolution) sub-network of the feature extraction network, that is, a sub-network for processing lower-resolution feature maps. Thus, it can be seen that the first low-resolution is higher than the second low-resolution.

[0092] For example, as Figure 4 shown, input the person image into the feature extraction network. The high-resolution sub-network of the feature extraction network (such as Figure 4The branches in the upper part) first extract features from the person image. When reaching the Stage module, a new branch (such as a low-resolution sub-network) is introduced. The low-resolution sub-network first downsamples the feature map to obtain a low-resolution feature map, and then processes this low-resolution feature map. During the processing, information exchange can be carried out across branches with the high-resolution sub-network; in addition, when reaching the low-resolution Stage module, a new branch (such as an even lower-resolution sub-network) can also be introduced. The even lower-resolution sub-network first downsamples the feature map to obtain an even lower-resolution feature map, and then processes this even lower-resolution feature map. During the processing, information exchange can be carried out across branches with the low-resolution sub-network and / or the high-resolution sub-network.

[0093] Step 204: Incorporate the position encoding vector into the feature vector to obtain the target feature vector.

[0094] Among them, the position encoding vector can be a vector composed of all position encodings, and the dimension size of this position encoding vector is the same as the dimension size of the feature vector.

[0095] In one embodiment, the terminal can first determine the position encoding vector; add the elements in the position encoding vector to the elements in the feature vector respectively to obtain the target feature vector.

[0096] Step 206: Encode the target feature vector to obtain an encoded feature vector with context relationship; encoding the target feature vector includes, during the encoding process, performing self-attention processing on the target feature vector using the self-attention mechanism to capture the context relationship.

[0097] Among them, the context relationship can be the dependency relationship between the elements in the target feature vector, which is used to reflect the internal association relationship between the pixels or pixel regions in the person image.

[0098] In one embodiment, the terminal can encode the target feature vector through the encoder of the Transformer model to obtain an encoded feature vector with context relationship.

[0099] Among them, the encoder of the Transformer model can refer to Figure 3 , and this encoder can include a multi-head self-attention layer, a feed-forward neural network layer, as well as residual connections and layer normalization.

[0100] In addition, during the encoding process, each layer of the encoder associates all the pixel points in the feature map with each other, thereby generating an enhanced feature sequence (i.e., the encoded feature vector).

[0101] In one embodiment, the terminal may input the target feature vector into the encoder of the Transformer model; for each head in the multi-head self-attention layer, the self-attention mechanism is respectively used to linearly process the target feature vector to obtain a query vector, a key vector, and a value vector, calculate the attention score according to the query vector and the key vector, perform normalization processing on the attention score, and obtain context information based on the normalized attention score and the value vector; perform a non-linear transformation on the context information through the feed-forward neural network layer to obtain non-linearly transformed features; through residual connection and layer normalization, fuse the non-linearly transformed features with the target feature vector, and perform normalization processing on the obtained fused features to obtain an encoded feature vector with context relationship.

[0102] Among them, in order to retain the spatial position information, a fixed position encoding is superimposed on the encoded feature vector input to the encoder, and the position encoding can be generated by sine and cosine functions. The context information can be information with context relationship.

[0103] Through the encoder, the features at each position (pixel point) can exchange information with the features at other positions, and this global dependency enables the model to capture complex context relationships in the image.

[0104] Step 208, decode the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and the stretching direction.

[0105] Among them, the stretching direction can also be the deformation direction, that is, the direction used to stretch the whole or part of the person to be processed in the person image, so that the whole or a certain part of the person to be processed is elongated or shortened, or enlarged or reduced proportionally.

[0106] In one embodiment, the terminal can decode the encoded feature vector through the decoder of the Transformer model to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and the stretching direction, as Figure 5 and Figure 6 shown, Figure 5 are the Mesh points of the buttocks of the person to be processed, where the points in the left figure are the predicted 2D Mesh points, and the gray points in the right figure correspond to the predicted 3D Mesh points; Figure 6 are the Mesh points of the abdomen of the person to be processed, where the points in the left figure are the predicted 2D Mesh points, and the gray points in the right figure correspond to the predicted 3D Mesh points.

[0107] Specifically, the terminal can input the encoded feature vector into the decoder of the Transformer model; through the decoder, cross-attention processing is performed on the encoded feature vector and the target query vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and the stretching direction.

[0108] Among them, the decoder part receives the feature sequence output from the encoder and predicts the position information and category of the target through the cross-attention mechanism.

[0109] The cross-attention processing can be a way of performing attention processing using the cross-attention mechanism. Under the cross-attention mechanism, at each layer of the decoder, the query vector will interact with the encoded feature vector output by the encoder.

[0110] For the query vector in the decoder, it can be a set of learning query vectors with a fixed number. These vectors serve as the implicit representation of the target points. After passing through the decoder, the query vectors perform cross-attention operations with the encoded feature vectors output by the encoder, thereby generating the final prediction result.

[0111] Step 210, based on the 2D Mesh point coordinates, 3D Mesh point coordinates, and the stretching direction, generate a figure with a beautiful body shape corresponding to the person to be processed.

[0112] In one embodiment, the terminal can use camera projection technology to generate a figure with a beautiful body shape corresponding to the person to be processed based on the 2D Mesh point coordinates, 3D Mesh point coordinates, and the stretching direction.

[0113] In the above embodiment, a feature vector is extracted from the person image containing the person to be processed; a position encoding vector is incorporated into the feature vector to obtain a target feature vector; the target feature vector is encoded to obtain an encoded feature vector with context relationships; encoding the training target feature vector includes, during the encoding process, performing self-attention processing on the target feature vector using the self-attention mechanism to capture context relationships; decoding the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and the stretching direction, thereby introducing three-dimensional technology into the body shaping of the person object, accurately capturing the three-dimensional space information of the human body. Therefore, even in the case of complex postures and occlusions, the 2D Mesh point coordinates, 3D Mesh point coordinates, and the stretching direction can be accurately predicted, and based on the 2D Mesh point coordinates, 3D Mesh point coordinates, and the stretching direction, a figure with a beautiful body shape corresponding to the person to be processed is generated, thereby effectively improving the body shaping effect of the person object.

[0114] In one embodiment, as Figure 7 shown, the method further includes:

[0115] S702. Extract the training feature vectors from the person image samples containing the target person.

[0116] S704. Incorporate the position encoding vectors into the training feature vectors to obtain the training target feature vectors.

[0117] S706. Encode the training target feature vectors through the encoder of the initial Transformer model to obtain the training encoded feature vectors with context relationships; encoding the training target feature vectors includes, during the encoding process, performing self-attention processing on the training feature vectors using the self-attention mechanism to capture context relationships.

[0118] S708. Decode the training encoded feature vectors through the decoder of the initial Transformer model to obtain the predicted 2D Mesh point coordinates, predicted 3D Mesh point coordinates, and predicted stretching directions.

[0119] S710. Determine the loss value based on the predicted 2D Mesh point coordinates, predicted 3D Mesh point coordinates, predicted stretching directions, and the corresponding labels.

[0120] Among them, the labels are the 2D Mesh point labels, 3D Mesh point labels, and stretching direction labels obtained by processing according to the SMPL human model.

[0121] Among them, SMPL (Skinned Multi-Person Linear Model) is a parametric model for 3D human body modeling. Through linear blend skinning (LBS) and shape-pose separation, it realizes accurate modeling of the human body shape and pose.

[0122] In this application, the SMPL human model is used as the human parameter model. This SMPL human model has 6890 vertices, that is , The dimension of is 10, that is , and its vertex coordinates are expressed as a function of shape parameters and pose parameters , specifically as follows:

[0123]

[0124] Among them, is the shape parameter, is the pose parameter, is the skinning weight, the function is the linear blend skinning function, is the template vertex after shape and pose deformation, is the joint position.

[0125] Perform forward calculation through shape parameters and pose parameters to output Mesh points, then perform normalization based on pelvic points, and cut human key points through predefined indexes to find the Mesh point information of the corresponding human parts, which is used as the label (GT) for real training. In addition, the stretching direction is also obtained as the label for real training.

[0126] S712. Optimize the initial Transformer model according to the loss value until the model converges.

[0127] Among them, for the specific implementation process of S702 - S710, reference can be made to S202 - S210, which will not be elaborated here.

[0128] In the above embodiments, training feature vectors are extracted from the human image samples containing the target person; position encoding vectors are incorporated into the training feature vectors to obtain training target feature vectors; the training target feature vectors are encoded by the encoder of the initial Transformer model to obtain training encoded feature vectors with context relationships. During the encoding process, the self - attention mechanism is used to perform self - attention processing on the training feature vectors to capture context relationships; the training encoded feature vectors are decoded by the decoder of the initial Transformer model to obtain predicted 2D Mesh point coordinates, predicted 3D Mesh point coordinates, and predicted stretching directions; the loss value is determined according to the predicted 2D Mesh point coordinates, predicted 3D Mesh point coordinates, predicted stretching directions, and the corresponding labels. Among them, the labels are 2D Mesh point labels, 3D Mesh point labels, and stretching direction labels obtained by processing according to the SMPL human model; the initial Transformer model is optimized according to the loss value until the model converges, so that model optimization can be achieved and the accuracy of model prediction can be improved. Therefore, the optimized model can accurately predict the 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching directions, and then generate a body - beautified person corresponding to the person to be processed, effectively improving the body - beautifying effect of the person object.

[0129] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0130] Based on the same inventive concept, an embodiment of the present application further provides a processing device for a human object in an image for implementing the above-mentioned human object processing method in the image. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the processing device for a human object in an image provided below can refer to the limitations on the human object processing method in the above text, and will not be repeated here.

[0131] In an exemplary embodiment, as Figure 8 shown, a processing device for a human object in an image is provided, including: an extraction module 802, a fusion module 804, an encoding module 806, a decoding module 808, and a generation module 810, where:

[0132] The extraction module 802 is configured to extract a feature vector from a human image containing a person to be processed;

[0133] The fusion module 804 is configured to integrate a position encoding vector into the feature vector to obtain a target feature vector;

[0134] The encoding module 806 is configured to encode the target feature vector to obtain an encoded feature vector with context relationship; encoding the target feature vector includes performing self-attention processing on the target feature vector by using a self-attention mechanism to capture the context relationship;

[0135] The decoding module 808 is configured to decode the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and a stretching direction;

[0136] The generation module 810 is configured to generate a body-beautiful person corresponding to the person to be processed based on the 2D Mesh point coordinates, 3D Mesh point coordinates, and the stretching direction.

[0137] In one embodiment, the extraction module 802 is further configured to extract a first feature map from a person image including a person to be processed; process the first feature map through a ResNet network to obtain a second feature map; and flatten the second feature map into a feature vector.

[0138] In one embodiment, the extraction module 802 is further configured to perform feature extraction on a person image including a person to be processed through a feature extraction network, and during the feature extraction process, perform multiple information exchanges between parallel sub-networks in the feature extraction network to obtain a first feature map.

[0139] In one embodiment, the extraction module 802 is further configured to input a person image including a person to be processed into a feature extraction network; the feature extraction network includes a first sub-network, a second sub-network, and a third sub-network in parallel; perform feature extraction on the person image through the feature extraction network, and during the feature extraction process, the first sub-network, the second sub-network, and the third sub-network respectively process feature maps with different resolutions, and during the processing process, perform cross-resolution information exchanges between different sub-networks to obtain a first feature map.

[0140] In one embodiment, the target feature vector is obtained by encoding through the encoder of the Transformer model, and the encoder includes a multi-head self-attention layer, a feed-forward neural network layer, as well as residual connections and layer normalization;

[0141] The encoding module 806 is further configured to input the target feature vector into the encoder of the Transformer model; for each head in the multi-head self-attention layer, respectively perform linear processing on the target feature vector through the self-attention mechanism to obtain a query vector, a key vector, and a value vector, calculate an attention score according to the query vector and the key vector, perform normalization processing on the attention score, and obtain context information based on the normalized attention score and the value vector; perform a non-linear transformation on the context information through the feed-forward neural network layer to obtain a non-linearly transformed feature; through residual connections and layer normalization, fuse the non-linearly transformed feature with the target feature vector, and perform normalization processing on the obtained fused feature to obtain an encoded feature vector with context relationships.

[0142] In one embodiment, the decoding module 808 is further configured to input the encoded feature vector into the decoder of the Transformer model;

[0143] Perform cross-attention processing on the encoded feature vector and the target query vector through the decoder to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and a stretching direction.

[0144] In one embodiment, as Figure 9 shown, the device further includes:

[0145] A training module 812 is configured to extract training feature vectors from person image samples including a target person; integrate position encoding vectors into the training feature vectors to obtain training target feature vectors; encode the training target feature vectors through an encoder of an initial Transformer model to obtain training encoded feature vectors with context relationships; encode the training target feature vectors, including performing self-attention processing on the training feature vectors by using a self-attention mechanism during the encoding process to capture context relationships; decode the training encoded feature vectors through a decoder of the initial Transformer model to obtain predicted 2D Mesh point coordinates, predicted 3D Mesh point coordinates, and predicted stretching directions; determine a loss value according to the predicted 2D Mesh point coordinates, the predicted 3D Mesh point coordinates, the predicted stretching directions, and corresponding labels, where the labels are 2D Mesh point labels, 3D Mesh point labels, and stretching direction labels obtained by processing according to the SMPL human body model; optimize the initial Transformer model according to the loss value until the model converges.

[0146] In the above embodiment, feature vectors are extracted from a person image including a person to be processed; position encoding vectors are integrated into the feature vectors to obtain target feature vectors; the target feature vectors are encoded to obtain encoded feature vectors with context relationships, and during the encoding process, self-attention processing is performed on the target feature vectors by using a self-attention mechanism to capture context relationships; the encoded feature vectors are decoded to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching directions, so as to introduce three-dimensional technology into the body shaping of a person object, accurately capture the three-dimensional spatial information of the human body. Therefore, even in the case of complex postures and occlusions, the 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching directions can be accurately predicted, and based on the 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching directions, a body-shaped person corresponding to the person to be processed is generated, thereby effectively improving the body shaping effect of the person object.

[0147] Each module in the above person object processing device in the image can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.

[0148] In an exemplary embodiment, a computer device is provided. The computer device may be a server and includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store image data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for processing human object in an image.

[0149] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as Figure 10 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (Near Field Communication, NFC), or other technologies. When the computer program is executed by the processor, it implements a method for processing human object in an image. The display unit of the computer device is used to form a visually visible picture, which may be a display screen, a projection device, or a virtual reality imaging device. The display screen may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.

[0150] Those skilled in the art can understand that Figure 10The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0151] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of processing the human object in the above image are implemented.

[0152] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of processing the human object in the above image are implemented.

[0153] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of processing the human object in the above image are implemented.

[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0155] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0156] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded in this application.

[0157] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A method for processing a human object in an image, characterized in that, The method includes: extracting a feature vector from a person image including a person to be processed; incorporating a positional encoding vector into the feature vector to obtain a target feature vector; encoding the target feature vector to obtain an encoded feature vector with context relationships; the encoding of the target feature vector includes, during the encoding process, performing self-attention processing on the target feature vector using a self-attention mechanism to capture the context relationships; decoding the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and a stretching direction; generating a body-beautiful person corresponding to the person to be processed based on the 2D Mesh point coordinates, the 3D Mesh point coordinates, and the stretching direction.

2. The method according to claim 1, wherein The extracting a feature vector from a person image including a person to be processed includes: extracting a first feature map from a person image including a person to be processed; processing the first feature map through a ResNet network to obtain a second feature map; flattening the second feature map into a feature vector.

3. The method according to claim 2, wherein The extracting a first feature map from a person image including a person to be processed includes: performing feature extraction on a person image including a person to be processed through a feature extraction network, and during the feature extraction process, performing multiple information exchanges between parallel sub-networks in the feature extraction network to obtain a first feature map.

4. The method according to claim 3, characterized in that, The performing feature extraction on a person image including a person to be processed through a feature extraction network, and during the feature extraction process, performing multiple information exchanges between parallel sub-networks in the feature extraction network to obtain a first feature map includes: inputting a person image including a person to be processed into the feature extraction network; the feature extraction network includes parallel first, second, and third sub-networks; performing feature extraction on the person image through the feature extraction network, and during the feature extraction process, the first, second, and third sub-networks respectively process feature maps of different resolutions, and during the processing, cross-resolution information exchanges are performed between different sub-networks to obtain a first feature map.

5. The method according to claim 1, characterized in that, The target feature vector is obtained by encoding through the encoder of a Transformer model, and the encoder includes a multi-head self-attention layer, a feed-forward neural network layer, as well as residual connections and layer normalization; The encoding the target feature vector to obtain an encoded feature vector with context relationships includes: inputting the target feature vector into the encoder of the Transformer model; for each head in the multi-head self-attention layer, respectively performing linear processing on the target feature vector using a self-attention mechanism to obtain a query vector, a key vector, and a value vector, calculating an attention score based on the query vector and the key vector, performing normalization processing on the attention score, and obtaining context information based on the normalized attention score and the value vector; performing a non-linear transformation on the context information through the feed-forward neural network layer to obtain a non-linearly transformed feature; Through the residual connection and layer normalization, the non-linear transformation features are fused with the target feature vector, and the obtained fused features are normalized to obtain an encoded feature vector with context relationship.

6. The method according to any one of claims 1 to 5, characterized in that The decoding of the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching directions includes: Inputting the encoded feature vector into the decoder of the Transformer model; Through the decoder, performing cross-attention processing on the encoded feature vector and the target query vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching directions.

7. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Extracting a training feature vector from a person image sample containing a target person; Incorporating a position encoding vector into the training feature vector to obtain a training target feature vector; Encoding the training target feature vector through the encoder of the initial Transformer model to obtain a training encoded feature vector with context relationship; the encoding of the training target feature vector includes, during the encoding process, performing self-attention processing on the training feature vector using the self-attention mechanism to capture the context relationship; Decoding the training encoded feature vector through the decoder of the initial Transformer model to obtain predicted 2D Mesh point coordinates, predicted 3D Mesh point coordinates, and predicted stretching directions; Determining a loss value according to the predicted 2D Mesh point coordinates, the predicted 3D Mesh point coordinates, the predicted stretching directions, and the corresponding labels; wherein, the labels are 2D Mesh point labels, 3D Mesh point labels, and stretching direction labels obtained by processing according to the SMPL human model; Optimizing the initial Transformer model according to the loss value until the model converges.

8. An apparatus for processing a human object in an image, characterized in that The device includes: An extraction module for extracting a feature vector from a person image containing a person to be processed; A fusion module for incorporating a position encoding vector into the feature vector to obtain a target feature vector; An encoding module for encoding the target feature vector to obtain an encoded feature vector with context relationship; the encoding of the target feature vector includes, during the encoding process, performing self-attention processing on the target feature vector using the self-attention mechanism to capture the context relationship; A decoding module for decoding the encoded feature vector to obtain 2D Mesh point coordinates, 3D Mesh point coordinates, and stretching directions; A generation module for generating a body-beautiful person corresponding to the person to be processed based on the 2D Mesh point coordinates, the 3D Mesh point coordinates, and the stretching directions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.