Human body analysis method and system based on Transformer

By using Transformer-based human analysis method in human body analysis, DeepLabV3+ and Transformer decoder to calculate the attention relationship between human body parts, the problems of boundary confusion, low accuracy and computational complexity in the existing technology are solved, and efficient and accurate real-time human body analysis is achieved.

CN115063586BActive Publication Date: 2025-05-23SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210664690.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-14
Publication Date
2025-05-23
Estimated Expiration
2042-06-14

AI Technical Summary

Technical Problem

The prior art has boundary confusion problems in human body analysis, the convolutional neural network is limited by the receptive field, and the use of human body prior information increases computational complexity, making it difficult to achieve real-time analysis.

Method used

The human body analysis method based on Transformer is adopted to extract multi-scale resolution image features through the DeepLabV3+ backbone network, and combine the Transformer decoder to calculate the attention relationship between human body parts to generate high-precision human body analysis results.

Benefits of technology

It improves the accuracy of human body analysis results, reduces computational complexity and time consumption, and achieves efficient real-time human body analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063586B_ABST
    Figure CN115063586B_ABST
Patent Text Reader

Abstract

The present invention discloses a human body parsing method based on Transformer. The method comprises: using DeepLabV3+ backbone network to extract low-resolution image features of pictures, performing upsampling operation on low-resolution image features to obtain multi-scale resolution image features, performing pixel-level decoding operation and Transformer decoder on different scale image features to obtain pixel-level embedding and embedding features respectively, performing inner product of pixel-level embedding and embedding features to obtain semantic segmentation map, and fusing semantic segmentation map to obtain human body parsing map. The present invention also discloses a human body parsing system based on Transformer. The present invention adopts a Transformer-based approach, and the attention mechanism of Transformer can model long-term dependencies and effectively capture global features, thereby improving the accuracy of human body parsing. Moreover, the present invention does not introduce prior information of human body, has fast calculation speed, low model complexity, and can perform real-time human body parsing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer vision, image processing and human body analysis, and in particular to a Transformer-based human body analysis method and system. Background Art

[0002] With the rapid development of machine learning technology in recent years, the technology in the field of computer vision and image processing has also developed rapidly and has been successfully applied in many fields. Images are everywhere in our daily lives and are an important tool for us to understand the world. By using machine learning and image processing technology, we can extract very diverse and meaningful information from images. The combination of machine learning technology and image processing technology has also given birth to many important applications, such as: face recognition, medical image analysis, human body detection, etc. Human-centered images account for a large proportion of many images, so the analysis and understanding of human semantics, that is, human body analysis, has great scientific research and commercial value.

[0003] Human body parsing, also known as human body semantic segmentation, is a hot topic in the field of computer vision research. It uses semantic segmentation technology to perform pixel-level classification operations on images or videos centered on the human body. Human body parsing divides the person captured in the image into different fine-grained semantic parts, such as head, torso, arms and legs. As a finer-grained semantic segmentation task, it is more challenging than finding the human body contour. There are also many applications in real life, such as: segmenting the clothes in the image with pixel-level accuracy through clothing parsing to improve the accuracy of clothing recommendation and corresponding retrieval algorithms; realizing pedestrian re-identification by extracting human features from images, etc.

[0004] With the development of deep learning technology, the research on human body analysis technology has made great progress and breakthroughs, and many models using convolutional neural networks have achieved certain results. However, due to the different sizes, lighting intensities, occluded parts of the characters in the images, the types and matching of clothing, the diversity of human postures and image backgrounds, etc., simple convolutional neural networks cannot achieve the purpose of human body analysis well. In addition, in terms of detail resolution of the human body, such as the depiction of finger contours, the recognition of jewelry, and the distinction between similar clothing, the prediction accuracy is still relatively low. At the same time, when the background in the picture is special and complex, the model will find it difficult to distinguish the foreground and background.

[0005] One of the current existing technologies is a graph learning-based method, which extracts rough human body parsing results through convolutional neural networks, embeds human body parts in the image as graph features, uses graph neural networks and attention mechanisms to infer graph features to obtain semantic connections between human body parts, and uses semantic connections between human body parts to refine the rough human body parsing results to obtain the final human body parsing graph. The disadvantage of this solution is that it cannot solve the problem of boundary confusion in the human body parsing graph, and the reasoning of graph features consumes a lot of computing time.

[0006] The second existing technology is a method based on a fully convolutional neural network. This method uses a convolutional neural network, such as resnet and DeepLab, to extract human features in an image, and then uses deconvolution to upsample and generate a human parsing map. The entire training process is supervised by a cross-entropy loss function. The disadvantage of this method is that the full use of a convolutional neural network is limited by the receptive field, and a single-layer convolution cannot capture long-distance features, so the final human parsing image is not accurate.

[0007] The third existing technology uses the prior information of the human body to assist in human body analysis. This method uses a convolutional network to extract a rough human body analysis map, and then uses a neural network to extract the edge map and posture map of the human body, and refines the rough human body analysis map through the edge map and posture map. The disadvantage of this method is that the use of prior information of the human body will introduce additional calculations, increase the complexity of the program and the complexity of the calculation, and is difficult to use in real-time analysis situations. Summary of the invention

[0008] The purpose of the present invention is to overcome the shortcomings of existing methods and propose a human body parsing method and system based on Transformer. The main problems solved by the present invention are: first, the existing graph learning-based methods cannot solve the boundary confusion problem in the human body parsing graph. Second, the method of the full convolutional neural network, the convolutional neural network is limited by the receptive field, so that the accuracy of the human body parsing image is not high. Third, the use of human body prior information to assist in human body parsing will introduce additional calculations, making it difficult to achieve real-time parsing scenarios.

[0009] In order to solve the above problems, the present invention proposes a human body parsing method based on Transformer, which includes:

[0010] Input human body pictures and parsed pictures, perform data augmentation on the input pictures, and process them into a uniform size;

[0011] Using the DeepLabV3+ backbone network, extract low-resolution image features of the image, perform multiple upsampling operations on the low-resolution image features, and obtain image features with resolutions of four scales: 1 / 32, 1 / 16, 1 / 8, and 1 / 4;

[0012] Input the 1 / 4 scale resolution image features into the pixel-level decoder for operation to obtain pixel-level embedding;

[0013] Input the resolution image features of 1 / 32, 1 / 16, and 1 / 8 scales and the number of N queries, use the Transformer decoder to calculate the attention relationship between different human body parts from the resolution image features of different scales, and obtain C*N dimensional embedding features, where C is the number of channels and N is the number of human body parts and clothing types to be segmented;

[0014] Perform an inner product of the embedding feature and the pixel-level embedding to obtain N H*W-dimensional binary semantic segmentation maps, where H and W represent the height and width of the binary semantic segmentation map, respectively. Each map represents the parsing result of a human body part, and the pixel value of the corresponding part is represented by 1, and the pixel values ​​of other areas are represented by 0;

[0015] N binary semantic segmentation maps are fused, that is, the 1 in the binary map is replaced with the label of the corresponding segmentation part, and the N replaced maps are added together to obtain the final human body analysis result.

[0016] Preferably, the input human body picture and the parsed picture perform data enhancement on the input picture and process it into a uniform size, specifically as follows:

[0017] Human body pictures and human body analysis pictures are input. Human body pictures are collected from the Internet, and analysis pictures are pictures of different parts of the human body and clothes marked with different colors manually. In order to make the trained model robust, the pictures are randomly rotated, horizontally mirrored, and randomly cropped for data enhancement. Finally, all pictures are scaled to a uniform size.

[0018] Preferably, the DeepLabV3+ backbone network is used to extract low-resolution image features of the picture, and multiple upsampling operations are performed on the low-resolution image features to obtain image features with resolutions of four scales: 1 / 32, 1 / 16, 1 / 8 and 1 / 4, specifically:

[0019] The image is input into the DeepLabV3+ backbone network, and the DeepLabV3+ backbone network performs 1*1 convolution, 3*3 dilated convolution with ratios of 6, 12 and 18 and image pooling operations on the image to obtain 5 feature maps. The 5 feature maps are cascaded and subjected to 1*1 convolution operations to obtain low-resolution image features. The low-resolution image features are upsampled to obtain image features with four scales of resolution: 1 / 32, 1 / 16, 1 / 8, and 1 / 4.

[0020] Preferably, the step of inputting the 1 / 4 scale resolution image features into a pixel-level decoder for operation to obtain pixel-level embedding is specifically as follows:

[0021] A pixel-level decoder is used to perform 1*1 convolution on the 1 / 4 scale resolution image features and the original image, and the 1 / 4 scale resolution image features are connected. Subsequently, upsampling operations are continuously performed through deconvolution operations to obtain multi-resolution image features of different scales.

[0022] Preferably, the input is the resolution image features of 1 / 32, 1 / 16, 1 / 8 scales and N query numbers, and the Transformer decoder is used to calculate the attention relationship between different human body parts from the resolution image features of different scales to obtain C*N dimensional embedding features, where C is the number of channels and N is the number of human body parts and clothing types to be segmented, specifically:

[0023] Input the 1 / 32, 1 / 16, 1 / 8 scale resolution image features and N queries to the Transformer decoder, where N is the type of human body parts and clothing to be segmented. First, calculate the cross attention:

[0024] X l =softmax(Q l K l )V l +X l-1

[0025] Where l is the subscript of the number of layers, X l is the query feature of the lth layer, Q l is the query input to layer l, V l and K l The image features of the first layer input are transformed by two different linear transformation functions f V and f K The transformed matrix is ​​then normalized for the cross-attention result and passed through a self-attention layer. The result calculated by the self-attention layer will be output as the final query feature after normalization through the feed-forward layer.

[0026] The Transformer decoder decodes the image features at three scales: 1 / 8, 1 / 16, and 1 / 32. The three decoding operations are repeated L times, i.e., a total of 3L decodings are performed. The decoded image will pass through a multi-layer perceptron to generate a C*N-dimensional embedding feature, where C is the number of channels and N is the number of human body parts and clothing to be segmented.

[0027] Accordingly, the present invention also provides a Transformer-based human body analysis system, comprising:

[0028] An image preprocessing unit, used to input human body images and parsed images, perform data enhancement on the input images, and process them into a uniform size;

[0029] A multi-scale resolution image feature unit, for extracting low-resolution image features of the image using the DeepLabV3+ backbone network, performing multiple upsampling operations on the low-resolution image features, and obtaining image features with resolutions of four scales: 1 / 32, 1 / 16, 1 / 8, and 1 / 4;

[0030] A pixel-level decoding unit, used to input the 1 / 4 scale resolution image features into a pixel-level decoder for operation to obtain pixel-level embedding;

[0031] A Transformer decoding unit is used to input the resolution image features of 1 / 32, 1 / 16, and 1 / 8 scales and N query numbers, and use the Transformer decoder to calculate the attention relationship between different human body parts from the resolution image features of different scales to obtain C*N dimensional embedding features, where C is the number of channels and N is the number of human body parts and clothing types to be segmented;

[0032] A semantic segmentation map acquisition unit is used to perform inner product on the embedded features and the pixel-level embedding to obtain N H*W-dimensional binary semantic segmentation maps, each of which represents the parsing result of a human body part, the pixel value of the corresponding part is represented by 1, and the pixel values ​​of other areas are represented by 0, and H and W represent the height and width of the binary semantic segmentation map respectively;

[0033] N binary semantic segmentation maps are fused, that is, the 1 in the binary map is replaced with the label of the corresponding segmentation part, and the N replaced maps are added together to obtain the final human body analysis result.

[0034] The implementation of the present invention has the following beneficial effects:

[0035] The present invention does not rely on any additional input data and prior information about the human body. Compared with the method of using prior information about the human body to assist in human body analysis, the present invention has the advantages of fast calculation speed and low model complexity. The present invention cleverly combines Transformer with convolutional neural network, fully exerts the ability of attention mechanism to capture global features, and maximizes the accuracy of human body analysis results. The input and output of each part of the present invention are interconnected to extract and integrate different features, which improves efficiency and makes the generated human body results more in line with people's expectations. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is an overall flow chart of a human body analysis method based on Transformer according to an embodiment of the present invention;

[0037] Figure 2 is a flow chart of a Transformer decoder according to an embodiment of the present invention;

[0038] Figure 3 It is a structural diagram of a Transformer-based human body analysis system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] Figure 1 is an overall flow chart of the Transformer-based human body analysis method according to an embodiment of the present invention, such as Figure 1 As shown, the method includes:

[0041] S1, input human body pictures and parsed pictures, perform data enhancement on the input pictures, and process them into a uniform size;

[0042] S2, using the DeepLabV3+ backbone network, extracting low-resolution image features of the image, performing multiple upsampling operations on the low-resolution image features, and obtaining image features with resolutions of four scales: 1 / 32, 1 / 16, 1 / 8, and 1 / 4;

[0043] S3, input the 1 / 4 scale resolution image features into the pixel-level decoder for operation to obtain pixel-level embedding;

[0044] S4, input the resolution image features of 1 / 32, 1 / 16, and 1 / 8 scales and the number of N queries, use the Transformer decoder to calculate the attention relationship between different human body parts from the resolution image features of different scales, and obtain C*N dimensional embedding features, where C is the number of channels and N is the number of human body parts and clothing types to be segmented;

[0045] S5, performing an inner product of the embedded feature and the pixel-level embedding to obtain N H*W-dimensional binary semantic segmentation maps, each of which represents the parsing result of a human body part, the pixel value of the corresponding part is represented by 1, and the pixel values ​​of other areas are represented by 0, and H and W respectively represent the height and width of the binary semantic segmentation map;

[0046] S6, fuse N binary semantic segmentation maps, that is, replace the 1 in the binary map with the label of the corresponding segmentation part, add the N replaced maps to obtain the final human body analysis result, and recommend clothing products of the same category or complementary clothing products of different categories that are similar to the current preference to the user

[0047] Step S1 is as follows:

[0048] S1-1, input human body pictures and human body analysis pictures. Human body pictures are collected from the Internet, and analysis pictures are pictures of different parts of the human body and clothes manually labeled with different colors. In order to make the trained model robust, the pictures are randomly rotated, horizontally mirrored, and randomly cropped for data enhancement. Finally, all pictures are scaled to a uniform size.

[0049] Step S2 is as follows:

[0050] S2-1, input the image into the DeepLabV3+ backbone network, the DeepLabV3+ backbone network performs 1*1 convolution, 3*3 dilated convolution with ratios of 6, 12 and 18 and image pooling operations on the image to obtain 5 feature maps, cascade the 5 feature maps, perform 1*1 convolution operations on them to obtain low-resolution image features, and upsample the low-resolution image features to obtain image features with four scales of resolution: 1 / 32, 1 / 16, 1 / 8, and 1 / 4.

[0051] Step S3 is as follows:

[0052] A pixel-level decoder is used to perform 1*1 convolution on the 1 / 4 scale resolution image features and the original image, and the 1 / 4 scale resolution image features are connected. Subsequently, upsampling operations are continuously performed through deconvolution operations to obtain multi-resolution image features of different scales.

[0053] Step S4, such as Figure 2 As shown, the details are as follows:

[0054] S4-1, input the 1 / 32, 1 / 16, 1 / 8 scale resolution image features and N queries into the Transformer decoder, where N is the type of human body parts and clothing to be segmented. First, calculate the cross attention:

[0055] X l =softmax(Q l K l )V l +X l-1

[0056] Where l is the subscript of the number of layers, X lis the query feature of the lth layer, Q l is the query input to layer l, V l and K l The image features of the first layer input are transformed by two different linear transformation functions f V and f K The transformed matrix is ​​then normalized for the cross-attention result and passed through a self-attention layer. The result calculated by the self-attention layer will be output as the final query feature after normalization through the feed-forward layer.

[0057] S4-2, the Transformer decoder decodes the image features with resolutions of 1 / 8, 1 / 16, and 1 / 32. The three decoding operations are repeated L times, that is, a total of 3L decodings are performed. The decoded image will pass through a multi-layer perceptron to generate a C*N-dimensional embedding feature, where C is the number of channels and N is the number of human body parts and clothing to be segmented.

[0058] Accordingly, the present invention also provides a human body analysis system based on Transformer, such as Figure 3 As shown, including:

[0059] The image preprocessing unit 1 is used to input a human body image and an analysis image, perform data enhancement on the input image, and process the image into a uniform size.

[0060] Specifically, human body pictures and human body analysis pictures are input. Human body pictures are collected from the Internet, and analysis pictures are pictures of different parts of the human body and clothes that are manually labeled with different colors. In order to make the trained model robust, the pictures are randomly rotated, horizontally mirrored, and randomly cropped for data enhancement. Finally, all pictures are scaled to a uniform size.

[0061] The multi-scale resolution image feature unit 2 is used to use the DeepLabV3+ backbone network to extract low-resolution image features of the image, perform multiple upsampling operations on the low-resolution image features, and obtain resolution image features of four scales: 1 / 32, 1 / 16, 1 / 8 and 1 / 4.

[0062] Specifically, the picture is input into the DeepLabV3+ backbone network, and the DeepLabV3+ backbone network performs 1*1 convolution, 3*3 dilated convolution with ratios of 6, 12 and 18 and image pooling operations on the picture to obtain 5 feature maps, and the 5 feature maps are cascaded and subjected to 1*1 convolution operations to obtain low-resolution image features, and the low-resolution image features are upsampled to obtain image features with four scales of resolution: 1 / 32, 1 / 16, 1 / 8, and 1 / 4.

[0063] The pixel level decoding unit 3 is used to input the 1 / 4 scale resolution image features into the pixel level decoder for operation to obtain pixel level embedding.

[0064] Specifically, a pixel-level decoder is used to perform 1*1 convolution on the 1 / 4 scale resolution image features and the original image, the 1 / 4 scale resolution image features are connected, and then upsampling operations are continuously performed through deconvolution operations to obtain multi-resolution image features of different scales.

[0065] The Transformer decoding unit 4 is used to input the resolution image features of 1 / 32, 1 / 16, and 1 / 8 scales and N query numbers, and use the Transformer decoder to calculate the attention relationship between different human body parts from the resolution image features of different scales to obtain C*N dimensional embedded features, where C is the number of channels and N is the number of human body parts and clothing types to be segmented.

[0066] Specifically, input the 1 / 32, 1 / 16, 1 / 8 scale resolution image features and N queries into the Transformer decoder, where N is the type of human body parts and clothing to be segmented. First, calculate the cross attention:

[0067] X l =softmax(Q l K l )V l +X l-1

[0068] Where l is the subscript of the number of layers, X l is the query feature of the lth layer, Q l is the query input to layer l, V l and K l The image features of the first layer input are transformed by two different linear transformation functions f V and f K The transformed matrix is ​​then normalized for the cross-attention result and passed through a self-attention layer. The result calculated by the self-attention layer will be output as the final query feature after normalization through the feed-forward layer.

[0069] The Transformer decoder decodes the image features at three scales: 1 / 8, 1 / 16, and 1 / 32. The three decoding operations are repeated L times, i.e., a total of 3L decodings are performed. The decoded image will pass through a multi-layer perceptron to generate a C*N-dimensional embedding feature, where C is the number of channels and N is the number of human body parts and clothing to be segmented.

[0070] The semantic segmentation map acquisition unit 5 is used to perform inner product on the embedded features and the pixel-level embedding to obtain N H*W dimensional binary semantic segmentation maps, each of which represents the analysis result of a human body part, the pixel value of the corresponding part is represented by 1, and the pixel values ​​of other areas are represented by 0, and H and W respectively represent the height and width of the binary semantic segmentation map.

[0071] The human body parsing map unit 6 is used to fuse N binary semantic segmentation maps, that is, to replace 1 in the binary map with the label of the corresponding segmentation part, and to add the N replaced maps to obtain the final human body parsing result.

[0072] Therefore, the present invention uses a human body parsing method based on Transformer, without the help of any additional input data and human body prior information. Compared with the method of using human body prior information to assist in human body parsing, it has the advantages of fast calculation speed and low model complexity; the present invention cleverly combines Transformer with convolutional neural network, fully exerts the ability of attention mechanism to capture global features, and maximizes the accuracy of human body parsing results; the input and output of each part in the present invention are interconnected, and different features are extracted and integrated, which improves efficiency and makes the generated human body results more in line with people's expectations.

[0073] The above is a detailed introduction to the Transformer-based human body analysis method and system provided in the embodiments of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A human body parsing method based on Transformer, It is characterized in that The method comprises: Input human body pictures and parsed pictures, perform data augmentation on the input pictures, and process them into a uniform size; Using the DeepLabV3+ backbone network, extract low-resolution image features of the image, perform multiple upsampling operations on the low-resolution image features, and obtain image features with resolutions of four scales: 1 / 32, 1 / 16, 1 / 8, and 1 / 4; Input the 1 / 4 scale resolution image features into the pixel-level decoder for operation to obtain pixel-level embedding; Input the resolution image features of 1 / 32, 1 / 16, and 1 / 8 scales and the number of N queries, use the Transformer decoder to calculate the attention relationship between different human body parts from the resolution image features of different scales, and obtain C*N dimensional embedding features, where C is the number of channels and N is the number of human body parts and clothing types to be segmented; Perform an inner product of the embedding feature and the pixel-level embedding to obtain N H*W-dimensional binary semantic segmentation maps, where H and W represent the height and width of the binary semantic segmentation map, respectively. Each map represents the parsing result of a human body part, and the pixel value of the corresponding part is represented by 1, and the pixel values ​​of other areas are represented by 0; Fuse N binary semantic segmentation maps, that is, replace the 1 in the binary map with the label of the corresponding segmentation part, and add the N replaced maps to obtain the final human body analysis result; The step of inputting the 1 / 4 scale resolution image features into the pixel-level decoder to obtain pixel-level embedding is as follows: A pixel-level decoder is used to perform a 1*1 convolution on the 1 / 4 scale resolution image features and the original image, the 1 / 4 scale resolution image features are connected, and then an upsampling operation is continuously performed through a deconvolution operation, so as to obtain multi-resolution image features of different scales; The input image features with resolutions of 1 / 32, 1 / 16, and 1 / 8 and the number of queries N are used, and the attention relationship between different human body parts is calculated from the image features with resolutions of different scales using the Transformer decoder to obtain C*N dimensional embedding features, where C is the number of channels and N is the number of human body parts and clothing types to be segmented, specifically: Input the 1 / 32, 1 / 16, 1 / 8 scale resolution image features and N queries to the Transformer decoder, where N is the type of human body parts and clothing to be segmented. First, calculate the cross attention: X l =softmax(Q l K l )V l +X l-1 Where l is the subscript of the number of layers, X l is the query feature of the lth layer, Q l is the query input to layer l, V l and K l The image features of the first layer input are transformed by two different linear transformation functions f V and f K The transformed matrix is ​​then normalized for the cross-attention result and passed through a self-attention layer. The result calculated by the self-attention layer will be output as the final query feature after normalization through the feed-forward layer. The Transformer decoder decodes the image features at three scales: 1 / 8, 1 / 16, and 1 / 32. The three decoding operations are repeated L times, i.e., a total of 3L decodings are performed. The decoded image will pass through a multi-layer perceptron to generate a C*N-dimensional embedding feature, where C is the number of channels and N is the number of human body parts and clothing to be segmented.

2. The human body analysis method based on Transformer according to claim 1, It is characterized in that The input human body picture and the parsed picture are input, data enhancement is performed on the input picture, and the picture is processed into a uniform size, specifically: Human body pictures and human body analysis pictures are input. Human body pictures are collected from the Internet, and analysis pictures are pictures of different parts of the human body and clothes marked with different colors manually. In order to make the trained model robust, the pictures are randomly rotated, horizontally mirrored, and randomly cropped for data enhancement. Finally, all pictures are scaled to a uniform size.

3. The human body analysis method based on Transformer as claimed in claim 1, It is characterized in that The DeepLabV3+ backbone network is used to extract low-resolution image features of the image, and multiple upsampling operations are performed on the low-resolution image features to obtain resolution image features of four scales: 1 / 32, 1 / 16, 1 / 8 and 1 / 4, specifically: The image is input into the DeepLabV3+ backbone network, and the DeepLabV3+ backbone network performs 1*1 convolution, 3*3 dilated convolution with ratios of 6, 12 and 18 and image pooling operations on the image to obtain 5 feature maps. The 5 feature maps are cascaded and subjected to 1*1 convolution operations to obtain low-resolution image features. The low-resolution image features are upsampled to obtain image features with four scales of resolution: 1 / 32, 1 / 16, 1 / 8, and 1 / 4.

4. A human body parsing system based on Transformer, It is characterized in that The system comprises: An image preprocessing unit, used to perform data enhancement on the input human body image and human body analysis image, and process them into a uniform size; A multi-scale resolution image feature unit, for extracting low-resolution image features of the image using the DeepLabV3+ backbone network, performing multiple upsampling operations on the low-resolution image features, and obtaining image features with resolutions of four scales: 1 / 32, 1 / 16, 1 / 8, and 1 / 4; A pixel-level decoding unit, used to input the 1 / 4 scale resolution image features into a pixel-level decoder for operation to obtain pixel-level embedding; A Transformer decoding unit is used to input the resolution image features of 1 / 32, 1 / 16, and 1 / 8 scales and N query numbers, and use the Transformer decoder to calculate the attention relationship between different human body parts from the resolution image features of different scales to obtain C*N dimensional embedding features, where C is the number of channels and N is the number of human body parts and clothing types to be segmented; A semantic segmentation map acquisition unit is used to perform inner product on the embedded features and the pixel-level embedding to obtain N H*W-dimensional binary semantic segmentation maps, each of which represents the parsing result of a human body part, the pixel value of the corresponding part is represented by 1, and the pixel values ​​of other areas are represented by 0, and H and W represent the height and width of the binary semantic segmentation map respectively; The human body parsing map unit is used to fuse N binary semantic segmentation maps, that is, replace the 1 in the binary map with the label of the corresponding segmentation part, and add the N replaced maps to obtain the final human body parsing result; The pixel-level decoding unit needs to use a pixel-level decoder to perform 1*1 convolution on the 1 / 4 scale resolution image features and the original image, connect the 1 / 4 scale resolution image features, and then continuously perform upsampling operations through deconvolution operations, so as to obtain multi-resolution image features of different scales; The Transformer decoding unit needs to input the 1 / 32, 1 / 16, 1 / 8 scale resolution image features and N queries into the Transformer decoder, where N is the type of human body parts and clothing to be segmented. First, the cross attention is calculated: X l =softmax(Q l K l )V l +X l-1 Where l is the subscript of the number of layers, X l is the query feature of the lth layer, Q l is the query input to layer l, V l and K l The image features of the first layer input are transformed by two different linear transformation functions f V and f K The transformed matrix is ​​then normalized for the cross-attention result and passed through a self-attention layer. The result calculated by the self-attention layer will be output as the final query feature after normalization through the feed-forward layer. The Transformer decoder decodes the image features at three scales: 1 / 8, 1 / 16, and 1 / 32. The three decoding operations are repeated L times, i.e., a total of 3L decodings are performed. The decoded image will pass through a multi-layer perceptron to generate a C*N-dimensional embedding feature, where C is the number of channels and N is the number of human body parts and clothing to be segmented.

5. The Transformer-based human body analysis system according to claim 4, It is characterized in that The image preprocessing unit needs to input human body pictures and human body analysis pictures. The human body pictures are collected from the Internet, and the analysis pictures are pictures of different parts of the human body and clothes manually marked with different colors. In order to make the trained model robust, the pictures are randomly rotated, horizontally mirrored, and randomly cropped for data enhancement, and finally all the pictures are scaled to a uniform size.

6. The Transformer-based human body analysis system according to claim 4, It is characterized in that The multi-scale resolution image feature unit needs to input the image into the DeepLabV3+ backbone network. The DeepLabV3+ backbone network performs 1*1 convolution, 3*3 dilated convolution with ratios of 6, 12 and 18 and image pooling operations on the image to obtain 5 feature maps. The 5 feature maps are cascaded and subjected to 1*1 convolution operations to obtain low-resolution image features. The low-resolution image features are upsampled to obtain image features with four scales of resolution: 1 / 32, 1 / 16, 1 / 8, and 1 / 4.