Potential diffusion virtual try-on research method based on cross attention hierarchical fusion
The virtual try-on method that combines HRNet and CLIP encoder solves the problem of alignment between clothing and human body in complex postures, achieves a natural and fitting virtual try-on effect, and enhances the model's understanding of complex postures and ability to retain details.
Patent Information
- Application Number
- CN202510612605.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional virtual try-on methods have shortcomings in handling the spatial alignment of clothing and the human body. Especially in complex or dynamic postures, clothing may overlap, penetrate or distort, resulting in unnatural virtual try-on effects.
The HRNet method is used to extract high-resolution feature maps of human posture, and the Warp-HR module is used to perform clothing warping adjustment. The CLIP text and image encoder is combined for natural language description and encoding. The CAF network model is used for multi-dimensional feature fusion and diffusion model to generate virtual try-on effects.
It improves the naturalness and fit of virtual try-on, enhances the ability to understand complex postures, ensures the spatial semantic alignment of clothing and the human body in different postures, and generates realistic and detail-retaining virtual try-on effects.
Smart Images

Figure CN120707998A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of latent diffusion virtual try-on technology, in particular to a research method of latent diffusion virtual try-on based on cross-attention hierarchical fusion. Background Art
[0002] Latent Diffusion Virtual Try-On research technology is a new technology designed to improve the realism and accuracy of virtual try-on. It combines HRNet for accurate human pose estimation, a three-layer cross-attention fusion mechanism to enhance the spatial alignment of human pose and target clothing, and introduces CLIP contrastive learning to accurately extract complex human poses.
[0003] In the field of potential diffusion virtual try-on, traditional virtual try-on methods have shortcomings in dealing with the spatial alignment of clothing and the human body, especially in complex or dynamic postures, where clothing may overlap, penetrate or distort, resulting in unnatural virtual try-on effects. Although diffusion models can generate more natural lighting, shadows and fine textures under dynamic backgrounds, they still face challenges in capturing the dynamic relationship between the human body and clothing, especially in extreme postures where the geometric deformation and stretching of clothing may lead to a decrease in the quality of the generated image. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a potential diffusion virtual try-on research method based on cross-attention hierarchical fusion to solve the problem that traditional virtual try-on methods have shortcomings in dealing with the spatial alignment of clothing and the human body, especially in complex or dynamic postures, clothing may overlap, penetrate or distort, resulting in the virtual try-on effect being unnatural.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a potential diffusion virtual try-on research method based on cross-attention hierarchical fusion, which includes:
[0008] The HRNet method is used to analyze human body images, extract high-resolution feature maps of human posture, and obtain the position information of 18 key points of the human body;
[0009] The Warp-HR module combines the position information of 18 key points of the human body with the clothing image, adjusts the local area of the clothing through affine transformation, completes the clothing warping, and obtains the clothing shape that matches the user's posture;
[0010] The CLIP text encoder is used to describe and encode complex human postures in natural language to obtain a text description of the human posture. At the same time, the CLIP image encoder is used to encode the human posture image to obtain a human posture image encoding.
[0011] Based on human posture image coding, accurate human posture information is obtained to solve the problem of virtual try-on under complex postures;
[0012] The CLIP encoder is used to convert the human image, target clothing, and human posture text description into a latent space, and features are extracted through three cross-attention modules to form a CAF network model;
[0013] Based on the CAF network model, human posture, clothing images and text description information are deeply mined and integrated from multiple dimensions to obtain a virtual try-on effect report.
[0014] As a preferred solution of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion described in the present invention, the HRNet method is used to analyze the human body image, extract the high-resolution feature map of the human body posture, and obtain the position information of 18 key points of the human body. The specific steps are as follows:
[0015] Collect images containing human bodies from product display images from e-commerce platforms and photos uploaded by users;
[0016] Use the HRNet network model to process the input image, extract detailed information from multiple scales, and output a high-resolution feature map of human posture;
[0017] Based on the high-resolution feature map of human body posture, the non-maximum suppression technology is used to determine the key points, and the position information of 18 key points of the human body is obtained (x i ,y i ), the expression is:
[0018]
[0019] Among them, S i (x,y) represents the response score map of the i-th key point at position (x,y), (x i ,y i ) is the position information of 18 key points of the human body obtained from the input image.
[0020] As a preferred solution of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion described in the present invention, the Warp-HR module is used to combine the position information of 18 key points of the human body with the clothing image, and the local area of the clothing is adjusted through affine transformation to complete the clothing warping and obtain the clothing shape that matches the user's posture. The specific steps are as follows:
[0021] Based on the high-resolution human posture feature map extracted by HRNet, the local area of the clothing is adjusted using affine transformation;
[0022] Based on the position information of 18 key points of the human body, the affine transformation matrix A is calculated so that the clothing image can be accurately warped to a state that matches the user's posture;
[0023] Applying the above affine transformation to the clothing image realizes clothing warping so that it fits naturally with the user's body contour and posture, and obtains the clothing shape that matches the user's posture.
[0024] As a preferred solution of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion described in the present invention, wherein: the CLIP text encoder is used to describe and encode complex human postures in natural language to obtain human posture text descriptions, and the CLIP image encoder is used to encode human posture images to obtain human posture image encodings. The specific steps are as follows:
[0025] Provide detailed natural language descriptions of complex human postures;
[0026] Input the above natural language description into the text encoder of CLIP and encode it to obtain the text feature vector v t (D P ), the expression is:
[0027] v t (D P )=Encode text (D P );
[0028] Among them, v t (D P ) represents the text feature vector obtained by encoding the natural language description of complex human posture through the CLIP text encoder, D P Representing natural language descriptions of complex human postures, Encode text represents the CLIP text encoder, which is used to convert natural language descriptions into feature vectors;
[0029] Encode the human body posture through the CLIP image encoder to obtain the image feature vector v i (I P ), the expression is:
[0030] v i (I P )=Encode image (I P );
[0031] Among them, vi (I P ) represents the image feature vector obtained by encoding the human posture image through the CLIP image encoder, I P Represents an image containing human posture, Encode image Represents the CLIP image encoder, which is used to convert an image into a feature vector.
[0032] As a preferred solution of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion described in the present invention, wherein: based on human posture image encoding, accurate human posture information is obtained to solve the virtual try-on problem under complex postures, the specific steps are as follows:
[0033] Use contrastive learning to measure the similarity between text descriptions and corresponding human pose images and minimize the difference between text and image representations in the shared multimodal space, expressed as:
[0034]
[0035] Among them, S tn,im Represents the similarity between text and image, S in,tm represents the similarity from image to text, Represents the similarity function between two vectors a and b, Represents the nth natural language description D P The feature vector after being encoded by the CLIP text encoder, Represents the mth human posture image I P Feature vector after being encoded by the CLIP image encoder;
[0036] Define the contrastive learning loss function, the expression is:
[0037]
[0038] Among them, L contrastive represents the contrastive learning loss function, N represents the number of samples, S tn,in Indicates the similarity of the positive sample to the text to the image, S tn,m Indicates the similarity between text and other images, S in,tn Indicates the similarity between the positive sample pair image and text, S in,tm Indicates the similarity between an image and other texts.
[0039] As a preferred solution of the latent diffusion virtual try-on research method based on cross-attention hierarchical fusion described in the present invention, the CLIP encoder is used to convert the human image, target clothing, and human posture text description into a latent space, and features are extracted through three cross-attention modules to form a CAF network model. The specific steps are as follows:
[0040] For human body images and target clothing, CLIP image encoder is used to encode them and obtain the corresponding feature vectors.
[0041] For the text description of human posture, use CLIP text encoder to encode and obtain the corresponding feature vector;
[0042] The first cross attention module uses the cosine similarity function to calculate the human posture text description data v t (D p ), human body posture image data v i (I p ) and target clothing data v c The similarity between (C) is expressed as:
[0043]
[0044] Among them, S t,i and S t,c Represents the similarity between text and image, and text and clothing respectively;
[0045] The second cross attention module combines human posture description and CLIP to extract the key features of clothing under different human postures through weighted clothing image features. The expression is:
[0046]
[0047] Among them, w i Based on the similarity S t,c The calculated weight coefficient is used to adjust the importance of each part of the clothing feature;
[0048] The third cross-attention module is modeled using a multi-head self-attention mechanism;
[0049] In the third cross-attention module, different “heads” can focus on different parts of the input, thereby better capturing local and global information, expressed as:
[0050]
[0051] Among them, Q, K, V represent query, key and value matrices respectively, d k is the dimension of the key vector;
[0052] Finally, the outputs of multiple heads are concatenated and subjected to a linear transformation to obtain the final output.
[0053] As a preferred solution of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion described in the present invention, the method is based on the CAF network model, which deeply mines and integrates human posture, clothing images and text description information from multiple dimensions to obtain a virtual try-on effect report. The specific steps are as follows:
[0054] After completing the above three cross-attention modules, the CAF network model will deeply fuse these features;
[0055] The deep fusion includes local feature enhancement, global semantic alignment, and feature splicing and combination;
[0056] The fused features are used as input into the diffusion model to generate the final virtual try-on effect image.
[0057] The diffusion model includes a denoising UNet structure, and a cross-attention mechanism is applied on this basis;
[0058] According to the generated virtual try-on effect diagram, the evaluation index is calculated to obtain the evaluation result;
[0059] Based on the above evaluation results, a virtual try-on effect report is generated.
[0060] As a preferred solution of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion described in the present invention, the evaluation indicators include realism score, detail retention rate and spatial alignment, and the virtual try-on effect report includes the input human posture image, target clothing image and corresponding text description, the generated virtual try-on effect diagram and the specific values of each evaluation indicator and its explanation.
[0061] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion as described in the first aspect of the present invention is implemented.
[0062] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion as described in the first aspect of the present invention.
[0063] The beneficial effects of the present invention are as follows: by using the HRNet network model to process the input image, extracting detail information from multiple scales and outputting a high-resolution feature map of human posture, accurate analysis of human posture is achieved. This step uses non-maximum suppression technology to determine key points, thereby obtaining the position information of 18 key points of the human body. The Warp-HR module performs clothing warping. The method can accurately calculate the affine transformation matrix and apply it to the clothing image, making the virtual try-on effect more natural and fitting the user's body shape and posture. By providing a detailed natural language description of complex human postures and encoding them using the CLIP text encoder and image encoder respectively, accurate human posture text descriptions and image encodings can be obtained, which not only enhances the ability to understand complex postures, but also promotes semantic consistency between text descriptions and actual images. Contrastive learning is used to measure the similarity between text descriptions and corresponding human posture images, and minimize the difference in text and image representations in the shared multimodal space. This process effectively solves the virtual try-on problem under complex postures. By defining and optimizing the contrastive learning loss function, the model can ensure spatial semantic alignment of clothing and human body in different postures while maintaining details. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0065] Figure 1 Flowchart of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion in Example 1.
[0066] Figure 2 This is the structural diagram of CAF-VTON in Example 1.
[0067] Figure 3 This is a structural diagram of Warp-HR in Example 1. DETAILED DESCRIPTION
[0068] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0069] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0070] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0071] Example 1, with reference to Figure 1 、 Figure 2 and Figure 3 , which is the first embodiment of the present invention, provides a potential diffusion virtual try-on research method based on cross-attention hierarchical fusion, including the following steps:
[0072] S1. Analyze human body images using the HRNet method, extract high-resolution feature maps of human body postures, and obtain the position information of 18 key points of the human body;
[0073] Furthermore, human body images are collected from product display pictures from e-commerce platforms and photos uploaded by users;
[0074] Use the HRNet network model to process the input image, extract detailed information from multiple scales, and output a high-resolution feature map of human posture;
[0075] Based on the high-resolution feature map of human body posture, the non-maximum suppression technology is used to determine the key points, and the position information of 18 key points of the human body is obtained (x i ,y i ), the expression is:
[0076]
[0077] Among them, S i (x,y) represents the response score map of the i-th key point at position (x,y), (x i ,y i ) is to obtain the position information of 18 key points of the human body from the input image;
[0078] It should be noted that the HRNet network can not only effectively capture the key point position information of human posture through multi-scale feature fusion, but also maintain sensitivity to details at different resolutions. Therefore, when using non-maximum suppression technology to determine key points, the position of various parts of the human body can be more accurately located, thereby improving the accuracy of subsequent clothing warping processing.
[0079] S2, using the Warp-HR module to combine the position information of 18 key points of the human body with the clothing image, adjust the local area of the clothing through affine transformation, complete the clothing warping, and obtain the clothing shape that matches the user's posture;
[0080] Furthermore, based on the high-resolution feature map of human posture extracted by HRNet, affine transformation is used to adjust the local area of the clothing;
[0081] Based on the position information of 18 key points of the human body, the affine transformation matrix A is calculated so that the clothing image can be accurately warped to a state that matches the user's posture;
[0082] Apply the above affine transformation to the clothing image to achieve clothing warping so that it fits naturally with the user's body contour and posture, and obtain a clothing shape that matches the user's posture;
[0083] It should be noted that the calculation of the affine transformation matrix is based on the position information of 18 key points of the human body. It can effectively adjust the local area of the clothing image to match the user's body contour and posture. This process not only takes into account the changes in human posture, but also adapts to the personalized needs of users of different body shapes, making the virtual try-on effect more natural and realistic.
[0084] S3, using the CLIP text encoder to perform natural language description and encoding of the complex human posture to obtain a human posture text description, and at the same time using the CLIP image encoder to encode the human posture image to obtain a human posture image encoding;
[0085] Going further, detailed natural language description of complex human postures;
[0086] Input the above natural language description into the text encoder of CLIP and encode it to obtain the text feature vector v t (D P ), the expression is:
[0087] v t (D P )=Encode text (D P );
[0088] Among them, v t (D P ) represents the text feature vector obtained by encoding the natural language description of complex human posture through the CLIP text encoder, D P Representing natural language descriptions of complex human postures, Encode text represents the CLIP text encoder, which is used to convert natural language descriptions into feature vectors;
[0089] Encode the human body posture through the CLIP image encoder to obtain the image feature vector v i (I P ), the expression is:
[0090] v i (I P )=Encode image (I P );
[0091] Among them, v i (I P ) represents the image feature vector obtained by encoding the human posture image through the CLIP image encoder, I P Represents an image containing human posture, Encode image Represents the CLIP image encoder, which is used to convert an image into a feature vector;
[0092] It should be noted that the application of CLIP text encoder and CLIP image encoder provides strong support for the description of complex human postures. The multimodal data processing method not only improves the ability to understand human postures, but also promotes the semantic consistency between text descriptions and actual images, thus laying the foundation for subsequent contrastive learning and cross-attention modules.
[0093] S4, based on human posture image coding, obtains accurate human posture information and solves the problem of virtual try-on under complex postures;
[0094] Furthermore, contrastive learning is used to measure the similarity between text descriptions and corresponding human pose images, and to minimize the difference between text and image representations in the shared multimodal space, as expressed by:
[0095]
[0096] Among them, S tn,im Represents the similarity between text and image, S in,tm represents the similarity from image to text, Represents the similarity function between two vectors a and b, Represents the nth natural language description D P The feature vector after being encoded by the CLIP text encoder, Represents the mth human posture image I P Feature vector after being encoded by the CLIP image encoder;
[0097] Define the contrastive learning loss function, the expression is:
[0098]
[0099] Among them, L contrastive represents the contrastive learning loss function, N represents the number of samples, S tn,in Indicates the similarity of the positive sample to the text to the image, S tn,m Indicates the similarity between text and other images, Sin,tn Indicates the similarity between the positive sample pair image and text, S in,tm Indicates the similarity between an image and other texts;
[0100] It should be noted that the design of the contrastive learning loss function aims to minimize the differences in text and image representations in a shared multimodal space, thereby improving the model's virtual try-on effect under complex postures. In this way, not only can the model's ability to understand different postures be enhanced, but also the realism and user experience of virtual try-on can be significantly improved.
[0101] S5, using the CLIP encoder to convert the human image, target clothing, and human posture text description into a latent space, and extract features through three cross-attention modules to form a CAF network model;
[0102] Furthermore, the human body image and target clothing are encoded using the CLIP image encoder to obtain the corresponding feature vectors.
[0103] For the text description of human posture, use CLIP text encoder to encode and obtain the corresponding feature vector;
[0104] The first cross attention module uses the cosine similarity function to calculate the human posture text description data v t (D p ), human body posture image data v i (I p ) and target clothing data v c The similarity between (C) is expressed as:
[0105]
[0106] Among them, S t,i and S t,c Represents the similarity between text and image, and text and clothing respectively;
[0107] The second cross attention module combines human posture description and CLIP to extract the key features of clothing under different human postures through weighted clothing image features. The expression is:
[0108]
[0109] Among them, w i Based on the similarity S t,c The calculated weight coefficient is used to adjust the importance of each part of the clothing feature;
[0110] The third cross-attention module is modeled using a multi-head self-attention mechanism;
[0111] In the third cross-attention module, different “heads” can focus on different parts of the input, thereby better capturing local and global information, expressed as:
[0112]
[0113] Among them, Q, K, V represent query, key and value matrices respectively, d k is the dimension of the key vector;
[0114] Finally, the outputs of multiple heads are concatenated and subjected to a linear transformation to obtain the final output;
[0115] It should be noted that the introduction of the three cross-attention modules greatly enhances the ability of the CAF network model to process multimodal data. In particular, the multi-head self-attention mechanism adopted in the third cross-attention module can better capture the local and global features of the input data, thereby generating more accurate and detailed virtual try-on renderings.
[0116] S6. Based on the CAF network model, the human body posture, clothing images and text description information are deeply mined and integrated from multiple dimensions to obtain a virtual try-on effect report;
[0117] Furthermore, after completing the above three cross-attention modules, the CAF network model will deeply fuse these features;
[0118] Deep fusion includes local feature enhancement, global semantic alignment, and feature splicing and combination;
[0119] The fused features are used as input into the diffusion model to generate the final virtual try-on effect image.
[0120] The diffusion model includes a denoising UNet structure and applies a cross-attention mechanism on this basis;
[0121] According to the generated virtual try-on effect diagram, the evaluation index is calculated to obtain the evaluation result;
[0122] Based on the above evaluation results, a virtual try-on effect report is generated;
[0123] Evaluation indicators include realism score, detail retention rate and spatial alignment. The virtual try-on effect report includes the input human posture image, target clothing image and corresponding text description, the generated virtual try-on effect image, and the specific values and explanations of each evaluation indicator.
[0124] It should be noted that the deep fusion strategy ensures the effectiveness of in-depth mining and integration of human posture, clothing images and text description information from multiple dimensions. The virtual try-on renderings generated by this method not only have high realism, but also retain the detailed features of the original clothing, providing users with a virtual try-on experience that is close to reality. In addition, the virtual try-on effect report generated according to the evaluation indicators helps to further optimize the model performance and provide users with valuable feedback information.
[0125] This embodiment also provides a computer device, which is applicable to the case of a potential diffusion virtual try-on research method based on cross-attention hierarchical fusion, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion proposed in the above embodiment.
[0126] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through Wi-Fi, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.
[0127] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, disk or optical disk.
[0128] In summary, the present invention uses the HRNet network model to process the input image, extracts detailed information from multiple scales, and outputs a high-resolution feature map of human posture, thereby achieving accurate analysis of human posture. This step uses non-maximum suppression technology to determine key points, thereby obtaining the position information of 18 key points of the human body. The Warp-HR module performs clothing warping. The method can accurately calculate the affine transformation matrix and apply it to the clothing image, making the virtual try-on effect more natural and fitting the user's body shape and posture. By providing a detailed natural language description of complex human postures and encoding them using the CLIP text encoder and image encoder respectively, accurate human posture text descriptions and image encodings can be obtained, which not only enhances the ability to understand complex postures but also promotes semantic consistency between text descriptions and actual images. Contrastive learning is used to measure the similarity between text descriptions and corresponding human posture images, and minimize the difference in text and image representations in the shared multimodal space. This process effectively solves the virtual try-on problem under complex postures. By defining and optimizing the contrastive learning loss function, the model can ensure the spatial semantic alignment of clothing and human body in different postures while maintaining details.
[0129] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A potential diffusion virtual try-on method based on cross-attention hierarchical fusion, characterized by: include: The HRNet method is used to analyze human body images, extract high-resolution feature maps of human posture, and obtain the position information of 18 key points of the human body; The Warp-HR module combines the position information of 18 key points of the human body with the clothing image, adjusts the local area of the clothing through affine transformation, completes the clothing warping, and obtains the clothing shape that matches the user's posture; The CLIP text encoder is used to describe and encode complex human postures in natural language to obtain a text description of the human posture. At the same time, the CLIP image encoder is used to encode the human posture image to obtain a human posture image encoding. Based on human posture image coding, accurate human posture information is obtained to solve the problem of virtual try-on under complex postures; The CLIP encoder is used to convert the human image, target clothing, and human posture text description into a latent space, and features are extracted through three cross-attention modules to form a CAF network model; Based on the CAF network model, human posture, clothing images and text description information are deeply mined and integrated from multiple dimensions to obtain a virtual try-on effect report.
2. The latent diffusion virtual try-on research method based on cross-attention layered fusion as claimed in claim 1 is characterized by: The HRNet method is used to analyze human body images, extract high-resolution feature maps of human body postures, and obtain the position information of 18 key points of the human body. The specific steps are as follows: Collect images containing human bodies from product display images from e-commerce platforms and photos uploaded by users; Use the HRNet network model to process the input image, extract detailed information from multiple scales, and output a high-resolution feature map of human posture; Based on the high-resolution feature map of human body posture, the non-maximum suppression technology is used to determine the key points, and the position information of 18 key points of the human body is obtained (x i ,y i ), the expression is: Among them, S i (x,y) represents the response score map of the i-th key point at position (x,y), (x i ,y i ) is the position information of 18 key points of the human body obtained from the input image.
3. The latent diffusion virtual try-on research method based on cross-attention layered fusion as claimed in claim 2 is characterized by: The Warp-HR module combines the position information of 18 key points of the human body with the clothing image, adjusts the local area of the clothing through affine transformation, completes the clothing warping, and obtains the clothing shape that matches the user's posture. The specific steps are as follows: Based on the high-resolution human posture feature map extracted by HRNet, the local area of the clothing is adjusted using affine transformation; Based on the position information of 18 key points of the human body, the affine transformation matrix A is calculated so that the clothing image can be accurately warped to a state that matches the user's posture; Applying the above affine transformation to the clothing image realizes clothing warping so that it fits naturally with the user's body contour and posture, and obtains the clothing shape that matches the user's posture.
4. The latent diffusion virtual try-on research method based on cross-attention layered fusion as claimed in claim 3 is characterized by: The CLIP text encoder is used to describe and encode complex human postures in natural language to obtain a human posture text description, and the CLIP image encoder is used to encode the human posture image to obtain a human posture image encoding. The specific steps are: Provide detailed natural language descriptions of complex human postures; Input the above natural language description into the text encoder of CLIP and encode it to obtain the text feature vector v t (D P ), the expression is: v t (D P )=Encode text (D P ); Among them, v t (D P ) represents the text feature vector obtained by encoding the natural language description of complex human posture through the CLIP text encoder, D P Representing natural language descriptions of complex human postures, Encode text represents the CLIP text encoder, which is used to convert natural language descriptions into feature vectors; Encode the human body posture through the CLIP image encoder to obtain the image feature vector v i (I P ), the expression is: v i (I P )=Encode image (I P ); Among them, v i (I P ) represents the image feature vector obtained by encoding the human posture image through the CLIP image encoder, I P Represents an image containing human posture, Encode image Represents the CLIP image encoder, which is used to convert an image into a feature vector.
5. The latent diffusion virtual try-on research method based on cross-attention layered fusion as claimed in claim 4 is characterized by: The method of obtaining accurate human posture information based on human posture image coding and solving the virtual try-on problem under complex postures is specifically carried out as follows: Use contrastive learning to measure the similarity between text descriptions and corresponding human pose images and minimize the difference between text and image representations in the shared multimodal space, expressed as: Among them, S tn,im Represents the similarity between text and image, S in,tm represents the similarity between image and text, Represents the similarity function between two vectors a and b, Represents the nth natural language description D P The feature vector after being encoded by the CLIP text encoder, Represents the mth human posture image I P Feature vector after being encoded by the CLIP image encoder; Define the contrastive learning loss function, the expression is: Among them, L contrastive represents the contrastive learning loss function, N represents the number of samples, S tn,in Indicates the similarity of the positive sample to the text to the image, S tn,m Indicates the similarity between text and other images, S in,tn Indicates the similarity between the positive sample pair image and text, S in,tm Indicates the similarity between an image and other texts.
6. The latent diffusion virtual try-on research method based on cross-attention layered fusion as claimed in claim 5 is characterized by: The CLIP encoder is used to convert the human image, target clothing, and human posture text description into a latent space, and features are extracted through three cross-attention modules to form a CAF network model. The specific steps are as follows: For human body images and target clothing, CLIP image encoder is used to encode them and obtain the corresponding feature vectors. For the text description of human posture, use CLIP text encoder to encode and obtain the corresponding feature vector; The first cross attention module uses the cosine similarity function to calculate the human posture text description data v t (D p ), human body posture image data v i (I p ) and target clothing data v c The similarity between (C) is expressed as: Among them, S t,i and S t,c Represents the similarity between text and image, and text and clothing respectively; The second cross attention module combines human posture description and CLIP to extract the key features of clothing under different human postures through weighted clothing image features. The expression is: Among them, w i Based on the similarity S t,c The calculated weight coefficient is used to adjust the importance of each part of the clothing feature; The third cross-attention module is modeled using a multi-head self-attention mechanism; In the third cross-attention module, different "heads" can focus on different parts of the input, thereby better capturing local and global information, expressed as: Among them, Q, K, V represent query, key and value matrices respectively, d k is the dimension of the key vector; Finally, the outputs of multiple heads are concatenated and subjected to a linear transformation to obtain the final output.
7. The latent diffusion virtual try-on research method based on cross-attention layered fusion as claimed in claim 6 is characterized by: Based on the CAF network model, the human body posture, clothing image and text description information are deeply mined and integrated from multiple dimensions to obtain a virtual try-on effect report. The specific steps are as follows: After completing the above three cross-attention modules, the CAF network model will deeply fuse these features; The deep fusion includes local feature enhancement, global semantic alignment, and feature splicing and combination; The fused features are used as input into the diffusion model to generate the final virtual try-on effect image. The diffusion model includes a denoising UNet structure, and a cross-attention mechanism is applied on this basis; According to the generated virtual try-on effect diagram, the evaluation index is calculated to obtain the evaluation result; Based on the above evaluation results, a virtual try-on effect report is generated.
8. The latent diffusion virtual try-on research method based on cross-attention layered fusion as claimed in claim 7 is characterized by: The evaluation indicators include realism score, detail retention rate and spatial alignment. The virtual try-on effect report includes the input human posture image, target clothing image and corresponding text description, the generated virtual try-on effect diagram, and the specific values of each evaluation indicator and its explanation.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the potential diffusion virtual try-on research method based on cross-attention hierarchical fusion according to any one of claims 1 to 8 are implemented.