Maskless semantic collapse clothing splitting method and system
Patent Information
- Application Number
- CN202611030392.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-11
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]该类方法在常规上衣、裤子等清晰闭合区域中具有一定效果,但面对真实人物图中的复杂服饰时仍存在不足
[0027]本发明通过无蒙版的语义坍缩与隐空间协同重绘,实现复杂服饰的完整拆分和非目标区域纯色化处理。本发明通过目标拆分指令生成目标服饰语义向量和非目标排除语义向量,不依赖固定分割类别或人工蒙版,能够响应细粒度服饰拆分需求;通过目标服饰注意力权重图、遮挡关联边和协同去噪重绘,使被遮挡或分散显示的同一服饰区域按照同一语义进行重建,减少漏拆和断裂;通过非目标替换权重图和纯色背景隐变量,将人体、背景及非目标服饰引导至统一背景表达;通过边界保护权重图和残留非目标特征回写,降低服饰边缘被误替换或非目标内容残留的情况。
Smart Images

Figure CN122821133A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method and system for decomposing clothing using semantic collapse without a mask. Background Technology
[0002] Existing methods for clothing image segmentation often rely on processes such as image matting, human body analysis, instance segmentation, or virtual try-on. Typically, a pixel-level mask is first generated, and then the target is extracted based on edges, transparency, connected components, or preset clothing categories.
[0003] While these methods are effective in clearly closed regions such as regular tops and pants, they still have shortcomings when dealing with complex clothing in real-life images. Firstly, the same garment often presents multiple discrete visible areas due to arms, hair, bag straps, wrinkles, shadows, or occlusion by outerwear. Traditional methods easily misclassify small areas of tulle, lace, tassels, and skirt edges as background noise or non-subject areas, leading to incomplete, fragmented, or only the largest connected portion retained in the segmentation results. Secondly, existing segmentation models are mostly limited by fixed category labels, making it difficult to respond to fine-grained natural language objectives such as "segment out the waistband," "segment out the outer tulle," and "segment out the puff-sleeved top." Even when using text editing models, they typically focus on overall image rewriting, without specifically addressing the complete preservation of the target garment, the uniform collapse of non-target human figures and backgrounds into solid-color backgrounds, and the collaborative reconstruction of discrete garment regions. Summary of the Invention
[0004] To address the aforementioned problems, embodiments of the present invention provide a method for semantically collapsing clothing without a mask, the method comprising:
[0005] Obtain the original image containing the character's clothing and the target segmentation instruction, perform clothing lexical parsing on the target segmentation instruction to obtain the target clothing semantic vector, and generate a non-target exclusion semantic vector based on the target clothing semantic vector;
[0006] The original image is encoded as image latent variables, and the semantic vector of the target clothing and the semantic vector of the non-target exclusion are input into a diffusion image editing network to generate a target clothing attention weight map and a non-target replacement weight map.
[0007] Based on the target clothing attention weight map, the first visible clothing region and the second visible clothing region are determined in the latent variables of the image, and occlusion association edges are generated;
[0008] Based on the occlusion association edge, the first visible clothing area and the second visible clothing area are collaboratively denoised and redrawn. Based on the non-target replacement weight map, the latent space features of the human body, background and non-target clothing are replaced with solid color background latent variables.
[0009] A boundary protection weight map is generated from the edge response of the target clothing attention weight map, and the replacement response of the non-target replacement weight map to the edge region of the target clothing is suppressed according to the boundary protection weight map to obtain the latent variables of the target clothing.
[0010] The latent variables of the target clothing are decoded and semantically back-encoded to extract residual non-target features. The residual non-target features are then written back to the non-target replacement weight map, and the resulting solid-color background clothing split image is output.
[0011] Furthermore, the clothing lexical parsing includes: extracting clothing name lexicals, clothing part lexicals, and clothing material lexicals from the target splitting instruction; establishing a main pointer term based on the clothing name lexicals, establishing a local limiting term based on the clothing part lexicals, and establishing a boundary limiting term based on the clothing material lexicals; concatenating and encoding the main pointer term, the local limiting term, and the boundary limiting term into the target clothing semantic vector; and generating a non-target exclusion semantic vector corresponding to the human body region, the background region, and the non-target clothing region based on the target clothing semantic vector.
[0012] Furthermore, the diffusion-based image editing network includes an attention layer; the method for generating the target clothing attention weight map and the non-target replacement weight map includes: converting the image latent variables into a visual word sequence, converting the target clothing semantic vector into a first conditional vector, and converting the non-target exclusion semantic vector into a second conditional vector; querying the visual word sequence with the first conditional vector in the attention layer to obtain the target clothing attention weight map; and querying the visual word sequence with the second conditional vector to obtain the non-target replacement weight map.
[0013] Furthermore, the method for generating the occlusion association edge includes: extracting a first response region as the first visible clothing region and extracting a second response region as the second visible clothing region from the target clothing attention weight map; extracting an occlusion response region located between the first visible clothing region and the second visible clothing region from the non-target replacement weight map; establishing the occlusion association edge connecting the first visible clothing region and the second visible clothing region based on the occlusion response region, and writing the occlusion association edge into the denoising path of the diffusion image editing network.
[0014] Furthermore, the collaborative denoising and redrawing of the first visible clothing region and the second visible clothing region includes: generating a first region redrawing condition and a second region redrawing condition based on the occlusion association edge; writing the first region redrawing condition into the latent space feature corresponding to the first visible clothing region, and writing the second region redrawing condition into the latent space feature corresponding to the second visible clothing region; and maintaining the first region redrawing condition and the second region redrawing condition sharing the target clothing semantic vector within the same denoising stage.
[0015] Furthermore, the method for generating the boundary protection weight map includes: extracting the edge region of the target clothing based on the edge response of the target clothing attention weight map; generating a material boundary response within the target clothing edge region based on the clothing material lexical; fusing the target clothing edge region and the material boundary response into the boundary protection weight map; and suppressing the non-target replacement weight map based on the boundary protection weight map when replacing the latent space features, so that the solid color background latent variable does not replace the latent space features corresponding to the target clothing edge region.
[0016] Furthermore, the diffusion-based image editing network generates the target clothing attention weight map and the non-target replacement weight map through a clothing splitting adaptation layer. The clothing splitting adaptation layer is obtained as follows: acquiring sample character clothing images, sample target splitting instructions, and sample solid-color background clothing images; generating sample target clothing semantic vectors and sample non-target exclusion semantic vectors based on the sample target splitting instructions; setting a trainable adaptation matrix in the attention layer of the diffusion-based image editing network; updating the trainable adaptation matrix based on target clothing redrawing constraints, non-target collapse constraints, and boundary protection constraints to obtain the clothing splitting adaptation layer.
[0017] Furthermore, the extraction and writing back of the residual non-target features includes: decoding the target clothing latent variable into a candidate clothing image; re-encoding the candidate clothing image into a candidate image latent variable; extracting residual association features from the candidate image latent variable based on the non-target exclusion semantic vector; encoding the residual association features into the residual non-target features; and writing the residual non-target features as exclusion conditions into the non-target replacement weight map, so that the updated non-target replacement weight map participates in the generation of the target clothing latent variable.
[0018] Furthermore, a method for decomposing clothing using semantic collapse without masking also includes: determining adjacent interference regions in the edge region of the target clothing, performing latent space difference verification on the adjacent interference regions based on the semantic vector of the target clothing and the semantic vector of the non-target exclusion, and correcting the boundary protection weight map and the non-target replacement weight map based on the verification results.
[0019] A maskless semantic collapse clothing segmentation system, the system includes:
[0020] Instruction parsing module: acquires the original image containing the character's clothing and the target segmentation instruction, performs clothing lexical parsing on the target segmentation instruction to obtain the target clothing semantic vector, and generates a non-target exclusion semantic vector based on the target clothing semantic vector;
[0021] Latent domain mapping module: Encodes the original image into image latent variables, inputs the target clothing semantic vector and the non-target exclusion semantic vector into a diffusion image editing network, and generates a target clothing attention weight map and a non-target replacement weight map;
[0022] Occlusion association module: Based on the target clothing attention weight map, determine the first visible clothing region and the second visible clothing region in the latent variables of the image, and generate occlusion association edges;
[0023] Latent domain redrawing module: Collaboratively denoise and redraw the first visible clothing area and the second visible clothing area based on the occlusion association edge, and replace the latent space features of the human body, background and non-target clothing with solid color background latent variables based on the non-target replacement weight map;
[0024] Boundary suppression module: Generates a boundary protection weight map from the edge response of the target clothing attention weight map, and suppresses the replacement response of the non-target replacement weight map to the edge region of the target clothing according to the boundary protection weight map, thereby obtaining the latent variables of the target clothing;
[0025] Residual correction module: decodes the latent variables of the target clothing and performs semantic back-encoding, extracts residual non-target features, writes the residual non-target features back to the non-target replacement weight map, and outputs a solid-color background clothing split image.
[0026] The technical effects and advantages of the maskless semantic collapse clothing splitting method provided by this invention are as follows:
[0027] This invention achieves complete segmentation of complex clothing and solid color processing of non-target areas through maskless semantic collapse and latent space collaborative redrawing. It generates target clothing semantic vectors and non-target exclusion semantic vectors through target segmentation instructions, without relying on fixed segmentation categories or manual masks, thus responding to fine-grained clothing segmentation needs. Through target clothing attention weight maps, occlusion-related edges, and collaborative denoising redrawing, it reconstructs occluded or scattered areas of the same clothing according to the same semantics, reducing omissions and breaks in segmentation. Through non-target replacement weight maps and solid color background latent variables, it guides the human body, background, and non-target clothing to a unified background representation. Through boundary protection weight maps and residual non-target feature rewriting, it reduces the possibility of clothing edges being mistakenly replaced or non-target content remaining. Attached Figure Description
[0028] Figure 1 This is a flowchart of a method for splitting clothing without masking semantic collapse, as shown in Example 1.
[0029] Figure 2 This is a flowchart of the latent space differential verification method for adjacent interference regions in Example 2;
[0030] Figure 3 This is a schematic diagram of a maskless semantic collapse clothing splitting system in Embodiment 3. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1:
[0032] Please see Figure 1 As shown, an embodiment of the present invention provides a method for decomposing clothing using semantic collapse without a mask, the method comprising:
[0033] S1. Obtain the original image containing the character's clothing and the target segmentation instruction, perform clothing lexical parsing on the target segmentation instruction to obtain the target clothing semantic vector, and generate a non-target exclusion semantic vector based on the target clothing semantic vector.
[0034] S2. Encode the original image into image latent variables, input the target clothing semantic vector and the non-target exclusion semantic vector into a diffusion image editing network, and generate a target clothing attention weight map and a non-target replacement weight map.
[0035] S3. Based on the target clothing attention weight map, determine the first visible clothing region and the second visible clothing region in the latent variables of the image, and generate occlusion association edges;
[0036] S4. Based on the occlusion association edge, the first visible clothing area and the second visible clothing area are collaboratively denoised and redrawn. Based on the non-target replacement weight map, the latent space features of the human body, background and non-target clothing are replaced with solid color background latent variables.
[0037] S5. Generate a boundary protection weight map from the edge response of the target clothing attention weight map, and suppress the replacement response of the non-target replacement weight map to the edge region of the target clothing according to the boundary protection weight map to obtain the latent variables of the target clothing.
[0038] S6. Decode the latent variables of the target clothing and semantically encode them back, extract the residual non-target features, write the residual non-target features back to the non-target replacement weight map, and output the solid color background clothing split image.
[0039] In this embodiment, clothing lexical parsing is used to convert the user-inputted target splitting instructions into semantic input that can be invoked by the diffusion image editing network. Specifically, after receiving the target splitting instructions, the system first normalizes the action expressions, conjunctions, and irrelevant modifiers in them, retaining the semantic fragments used to define the clothing to be split. Action expressions include instructional words such as "split out," "extract," and "generate separately." These words are only used to confirm that the current task is a clothing splitting task and are not used as the main content of the target clothing semantic vector.
[0040] After normalization, the system performs sequential matching and dependency binding on semantic fragments based on pre-set clothing terminology, clothing part terminology, and clothing material terminology to obtain clothing name terminology, clothing part terminology, and clothing material terminology. Clothing name terminology is used to determine the main category of the clothing to be split, such as skirt, top, shawl, and coat; clothing part terminology is used to determine the local position or structural range within the main category, such as cuff, collar, hem, and waist; clothing material terminology is used to determine the edge shape or visual texture, such as tulle, lace, tassel, and leather. These terminology do not participate in encoding in isolation, but are bound according to the relationship of "clothing name terminology as the center, clothing part terminology limiting clothing name terminology, and clothing material terminology limiting clothing name terminology or clothing part terminology".
[0041] The system establishes a primary reference term based on clothing name terms, which specifies the clothing object to be retained in this split; it establishes a local constraint term based on clothing part terms, which specifies the local areas within the clothing object that should be retained first; and it establishes a boundary constraint term based on clothing material terms, which specifies the texture, translucency, or fine contour features that should be retained at the edges of the clothing object. Subsequently, the system concatenates the primary reference term, local constraint term, and boundary constraint term in that order, and inputs the concatenation result into a text encoder to generate a semantic vector of the target clothing. The purpose of this order is to first determine the main body of the clothing, and then constrain it with local areas and boundary shapes, so as to avoid material terms or part terms being misinterpreted as background decorations or human appendages because they are detached from the main body of the clothing.
[0042] After generating the target clothing semantic vector, the system further generates a non-target exclusion semantic vector based on the target clothing semantic vector. Specifically, the system marks the clothing object corresponding to the target clothing semantic vector as the retained object, and marks the skin, hair, limbs, face, background objects, scene objects, and other clothing objects not covered by the target clothing semantic vector as excluded objects. Then, it generates exclusion tokens corresponding to the human body region, background region, and non-target clothing region, and encodes the exclusion tokens into non-target exclusion semantic vectors. In subsequent steps, the non-target exclusion semantic vectors correspond to the non-target replacement weight map and are used to limit the latent space features of the human body, background, and non-target clothing to the replacement path of the latent variable of the solid color background.
[0043] For example: when the target splitting instruction is "split out the outer tulle skirt of the person", the system parses "skirt" or "skirt hem" into clothing name units and clothing part units, and parses "outer layer" and "tulle" into clothing material units and boundary limiting content, and generates a target clothing semantic vector accordingly; at the same time, the person's body, inner clothing, background environment and clothing parts that do not belong to the outer tulle skirt hem are encoded into non-target exclusion semantic vectors. Thus, subsequent steps can retain the skirt hem area that is occluded or scattered based on the same target clothing semantic vector, and exclude the human body and other clothing areas based on the non-target exclusion semantic vectors.
[0044] In this embodiment, the diffusion-based image editing network includes an attention layer for receiving image latent variables and text conditions. The image latent variables are the latent space representations formed after the original image is compressed by the encoder. They still retain the arrangement relationship corresponding to the spatial position of the original image. In order to facilitate matching with the text conditions, the system expands the image latent variables into a sequence of visual lexical units according to their spatial positions. Each visual lexical unit corresponds to a local latent space feature in the image latent variables, which is used to characterize the color, contour, texture and semantic information of the corresponding position in the original image.
[0045] After obtaining the semantic vectors of the target clothing and the exclusion semantic vectors, the system performs conditional mapping on the two, converting them into a first conditional vector and a second conditional vector that can be input into the attention layer. The first conditional vector is used to express the semantic orientation of the clothing to be retained, such as the clothing name, clothing part, and material boundary; the second conditional vector is used to express the exclusion objects that do not belong to the clothing to be retained, including the human body region, background region, and non-target clothing region. The conditional mapping can be completed by the linear mapping layer or conditional adaptation layer after the text encoder. Its function is to make the text side vectors and the image side visual lexical units in a computable matching relationship.
[0046] In the attention layer, the system uses the first condition vector as the target query condition and performs correlation calculation with the visual word sequence to obtain the response value of each visual word relative to the clothing to be segmented. Then, the response value is restored according to the spatial arrangement relationship of the latent variables of the image to form the target clothing attention weight map. The target clothing attention weight map is not a manually labeled clothing mask, but a continuous weight result obtained by matching the first condition vector and the visual word sequence in the latent space. It is used to indicate the clothing feature regions that should be retained and reconstructed first during subsequent denoising and redrawing.
[0047] Accordingly, the system uses the second condition vector as the exclusion query condition and performs correlation calculation with the visual word sequence to obtain the response value of each visual word relative to the excluded object. Then, the response value is restored according to the spatial arrangement relationship of the image latent variables to form a non-target replacement weight map. The non-target replacement weight map is used to indicate the latent spatial features corresponding to the human body, background and non-target clothing in the replacement path of the solid background latent variable in subsequent steps. Since the non-target replacement weight map comes from the matching result of the exclusion semantic condition and the visual word sequence, it does not need to generate a pixel-level human body segmentation map or clothing mask in advance.
[0048] For example: when the target splitting instruction is "split out the outer tulle skirt", the system converts the target clothing semantic vector corresponding to "outer tulle skirt" into a first conditional vector, and accordingly enhances the response of the skirt, tulle texture and its edge position in the visual word sequence; at the same time, it converts the non-target exclusion semantic vectors corresponding to the human body, inner clothing and background into a second conditional vector, and accordingly enhances the replacement response of the above non-target areas. The resulting target clothing attention weight map and non-target replacement weight map are respectively entered into the subsequent collaborative denoising and redrawing and solid color background replacement processes.
[0049] In this embodiment, the occlusion association edge is used to represent the latent space relationship formed in the original image by the occlusion of the same garment to be split due to the human body, background objects or non-target garments. The occlusion association edge is not a pixel line or a pre-annotated garment mask, but a denoising guidance information formed after corresponding analysis between the target garment attention weight map and the non-target replacement weight map.
[0050] Specifically, the system first reads the response distribution corresponding to the semantic vector of the target clothing in the attention weight map of the target clothing. If the response distribution presents a first response region and a second response region that are separated from each other in the spatial arrangement of the latent variables of the image, then the first response region is determined as the first visible clothing region and the second response region is determined as the second visible clothing region. The first visible clothing region and the second visible clothing region both originate from the same semantic vector of the target clothing. Therefore, they belong to the same clothing to be separated semantically, but they may be separated by arms, hair, bag straps, coats or background objects in the image space.
[0051] Subsequently, the system reads the response distribution between the first visible clothing area and the second visible clothing area in the non-target replacement weight map. When the response distribution corresponds to the human body area, the background area, or the non-target clothing area, it is identified as the occlusion response area. The occlusion response area is used to explain that the disconnect between the first visible clothing area and the second visible clothing area is not a semantic interruption of the target clothing, but is caused by the occlusion of the non-target object.
[0052] After determining the occlusion response region, the system establishes an occlusion association edge based on the relative positional relationship between the first visible clothing region, the second visible clothing region, and the occlusion response region. The occlusion association edge carries the association information that the first visible clothing region and the second visible clothing region share the same target clothing semantic vector, and records that there is a non-target occlusion response between them. After the occlusion association edge is written into the denoising path of the diffusion image editing network, it is used in the subsequent denoising process to constrain the first visible clothing region and the second visible clothing region to be redrawn collaboratively according to the same clothing semantics. At the same time, it makes the occlusion response region continue to be constrained by the non-target replacement weight map and enter the replacement path of the solid color background latent variable.
[0053] For example: In the original image, when the skirt is divided into a left skirt region and a right skirt region by the person's arm, the target clothing attention weight map responds in the left skirt region and the right skirt region respectively, and the non-target replacement weight map responds in the arm region between the two. Based on this, the system regards the left skirt region as the first visible clothing region, the right skirt region as the second visible clothing region, and the arm region as the occlusion response region, and generates an occlusion association edge connecting the two skirt regions. Thus, the subsequent denoising and redrawing will not mistake the two skirts for two unrelated objects, nor will it mistakenly retain the middle arm as part of the target clothing.
[0054] In this embodiment, after generating the occlusion association edge, the system does not directly stitch the first visible clothing area and the second visible clothing area in the pixel space. Instead, during the denoising process of the diffusion image editing network, the system imposes conditional constraints on the latent space features corresponding to the two. The first region redrawing condition and the second region redrawing condition are both generated by the occlusion association edge, which are used to indicate that the first visible clothing area and the second visible clothing area belong to the same clothing to be split, and the spatial disconnect between the two is caused by the occlusion response area.
[0055] Specifically, the system reads the latent space feature positions corresponding to the first visible clothing area and the second visible clothing area based on the occlusion association edges, and binds the target clothing semantic vector to the two positions respectively to form the first region redrawing condition and the second region redrawing condition. The first region redrawing condition is used to constrain the first visible clothing area to maintain the outline, texture and material orientation of the target clothing during the denoising process; the second region redrawing condition is used to constrain the second visible clothing area to be redrawn according to the same target clothing semantics. The only difference between the two is that the corresponding latent space feature positions are different, rather than corresponding to different clothing objects.
[0056] During the denoising stage, the system writes the first region redrawing condition into the latent space feature corresponding to the first visible clothing region, and writes the second region redrawing condition into the latent space feature corresponding to the second visible clothing region. Writing means that in the conditional denoising process of the diffusion image editing network, the region redrawing condition is used as the conditional input of the latent space feature of the region, so that the region is constrained by the semantic vector of the target clothing in each round of denoising update. Since the first region redrawing condition and the second region redrawing condition share the same semantic vector of the target clothing, the first visible clothing region and the second visible clothing region will not be redrawn as unrelated clothing fragments respectively.
[0057] For example: When the target splitting instruction is "split out the outer tulle skirt", and the skirt is obscured by the arm into a left skirt area and a right skirt area, the system writes the redrawing conditions corresponding to "outer tulle skirt" to the left skirt area and the right skirt area respectively, and uses the same target clothing semantic vector to constrain the two in the same denoising stage. The middle arm area is still controlled by the non-target replacement weight map and does not participate in the skirt redrawing path. Thus, the left skirt area and the right skirt area can complete collaborative redrawing according to the same clothing semantics.
[0058] In this embodiment, the boundary protection weight map is used to preserve the fine contours, translucent textures, and material transition information at the edge of the target clothing during non-target replacement. The boundary protection weight map is not a pre-generated pixel-level mask, but rather latent space constraint information obtained jointly from the edge response of the target clothing attention weight map and the clothing material lexical.
[0059] Specifically, after obtaining the attention weight map of the target garment, the system extracts the edge region of the target garment along the area where the response changes significantly in the weight map. The edge region of the target garment is used to represent the position where the target garment features transition from strong response to weak response. It usually corresponds to the edge of the skirt, the edge of the cuff, the edge of the neckline, the end of the tassel, the edge of the lace cutout, or the outer contour of the tulle. The edge region of the target garment still corresponds to the spatial arrangement relationship of the latent variables in the image, so the corresponding latent space features can be located during subsequent latent space replacement.
[0060] After extracting the edge region of the target garment, the system reads the garment material terms obtained from the aforementioned garment terminology parsing and generates a material boundary response within the edge region of the target garment based on the garment material terms. The material boundary response is used to distinguish the visual forms that different materials should retain at the boundary. For example, tulle corresponds to light transmission and soft transition features, lace corresponds to openwork and fine texture features, and tassels correspond to slender and discontinuous outline features. The system merges the edge region of the target garment with the material boundary response to obtain a boundary protection weight map, which includes both the spatial boundary position of the target garment and the boundary retention tendency related to the material.
[0061] When replacing the latent space features of the human body, background, and non-target clothing with solid background latent variables based on the non-target replacement weight map, the system simultaneously reads the boundary protection weight map. For latent space features that are located in the edge area of the target clothing and have material boundary response, the system reduces the replacement effect of the non-target replacement weight map at that location, so that the solid background latent variables do not directly cover the latent space features corresponding to the edge area of the target clothing. For non-target areas not covered by the boundary protection weight map, the system still follows the non-target replacement weight map to enter the replacement path of solid background latent variables.
[0062] For example: When the target splitting instruction is "split out the outer tulle skirt", the system extracts the outer contour of the skirt according to the target clothing attention weight map, and generates a material boundary response at the outer contour according to the clothing material term "tulle". When non-target replacement is performed later, the human body and background area are replaced with solid color background latent variables, while the light transmission transition and thin contour at the edge of the tulle skirt are constrained by the boundary protection weight map and are not collapsed as background.
[0063] In this embodiment, the clothing splitting adaptation layer is set in the attention layer of the diffusion image editing network. It is used to adjust the attention mapping relationship between the target splitting instruction and the latent variables of the image. The clothing splitting adaptation layer is not an independent segmentation model, nor does it output a pixel-level clothing mask. Instead, it changes the matching weights between the target clothing semantic vector, the non-target exclusion semantic vector and the visual word sequence in the attention layer through a trainable adaptation matrix, so that the network can generate a target clothing attention weight map and a non-target replacement weight map during subsequent inference.
[0064] When training the clothing splitting adaptation layer, the system obtains sample human clothing images, sample target splitting instructions, and sample solid-color background clothing images as training groups. The sample human clothing images are used to provide input content including human body, background, target clothing, and non-target clothing; the sample target splitting instructions are used to indicate which clothing objects need to be retained in the training group; and the sample solid-color background clothing images are used to provide the corresponding target output, where the target clothing is retained, and the human body, background, and non-target clothing are processed into solid-color backgrounds.
[0065] The system processes the sample target segmentation instruction according to the aforementioned clothing lexical parsing method, obtaining the sample target clothing semantic vector and the sample non-target exclusion semantic vector. Subsequently, the sample person clothing image is encoded as a sample image latent variable, and the sample target clothing semantic vector and the sample non-target exclusion semantic vector are input into the attention layer of the diffusion image editing network. A trainable adaptation matrix is set in the attention layer to participate in the attention calculation of text conditions and visual lexical sequences. During training, the basic generation structure of the diffusion image editing network remains unchanged, and the trainable adaptation matrix is updated to obtain the clothing segmentation adaptation layer.
[0066] The update of the trainable adaptation matrix is based on target clothing redrawing constraints, non-target collapse constraints, and boundary protection constraints. Target clothing redrawing constraints are used to ensure that the visual units corresponding to the semantic vectors of the target clothing in the sample retain the outline, texture, and material features of the target clothing during the denoising process. Non-target collapse constraints are used to make the human body, background, and non-target clothing features corresponding to the non-target excluded semantic vectors of the sample tend to be expressed as a solid color background in the sample solid color background clothing image. Boundary protection constraints are used to ensure that the material transitions, hollow textures, or fragmented contours at the edges of the target clothing are not covered by the non-target collapse path. All of the above constraints apply to the update process of the trainable adaptation matrix, without the need to introduce an additional human body parsing model or manually labeled masks.
[0067] After training is completed, the training adaptation matrix can be solidified into a clothing segmentation adaptation layer. During inference, when a new original image and target segmentation instruction are input, the clothing segmentation adaptation layer participates in the attention layer calculation, so that the target clothing semantic vector responds more concentratedly to the clothing region to be retained, and the non-target exclusion semantic vector responds more concentratedly to the human body, background and non-target clothing regions, thereby generating a target clothing attention weight map and a non-target replacement weight map.
[0068] In this embodiment, after obtaining the latent variables of the target clothing, the system first inputs the latent variables of the target clothing into the decoder to generate candidate clothing images. The candidate clothing images are used to represent the intermediate output results after the target clothing is redrawn and the non-target clothing is collapsed. They may still have a small number of human body, background or non-target clothing features remaining at the clothing edge, occlusion boundary or semi-transparent material neighborhood. In order to avoid directly using the candidate result as the final output, the system performs semantic verification on the candidate clothing images.
[0069] Specifically, the system uses an encoder corresponding to the original image encoding process to re-encode the candidate clothing image into candidate image latent variables. The candidate image latent variables have the same latent space representation as the image latent variables. This latent space is used to check whether there are still features in the candidate clothing image that are related to the non-target exclusion semantic vector. The re-encoding is not to re-perform clothing segmentation, nor to generate a new pixel-level mask, but to convert the candidate clothing image back into a latent space representation that can be matched with the non-target exclusion semantic vector.
[0070] Subsequently, the system uses the non-target exclusion semantic vector as the exclusion query condition to perform association extraction in the latent variables of the candidate image. When there is a correspondence between the local latent space features in the latent variables of the candidate image and the non-target exclusion semantic vector, the local latent space features are identified as residual association features. The residual association features are used to represent non-target content in the candidate clothing image that has not been fully replaced by the solid color background latent variable, such as human body edge afterimages, background texture residues, or non-target clothing fragments.
[0071] The system encodes residual associated features into residual non-target features after conditional mapping, and writes the residual non-target features into the non-target replacement weight map as exclusion conditions. Writing means enhancing the replacement response at the corresponding position of the residual associated features in the non-target replacement weight map, so that the position continues to enter the replacement path of the solid background latent variable when the target clothing latent variable is generated in the future. For the target clothing edge region jointly defined by the target clothing attention weight map and the boundary protection weight map, the system retains its target clothing redrawing path to avoid replacing the target clothing edges such as tulle, lace, and tassels as non-target residues.
[0072] In one alternative implementation, to unify the description of target clothing redrawing, occlusion-coordinated redrawing, non-target collapse replacement, and residual non-target feature write-back in the same denoising stage, the joint update is crucial.
[0073] The following expression can be used to update the latent variables of the target clothing:
[0074] ;
[0075] In the formula, For the first The target clothing hidden variable after the round of updates, For the first Image latent variables in the round denoising stage For the target clothing semantic vector, This is a semantic vector for non-target clothing. Attention weighting chart for target clothing Replace the non-target weight map. The occlusion association weight is formed by the occlusion association edges. For boundary protection weight map, For a solid color background, there is a hidden variable. For residual non-target features, To redraw the target based on the semantic vector of clothing, For collaborative redraw mapping based on occlusion-related edges, For exclusion mapping based on non-target exclusion semantic vectors, For element-wise multiplication, where It is a matrix of all ones with the same dimension as the weight graph. , , , These are adjustment coefficients corresponding to the intensity of each action. For the denoising round index, the target clothing attention weight map corresponds to the weight result generated by the target clothing semantic vector, the non-target replacement weight map corresponds to the replacement response generated by the non-target exclusion semantic vector, the occlusion association weight corresponds to the conditional constraint of the occlusion association edge in the denoising path, the boundary protection weight map corresponds to the protection constraint formed based on the edge region and material boundary response of the target clothing, and the residual non-target features correspond to the residual association information extracted and written back from the latent variables of the candidate image. This expression is used to describe the combination relationship of each latent space feature in a denoising update, without limiting the specific number of network layers, the number of denoising rounds, or the value of the adjustment coefficient.
[0076] For example: when the outer tulle skirt has been preserved in the candidate clothing image, but there is still residual arm color on the side of the skirt, the system re-encodes the candidate clothing image into candidate image latent variables, and extracts residual association features related to the arm based on the non-target exclusion semantic vector; then, the residual association features are encoded into residual non-target features and written into the corresponding position in the non-target replacement weight map, so that when the target clothing latent variable is generated again, this position is guided to the replacement path of the solid color background latent variable. Thus, the final output solid color background clothing split image can preserve the target clothing and reduce the residue of the human body, background and non-target clothing. Example 2:
[0077] like Figure 2 As shown, this embodiment further improves upon the design of Embodiment 1. The difference lies in the fact that, in the actual operation of Embodiment 1, it was found that when the target garment and non-target garments are similar in color, material, or edge texture, the attention weight map of the target garment may generate local responses at adjacent non-target garment locations. This causes edge fragments of the non-target garments to be mistakenly incorporated into the redrawing path of the target garment, failing to adequately exclude adjacent interference areas of the same color or material while preserving the edge details of the target garment. Based on this, a maskless semantic collapse garment splitting method further includes a latent space difference verification step for adjacent interference areas.
[0078] Specifically, after generating the target clothing attention weight map and the non-target replacement weight map, the system reads the low-stability response portion of the target clothing attention weight map that is adjacent to the edge region of the target clothing, and reads the exclusion response portion of the non-target replacement weight map that intersects with or is adjacent to the low-stability response portion. The two are jointly identified as adjacent interference regions. The adjacent interference regions are used to represent the regions in the latent space that are simultaneously affected by the semantic vector of the target clothing and the semantic vector of the non-target exclusion. They usually correspond to the inner layer of the same color, the edge of the adjacent coat, the waist belt of similar material, or the background texture close to the target clothing.
[0079] After identifying adjacent interference regions, the system uses the target clothing semantic vector and the non-target exclusion semantic vector to perform correlation verification on visual words in the adjacent interference regions. For visual words that better match the target clothing semantic vector and have an occlusion association with the first visible clothing region or the second visible clothing region, the system marks them as target homologous features. For visual words that better match the non-target exclusion semantic vector and do not form a shared target clothing semantic with the occlusion association, the system marks them as interference exclusion features. Target homologous features are used to enter the collaborative denoising and redrawing path, while interference exclusion features are used to enter the replacement path for the solid color background latent variable.
[0080] Subsequently, the system corrects the boundary protection weight map based on the target homology features, so that the boundary protection weight map only maintains the protection effect in the edge area of the target clothing and its target homology features; at the same time, it enhances the non-target replacement weight map based on the interference exclusion features, so that adjacent non-target clothing or background textures are not retained due to similar colors or materials. The above processing is completed in the latent space corresponding to the latent variables of the image, without the need to pre-generate pixel-level clothing masks or call the human body analysis model to determine the clothing category boundary.
[0081] For example: When the target splitting instruction is "split out the outer tulle skirt", and there is an inner skirt with a similar color next to the outer tulle skirt, Implementation 1 may generate target response and non-target response at the adjacent edges of the two. In this implementation, the part near the edge is determined as the adjacent interference region, and latent space difference verification is performed according to the target clothing semantic vector corresponding to the "outer tulle skirt" and the non-target exclusion semantic vector corresponding to the inner clothing. If a local feature is consistent with the occlusion association edge and material boundary response of the outer tulle skirt, its target homology feature is retained. If a local feature only matches the inner skirt, it is written as an interference exclusion feature into the non-target replacement weight map. Thus, while maintaining the edge transition of the outer tulle, the situation where the inner clothing is mistakenly retained is reduced. Example 3:
[0082] like Figure 3As shown, based on the same inventive concept as the maskless semantic collapse clothing segmentation method in the foregoing embodiments, this application provides a maskless semantic collapse clothing segmentation system. The system and method embodiments in this application are based on the same inventive concept. The system includes:
[0083] Instruction parsing module: acquires the original image containing the character's clothing and the target segmentation instruction, performs clothing lexical parsing on the target segmentation instruction to obtain the target clothing semantic vector, and generates a non-target exclusion semantic vector based on the target clothing semantic vector;
[0084] Latent domain mapping module: Encodes the original image into image latent variables, inputs the target clothing semantic vector and the non-target exclusion semantic vector into a diffusion image editing network, and generates a target clothing attention weight map and a non-target replacement weight map;
[0085] Occlusion association module: Based on the target clothing attention weight map, determine the first visible clothing region and the second visible clothing region in the latent variables of the image, and generate occlusion association edges;
[0086] Latent domain redrawing module: Collaboratively denoise and redraw the first visible clothing area and the second visible clothing area based on the occlusion association edge, and replace the latent space features of the human body, background and non-target clothing with solid color background latent variables based on the non-target replacement weight map;
[0087] Boundary suppression module: Generates a boundary protection weight map from the edge response of the target clothing attention weight map, and suppresses the replacement response of the non-target replacement weight map to the edge region of the target clothing according to the boundary protection weight map, thereby obtaining the latent variables of the target clothing;
[0088] Residual correction module: decodes the latent variables of the target clothing and performs semantic back-encoding, extracts residual non-target features, writes the residual non-target features back to the non-target replacement weight map, and outputs a solid-color background clothing split image.
[0089] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0090] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present application, based on the technical solution and concept of the present application, should be covered within the scope of protection of the present application.
Claims
1. A method for decomposing clothing using semantic collapse without masking, characterized in that, The methods include: Obtain the original image containing the character's clothing and the target segmentation instruction, perform clothing lexical parsing on the target segmentation instruction to obtain the target clothing semantic vector, and generate a non-target exclusion semantic vector based on the target clothing semantic vector; The original image is encoded as image latent variables, and the semantic vector of the target clothing and the semantic vector of the non-target exclusion are input into a diffusion image editing network to generate a target clothing attention weight map and a non-target replacement weight map. Based on the target clothing attention weight map, the first visible clothing region and the second visible clothing region are determined in the latent variables of the image, and occlusion association edges are generated; Based on the occlusion association edge, the first visible clothing area and the second visible clothing area are collaboratively denoised and redrawn. Based on the non-target replacement weight map, the latent space features of the human body, background and non-target clothing are replaced with solid color background latent variables. A boundary protection weight map is generated from the edge response of the target clothing attention weight map, and the replacement response of the non-target replacement weight map to the edge region of the target clothing is suppressed according to the boundary protection weight map to obtain the latent variables of the target clothing. The latent variables of the target clothing are decoded and semantically back-encoded to extract residual non-target features. The residual non-target features are then written back to the non-target replacement weight map, and the resulting solid-color background clothing split image is output.
2. The method for decomposing clothing using semantic collapse without masking, as described in claim 1, is characterized in that... The clothing lexical parsing includes: extracting clothing name lexicals, clothing part lexicals, and clothing material lexicals from the target splitting instruction; establishing a main pointer term based on the clothing name lexicals, establishing a local limiting term based on the clothing part lexicals, and establishing a boundary limiting term based on the clothing material lexicals; concatenating and encoding the main pointer term, the local limiting term, and the boundary limiting term into the target clothing semantic vector; and generating a non-target exclusion semantic vector corresponding to the human body region, the background region, and the non-target clothing region based on the target clothing semantic vector.
3. The method for decomposing clothing using semantic collapse without masking, as described in claim 1, is characterized in that... The diffusion-based image editing network includes an attention layer; the method for generating the target clothing attention weight map and the non-target replacement weight map includes: converting the image latent variables into a visual word sequence, converting the target clothing semantic vector into a first conditional vector, and converting the non-target exclusion semantic vector into a second conditional vector; querying the visual word sequence with the first conditional vector in the attention layer to obtain the target clothing attention weight map; and querying the visual word sequence with the second conditional vector to obtain the non-target replacement weight map.
4. The method for decomposing clothing using semantic collapse without masking as described in claim 1, characterized in that, The method for generating occlusion association edges includes: extracting a first response region as the first visible clothing region and extracting a second response region as the second visible clothing region from the target clothing attention weight map; extracting an occlusion response region located between the first visible clothing region and the second visible clothing region from the non-target replacement weight map; establishing the occlusion association edge connecting the first visible clothing region and the second visible clothing region based on the occlusion response region, and writing the occlusion association edge into the denoising path of the diffusion image editing network.
5. The method for decomposing clothing using semantic collapse without masking, as described in claim 4, is characterized in that... Collaborative denoising and redrawing of the first visible clothing region and the second visible clothing region includes: generating a first region redrawing condition and a second region redrawing condition based on the occlusion association edge; writing the first region redrawing condition into the latent space feature corresponding to the first visible clothing region, and writing the second region redrawing condition into the latent space feature corresponding to the second visible clothing region; and maintaining the first region redrawing condition and the second region redrawing condition sharing the target clothing semantic vector within the same denoising stage.
6. The method for decomposing clothing using semantic collapse without masking, as described in claim 1, is characterized in that... The method for generating the boundary protection weight map includes: extracting the edge region of the target clothing based on the edge response of the target clothing attention weight map; generating a material boundary response within the target clothing edge region based on the clothing material terms; fusing the target clothing edge region and the material boundary response into the boundary protection weight map; and suppressing the non-target replacement weight map based on the boundary protection weight map when replacing the latent space features, so that the solid color background latent variable does not replace the latent space features corresponding to the target clothing edge region.
7. The method for decomposing clothing using semantic collapse without masking, as described in claim 1, is characterized in that... The diffusion-based image editing network generates the target clothing attention weight map and the non-target replacement weight map through a clothing splitting adaptation layer. The clothing splitting adaptation layer is obtained as follows: acquiring sample character clothing images, sample target splitting instructions, and sample solid-color background clothing images; generating sample target clothing semantic vectors and sample non-target exclusion semantic vectors based on the sample target splitting instructions; setting a trainable adaptation matrix in the attention layer of the diffusion-based image editing network; updating the trainable adaptation matrix based on target clothing redrawing constraints, non-target collapse constraints, and boundary protection constraints to obtain the clothing splitting adaptation layer.
8. The method for decomposing clothing using semantic collapse without masking, as described in claim 1, is characterized in that... The extraction and writing back of residual non-target features includes: decoding the target clothing latent variable into candidate clothing images; re-encoding the candidate clothing images into candidate image latent variables; extracting residual association features from the candidate image latent variables based on the non-target exclusion semantic vector; encoding the residual association features into residual non-target features; and writing the residual non-target features as exclusion conditions into the non-target replacement weight map, so that the updated non-target replacement weight map participates in the generation of the target clothing latent variable.
9. The method for decomposing clothing using semantic collapse without masking according to claim 1, characterized in that, Also includes: Adjacent interference regions are identified in the edge region of the target clothing. Latent space difference verification is performed on the adjacent interference regions based on the semantic vector of the target clothing and the semantic vector of the non-target exclusion. The boundary protection weight map and the non-target replacement weight map are then corrected based on the verification results.
10. A maskless semantic collapse clothing splitting system, characterized in that, The system includes: Instruction parsing module: acquires the original image containing the character's clothing and the target segmentation instruction, performs clothing lexical parsing on the target segmentation instruction to obtain the target clothing semantic vector, and generates a non-target exclusion semantic vector based on the target clothing semantic vector; Latent domain mapping module: Encodes the original image into image latent variables, inputs the target clothing semantic vector and the non-target exclusion semantic vector into a diffusion image editing network, and generates a target clothing attention weight map and a non-target replacement weight map; Occlusion association module: Based on the target clothing attention weight map, determine the first visible clothing region and the second visible clothing region in the latent variables of the image, and generate occlusion association edges; Latent domain redrawing module: Collaboratively denoise and redraw the first visible clothing area and the second visible clothing area based on the occlusion association edge, and replace the latent space features of the human body, background and non-target clothing with solid color background latent variables based on the non-target replacement weight map; Boundary suppression module: Generates a boundary protection weight map from the edge response of the target clothing attention weight map, and suppresses the replacement response of the non-target replacement weight map to the edge region of the target clothing according to the boundary protection weight map, thereby obtaining the latent variables of the target clothing; Residual correction module: decodes the latent variables of the target clothing and performs semantic back-encoding, extracts residual non-target features, writes the residual non-target features back to the non-target replacement weight map, and outputs a solid-color background clothing split image.