Deep learning based automatic label generation method for clothing categories
By constructing a parameterized clothing pattern topology skeleton library and combining causal intervention calculation and dynamic arbitration, the problem of easily confusing visually similar but semantically different items in the automatic generation of clothing category labels using deep learning methods is solved. This improves the accuracy and applicability of label generation, adapts to complex scenarios, and meets the needs of e-commerce retail and digital management of clothing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YISHANG CHUANGZHAN (SHANGHAI) TECHNOLOGY CO LTD
- Filing Date
- 2026-02-14
- Publication Date
- 2026-05-29
AI Technical Summary
Existing deep learning methods are prone to confusing visually similar but semantically different clothing categories in automatic clothing category label generation, leading to label generation errors, especially in scenarios with complex backgrounds, varied clothing poses, or partial occlusion, where the accuracy is insufficient.
A parametric clothing pattern topology skeleton library is constructed. By combining skeleton fitting, texture evolution and perturbation comparison image generation with causal intervention calculation and dynamic arbitration, the accuracy and interpretability of label generation are improved.
It effectively distinguishes visually similar but semantically different clothing categories, improves the accuracy of tag generation, and achieves traceability and verifiability of results through a structured decision evidence chain, adapting to complex scenarios and improving the applicability of e-commerce retail and digital management of clothing.
Smart Images

Figure CN122116335A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital management of clothing, and more specifically, to a method for automatically generating clothing category labels based on deep learning. Background Technology
[0002] In e-commerce retail and digital management of apparel, automatic tag generation for apparel categories is one of the main methods to achieve efficient product retrieval, classification management, and accurate recommendations. Apparel category tags typically present a complex hierarchical structure, such as "clothing > tops > shirts > denim shirts." This hierarchical structure requires tag generation to accurately identify fine-grained categories of clothing, achieving precise positioning from top-level broad categories to bottom-level subcategories. With the rapid development of the e-commerce industry, the number of clothing images accumulated by platforms has grown exponentially. These images generally have problems such as complex backgrounds, varied clothing poses, and partial occlusion, which places higher demands on the accuracy and efficiency of automatic tag generation for apparel categories. Currently, automatic tag generation technology for apparel categories is mainly based on deep learning methods. By constructing deep neural network models and training them on massive amounts of clothing image data, the models learn visual features in the images, such as texture, color, and outline, thereby completing the classification and tag generation of apparel categories. However, current deep learning-based automatic tag generation technology for apparel categories struggles to effectively distinguish between clothing items with highly similar visual features but different category semantics, easily leading to semantic confusion and tag generation errors.
[0003] In the process of fine-grained classification of clothing, models often confuse categories with similar visual features but fundamental semantic differences. For example, they may misclassify a hoodie as a sweatshirt, an A-line skirt as an umbrella skirt, or a loafer as an Oxford shoe. In real-world applications with complex backgrounds, varied clothing poses, or partial occlusion, this semantic confusion is even more pronounced. Models often over-rely on local texture features, such as lace or denim textures, while ignoring factors that determine the essence of clothing categories, such as overall fit and cut. This leads to misclassification in label generation. The main reason is that models passively summarize statistical patterns from massive amounts of image pixels. The features they learn are essentially complex combinations of local visual patterns such as texture and color, rather than high-level, structured semantic concepts such as fit, cut intention, and category definition as understood by humans. Existing deep learning models lack topological prior knowledge about how clothing is constructed and cannot fundamentally understand the semantic connotation of clothing categories. They can only make classification judgments based on combinations of local visual features.
[0004] In e-commerce platforms, where massive amounts of clothing images with complex backgrounds and diverse poses are automatically and precisely labeled, existing technologies, in pursuit of pixel-level feature discrimination, often overemphasize the learning of local visual features. This inevitably leads to overfitting of the model to local texture noise, neglecting the overall pattern topology that determines the essence of clothing categories. As a result, the accuracy of automatic clothing category label generation is difficult to meet the needs of practical applications. Summary of the Invention
[0005] To address the problems existing in the prior art, the present invention aims to provide a method for automatically generating clothing category tags based on deep learning. This method can construct a parameterized clothing pattern topology skeleton library, perform causal intervention calculations and dynamic arbitration, effectively solving the problem that existing deep learning methods easily confuse visually similar but semantically different clothing categories, and improving the accuracy of automatic clothing category tag generation.
[0006] To solve the above problems, the present invention adopts the following technical solution: A deep learning-based method for automatically generating clothing category tags includes: Define a parametric clothing pattern topology skeleton library that covers the target product category. Each skeleton is a two-dimensional graph structure composed of key points and elastic constraint edges. The input clothing image is fitted with the topological skeleton of the clothing pattern, and the skeleton outline is matched with the clothing outline in the image by deformation to obtain the skeleton instance and deformation parameters. Perform graph similarity matching between the skeleton instance and the skeleton in the skeleton library to determine the matching category and output the deformation parameter vector; The skeleton of the matching product category is deformed based on the deformation parameter vector, and combined with the texture information extracted from the input image to generate a virtual try-on image; A perturbation skeleton is generated by applying a predetermined offset to the key points of the deformed skeleton and then combining the texture information to generate a set of perturbation comparison images. The topological features, texture features, and perturbation contrast images of the skeleton instance are input into a dual-channel decision network, which outputs category labels and decoupled confidence vectors.
[0007] Furthermore, a parametric garment pattern topology skeleton library covering the target product category is defined, including: Define a set of clothing structural primitives, each primitive encapsulating the biomechanical interaction logic associated with a specific human body part and the dynamic response rules of the fabric; Based on the target product category, relevant primitives are selected from the primitives. A dynamic topology network is formed through physical coupling negotiation between primitives. The stiffness distribution function of the connection points and the boundary deformation transfer coefficient are calculated to obtain the set of physical parameters of the target product category. The design semantics are obtained, and a modulation instruction sequence is generated based on the design semantics. The physical parameter set is then subjected to gradient modulation to generate the pattern topology skeleton of the target product category.
[0008] Furthermore, the construction of the pattern topology skeleton includes: Define a field node, which has a perceptual function that scans the surrounding image space to output a set of orientation fit vectors; Compare the directional consistency vector sets of two field nodes to find the consensus direction, generate virtual connection paths based on the consensus direction, and define a dynamic stiffness protocol for the path. A two-dimensional external potential field is generated based on the input image. Field nodes and virtual connection paths are placed in the two-dimensional external potential field. Dynamic equilibrium is achieved through iterative calculation, and the equilibrium two-dimensional graph structure is output.
[0009] Further, fitting the input clothing image to the topological skeleton of the clothing pattern includes: The orientation sensing function of the field node is activated to generate a detection signal, which generates a multi-level image response field of the input clothing image. The detection signal is coupled with the image response field to obtain an initial coupling strength vector set. Deformation intention is generated based on the initial coupling strength vector group. Local deformation proposal is calculated according to the deformation intention and dynamic stiffness protocol. After simulating the execution of the proposal, the new coupling strength is obtained. The image correction feedback signal is generated by comparison to obtain the proposal and feedback pair sequence. Based on the sequence of proposals and feedback, a stable consensus proposal set is found, the consensus proposal set is encoded into a deformation protocol, and the deformation protocol is replayed to drive the deformation of the topology skeleton, thus obtaining the skeleton instance and deformation parameters.
[0010] Further, the matching product category is determined and the deformation parameter vector is output, including: The deformation parameter vector is applied to each candidate skeleton in the clothing pattern topology skeleton library, and physical deformation and relaxation simulation is performed to obtain a set of relaxed candidate skeletons. Key response patterns are extracted from relaxed candidate skeletons, feature signs are analyzed from input clothing images, key response patterns are compared with feature signs, and causal correlation scores are calculated. The relaxed candidate skeleton is projected onto the preset design semantic constraint space for verification, and the deformation parameter vector is checked against the design intent of the candidate category. Figure One Consistency is determined by obtaining a semantic constraint conflict score. Candidates are selected based on causal correlation score and semantic constraint conflict score. The deformation parameters of the selected candidate skeletons are fine-tuned and calibrated to determine the best matching category and output the calibrated deformation parameter vector.
[0011] Furthermore, virtual try-on images are generated, including: Establish a continuous texture field for the input clothing area, analyze the texture field to assign it physical properties and record the original lighting reference; A geometric deformation field is generated based on the calibrated deformation parameter vector. The non-rigid response of the texture field under the action of the geometric deformation field is calculated based on the physical properties of the texture field, and the evolved texture field is obtained. A virtual lighting environment is established based on the original lighting reference. The light and shadow information of the deformed skeleton geometry is calculated. The color information of the evolved texture field is combined with the light and shadow information to generate a virtual try-on image.
[0012] Furthermore, a predetermined offset is applied to the key points of the deformed skeleton to generate a perturbed skeleton, including: The classification gradient of the dual-channel decision network is calculated based on virtual try-on images, and the gradient is inversely mapped into a key point displacement proposal field for the deformed skeleton. Identify the semantic structural vulnerabilities of the current category from the semantic constraint space of the category design, and combine the dynamic stiffness protocol to filter and strengthen the displacement suggestion field to generate a directional perturbation scheme. A structured perturbation sequence is designed based on the directional perturbation scheme, generating a perturbation skeleton and corresponding explanatory labels. A family of perturbation contrast images is then synthesized based on the evolved texture field.
[0013] Furthermore, a set of perturbation contrast images is generated by combining texture information, including: The geometric deformation of the perturbation skeleton is deconstructed into a local strain field applied to the texture field. Based on the physical properties of the texture field, its state change is predicted by a micromaterial response model to obtain the texture response prediction field. The preliminary results of the texture response prediction field and the perturbation skeleton are combined and input into a dual-channel decision network to obtain intermediate perceptual features. Based on the comparison with the reference features, the texture response prediction field and the perturbation skeleton are adjusted collaboratively and iteratively to obtain the optimized pairing. Based on the explanatory labels, the parameters of the virtual lighting environment are adjusted, and the representation of the micro-geometric details of the surface in the region with pattern density variation in the texture response prediction field is increased during the rendering process, and a perturbation contrast image is synthesized.
[0014] Furthermore, the output category labels and decoupling confidence vectors include: Analyze the explanatory labels corresponding to the perturbation skeleton and select the relevant perturbation image subsets. Reconstruct the counterfactual feature representation of the competing categories through feature inverse mapping. Establish a causal intervention calculation layer, which receives the topological features, texture features, counterfactual feature representations and semantic constraint conflict scores of skeleton instances, performs feature replacement interventions and calculates the causal effect matrix.
[0015] Furthermore, the output category labels and decoupled confidence vectors also include: A dynamic arbitrator is set up to check logical conflicts based on the causal effect matrix and semantic constraint conflict score, dynamically calibrate the original classification probability, and output the calibrated category label and the decoupling confidence vector based on the causal effect strength. Encapsulate a structured chain of decision-making evidence that includes key intermediate data and reasoning processes.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) This solution effectively solves the problem that existing deep learning methods are prone to confusing visually similar but semantically different clothing categories by constructing a parameterized clothing pattern topology skeleton library, performing causal intervention calculation and dynamic arbitration, and improves the accuracy of automatic clothing category label generation. It can accurately distinguish similar subcategories such as hoodies and sweatshirts, A-line skirts and umbrella skirts.
[0017] (2) This solution encapsulates the structured decision evidence chain, integrates the key intermediate data and reasoning process of the entire label generation process, and realizes the traceability, verifiability and interpretability of the classification results, meets the reproducibility requirements of patent technology, and facilitates subsequent objection verification of label results and optimization of technical solutions.
[0018] (3) This solution is designed through a multi-stage collaborative approach, including skeleton fitting, texture evolution, and perturbation contrast image generation. It can adapt to real-world application scenarios such as complex backgrounds, varied clothing poses, and partial occlusion, thereby reducing the impact of complex scenarios on the accuracy of label generation and improving the applicability of the technical solution in e-commerce retail, digital management of clothing, and other scenarios.
[0019] (4) This solution adopts a parametric and structured design approach. The constructed clothing pattern topology skeleton library can be flexibly adapted to different target clothing categories. The technical modules of each step are closely connected and highly reusable, which facilitates subsequent expansion to adapt to more clothing subcategories and further enhances the practicality and scalability of the technical solution. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0021] Figure 1 This is a flowchart of the method for automatically generating clothing category tags based on deep learning, as described in this invention. Detailed Implementation
[0022] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0023] Please see Figure 1 A method for automatically generating clothing category tags based on deep learning, which includes: Step 1: Define a parametric clothing pattern topology skeleton library covering the target product category. Each skeleton is a two-dimensional graph structure composed of key points and flexible constraint edges. The specific operations are as follows: Before constructing this skeleton library, it is necessary to define the scope of the target clothing categories, covering various types of clothing commonly used in e-commerce retail and digital clothing management, to ensure that the coverage of the skeleton library can meet the actual application needs. Each defined clothing pattern topology skeleton adopts a two-dimensional graph structure, which is composed of key points and elastic constraint edges. Key points are used to locate the key structural positions of the clothing, such as the collar endpoint, cuff endpoint, hem turning point, waist key point, etc. Each key point has adjustable parameter attributes to adapt to different sizes and patterns of the same type of clothing. Elastic constraint edges are used to connect adjacent key points. Their elastic characteristics can simulate the deformation ability of clothing fabric, allowing the skeleton to adapt to the contour of the actual clothing image, thereby achieving accurate matching with the clothing image. Through parametric design, the key point position and the stiffness of the elastic constraint edges of each skeleton can be quantitatively adjusted, ultimately forming a set of clothing pattern topology skeletons that include various target clothing categories, are flexibly deformable, and are parametrically controllable, i.e., the parametric clothing pattern topology skeleton library.
[0024] The definition of a parameterized clothing pattern topology skeleton library covering the target product category also includes the following steps: Step 11: Define a set of clothing structural primitives. Each primitive encapsulates the biomechanical interaction logic associated with a specific human body part and the dynamic response rules of the fabric. The specific operations are as follows: The basic units constituting the topological skeleton of a garment pattern are defined as garment structural primitives. Each garment structural primitive corresponds to a specific part of the human body, such as a shoulder and sleeve primitive corresponding to the shoulder, a body primitive corresponding to the torso, a waist primitive corresponding to the waist, and a cuff primitive corresponding to the cuff. This correspondence ensures that the primitives can accurately simulate the fit and interaction characteristics between the garment and the relevant parts of the human body. Within each structural primitive, two types of logical rules are encapsulated: biomechanical interaction logic and fabric dynamic response rules. The biomechanical interaction logic is used to simulate the mechanical action and mechanical response of the corresponding parts of the garment when the relevant parts of the human body move. For example, when the human arm swings, the shoulder and sleeve primitive is subjected to tension, torsional force, and corresponding deformation trend. The fabric dynamic response rules are used to simulate the physical properties of the garment fabric itself, including the elasticity, toughness, breathability, and wrinkle formation rules of the fabric. For example, the difference in deformation between denim and cotton fabrics under stress, and the difference in wrinkle recovery ability between silk and wool fabrics. By encapsulating these two types of rules in structural primitives, each primitive can independently simulate the physical behavior of clothing associated with the corresponding human body parts.
[0025] Step 12: Select relevant primitives from the primitives according to the target category, form a dynamic topology network through physical coupling negotiation between primitives, and calculate the stiffness distribution function and boundary deformation transfer coefficient of the connection points to obtain the set of physical parameters of the target category. The specific operations are as follows: Based on the clothing structure primitives defined in step 11, and combined with the requirements of the target clothing category, a dynamic topology network corresponding to this category is constructed, and a set of physical parameters supporting the generation of the pattern topology skeleton is obtained. First, according to the specific characteristics of the target clothing category, primitives related to it are selected from the defined clothing structure primitives. The selection logic is based on the structural composition of the target category. For example, when generating the pattern topology skeleton of a denim shirt, primitives related to the shirt structure, such as shoulder and sleeve primitives, torso primitives, collar primitives, and cuff primitives, need to be selected, while primitives unrelated to the shirt, such as skirt primitives and trouser leg primitives, are excluded. After selection, physical coupling negotiation is performed on these related primitives. Physical coupling negotiation refers to simulating the physical connection relationships and interaction laws between the primitives. For example... The connection methods between shoulder and sleeve primitives and torso primitives, the connection angle between collar primitives and torso primitives, and the force transmission relationship are all negotiated to form a dynamic topological network that can fully simulate the structure and physical characteristics of the target garment category. This network can reflect the dynamic interaction between the primitives. After the dynamic topological network is formed, the stiffness distribution function and boundary deformation transfer coefficient of the connection points are obtained through mechanical simulation and parameter calculation. The stiffness distribution function is used to describe the stiffness distribution of each primitive connection point in the topological network, reflecting the ability of the connection point to resist deformation. The boundary deformation transfer coefficient is used to describe how the deformation at the boundary of the primitive is transmitted between adjacent primitives. These two parameters and other relevant physical parameters together constitute the set of physical parameters of the target garment category.
[0026] Step 13: Obtain design semantics, generate modulation instruction sequences based on design semantics, perform gradient modulation on the physical parameter set, and generate the pattern topology skeleton of the target product category. The specific operations are as follows: Combining design semantics, the set of physical parameters obtained in step 12 is modulated to ultimately generate the pattern topology skeleton of the target clothing category. First, the design semantics of the target clothing category are obtained. The design semantics encompass information such as the design concept, pattern requirements, and style characteristics of the clothing category. For example, the design semantics of a casual shirt may include a loose fit, a lapel design, and straight cuffs, while the design semantics of a fitted dress may include a waist-cinching fit, an A-line skirt, and short sleeves. These design semantics can be obtained through design standards in the clothing industry, designer requirement inputs, etc. Based on the obtained design semantics, a corresponding modulation instruction sequence is generated through a semantic parsing algorithm. Each instruction in this sequence corresponds to a design semantic requirement, used to clarify the requirements for physical parameters. The adjustment direction and magnitude of specific parameters in the set are then determined. Gradient modulation is then used to apply the modulation instruction sequence to the physical parameter set. Gradient modulation enables smooth and precise adjustment of physical parameters, ensuring that the adjusted parameters accurately match the design semantic requirements. For example, based on the design semantics of the waist-cinching pattern, gradient modulation reduces the stiffness distribution value of the waist basic element connection points, allowing the waist to form a waist-cinching deformation trend. Based on the semantics of the lapel design, the physical parameters of the collar basic element are adjusted to make the collar form a folded pattern feature. Through the above gradient modulation process, the design semantics are transformed into physical parameter adjustment instructions for the pattern topology skeleton, ultimately generating a pattern topology skeleton that meets the design requirements of the target category and has a complete structure and physical characteristics.
[0027] The construction of the pattern topology skeleton also includes the following steps: Step 14: Define field nodes. Field nodes have a perceptual function that scans the surrounding image space to output a set of directional fit vectors. The specific operations are as follows: Field nodes are the basic sensing units for constructing the pattern topology skeleton. They are distributed within the spatial range corresponding to the garment image and are used to sense relevant feature information in the image space. Each field node is equipped with a dedicated sensing function, which can perform omnidirectional scanning of the image space around the field node. The scanning range can be adaptively adjusted according to the size, resolution, and target precision of the garment image. The function of the sensing function is to identify the directional information of the garment contour and texture in the surrounding image space, such as the direction of the garment edge and the extension direction of the texture, and quantify this directional information to output a directional matching vector set. The directional matching vector set consists of multiple vectors, each corresponding to a specific direction. The value of the vector is used to represent the degree of matching of the garment features perceived by the field node in that direction. The higher the value, the higher the degree of matching between that direction and the garment features. This vector set can clearly reflect the directional distribution of garment features around the field node.
[0028] Step 15: Compare the directional consistency vector sets of the two field nodes to find the consensus direction, generate a virtual connection path based on the consensus direction, and define a dynamic stiffness protocol for the path. The specific operations are as follows: Based on the field nodes and their output direction similarity vector sets defined in step 14, a virtual connection path is generated and a dynamic stiffness protocol is defined to further improve the structural foundation of the pattern topology skeleton. First, any two field nodes are selected, and their output direction similarity vector sets are compared and analyzed. The comparison process uses a vector similarity calculation method to find vectors with higher values and consistent directions in the two vector sets. The directions corresponding to these vectors are the consensus directions between the two field nodes. The consensus direction reflects the common direction of the clothing features perceived by the two field nodes and is also the most suitable direction for building a connection between the two field nodes. Based on the found consensus direction, a virtual connection path connecting the two field nodes is generated. The direction of this path is consistent with the consensus direction. To maintain consistency and ensure that the path conforms to the distribution pattern of clothing features, the length of the virtual connection path is adaptively adjusted based on the distance between the two field nodes and the distribution of clothing features. After the virtual connection path is generated, a dynamic stiffness protocol is defined for the path. The dynamic stiffness protocol is used to specify the stiffness characteristics of the path, including the initial stiffness value of the path, the law of stiffness change with deformation, and the limit threshold of stiffness. For example, the initial stiffness value of the virtual connection path near the edge of the clothing outline is set higher to ensure that it can stably support the clothing outline, while the stiffness value of the path located in the internal texture area of the clothing is set lower to simulate the flexibility of the fabric. The definition of the dynamic stiffness protocol provides a stiffness constraint basis for the deformation of the topological skeleton.
[0029] Step 16: Generate a two-dimensional external potential field based on the input image, place the field nodes and virtual connection paths in the two-dimensional external potential field, perform iterative calculations to achieve dynamic equilibrium, and output the balanced two-dimensional graph structure. The specific operations are as follows: By constructing and iteratively calculating a two-dimensional external potential field, the field nodes and virtual connection paths achieve a dynamic balance, ultimately outputting a two-dimensional graph structure of the pattern topology skeleton. First, a two-dimensional external potential field is generated based on the input clothing image. This generation is based on information such as pixel grayscale values, contour features, and texture distribution of the clothing image. The pixel differences between the clothing area and the background area in the image are converted into potential energy differences in the potential field. The potential energy is higher at the clothing contour edges and lower in the background area. This potential energy distribution allows the field nodes and virtual connection paths to converge towards the clothing contour and feature areas. Subsequently, the field nodes defined in step 14 and the virtual connection paths generated in step 15 are placed entirely within this two-dimensional external potential field. The virtual connection path is subjected to potential energy in the potential field, resulting in motion and deformation tendencies. Through continuous iterative calculations, the motion process of the field nodes and virtual connection paths in the potential field is simulated. Each iteration adjusts the position of the field nodes and the shape of the virtual connection paths according to the current potential energy distribution and the dynamic stiffness protocol of the virtual connection paths, until the position of the field nodes no longer changes significantly and the deformation of the virtual connection paths reaches a stable state, that is, the entire system reaches dynamic equilibrium. At this time, the structure formed by the field nodes and virtual connection paths can accurately fit the outline and feature distribution of the clothing. The output of the balanced structure is the two-dimensional graph structure of the topological skeleton of the target clothing category pattern, completing the final construction of the pattern topological skeleton.
[0030] In a preferred embodiment of the present invention, step 2 is further included: fitting the input clothing image with the clothing pattern topological skeleton, and matching the skeleton outline with the clothing outline in the image through deformation to obtain a skeleton instance and deformation parameters. The specific operation is as follows: The process involves accurately fitting the input clothing image to the constructed clothing pattern topology skeleton. Adaptive deformation of the skeleton ensures its contour matches the actual clothing contour in the image, resulting in a skeleton instance adapted to the current input image and its corresponding deformation parameters. Before fitting, the parameterized clothing pattern topology skeleton library constructed in step 1 is called. Based on the initial visual features of the input clothing image, candidate topology skeletons that preliminarily match the clothing category are selected as initial templates for fitting. This selection process can be quickly completed based on clothing category features to improve fitting efficiency. The fitting logic eliminates the deviation between the skeleton contour and the clothing image contour through dynamic deformation of the skeleton. The deformation process must adhere to the dynamic stiffness defined in step 15. The protocol ensures that the deformation conforms to the physical properties and pattern structure of the garment fabric, avoiding deformed deformations that do not conform to the actual garment structure. During the deformation process, the matching degree between the skeleton outline and the garment image outline is monitored in real time. By continuously adjusting the key point positions of the skeleton and the shape of the elastic constraint edges, the deviation between the two is gradually reduced until the skeleton outline can completely and accurately fit the outline of the garment in the image. At this point, the deformation stops. After the deformation stops, the topological skeleton currently in the fitting state is the skeleton instance adapted to the input garment image. All parameter adjustment data generated during the deformation process, including the displacement of key points, the degree of deformation of elastic constraint edges, and the dynamic adjustment value of stiffness parameters, will be integrated and quantified to form deformation parameters.
[0031] Fitting the input clothing image to the clothing pattern topology skeleton also includes the following steps: Step 21: Activate the orientation sensing function of the field node to generate a probe signal, generate a multi-level image response field of the input clothing image, and couple the probe signal with the image response field to obtain an initial coupling strength vector set. The specific operation is as follows: The orientation sensing function of the field node defined in step 14 is activated. Once activated, this function generates a detection signal with a specific frequency and direction. The generation pattern of the detection signal matches the perception range of the field node, which can fully cover the corresponding area of the input clothing image. This is used to detect the orientation and distribution information of visual features such as the outline and texture of the clothing in the image. Subsequently, a multi-level image response field is generated based on the input clothing image. The division of the levels is based on the feature scale of the clothing image. Different levels correspond to different scales of clothing features. The shallow level mainly responds to the local texture, fine edges and other detailed features of the clothing, while the deep level mainly responds to the overall outline, pattern structure and other macroscopic features of the clothing. The system uses a multi-level setup to ensure comprehensive capture of various visual features of the clothing. Next, the generated detection signal is coupled with the multi-level image response field. The coupling process involves matching the detection signal with the feature information in each level of the response field and calculating the degree of fit between the detection signal and the features at different locations in each level of the response field. This degree of fit is quantified into a specific value. Multiple values are arranged in a certain order to form an initial coupling strength vector group. Each vector in this vector group corresponds to the coupling result between a field node and a certain level of the image response field. The magnitude of the vector value directly reflects the degree of matching between the field node detection signal and the corresponding image feature.
[0032] Step 22: Generate deformation intention based on the initial coupling strength vector set. Calculate local deformation proposals according to the deformation intentions and dynamic stiffness protocol. Simulate the execution of the proposals to obtain new coupling strengths. By comparing the generated image correction feedback signals, obtain the proposal and feedback pair sequence. The specific operations are as follows: The initial coupling strength vector set is analyzed and processed to identify regions with low values. These regions correspond to locations where the matching degree between the field node detection signal and image features is low, and are also key areas requiring deformation adjustment. Simultaneously, based on the distribution pattern of the vector values, the deformation direction and approximate magnitude of each region are determined, thereby generating an overall deformation intention. The deformation intention clarifies the parts, directions, and initial magnitudes of the topological skeleton that need adjustment, providing guidance for the calculation of local deformation proposals. Subsequently, based on the generated deformation intention and the dynamic stiffness protocol defined in step 15, local deformation proposals are calculated. A local deformation proposal is a specific deformation scheme formulated for each region requiring adjustment, including parameters such as the displacement of the field nodes in that region, the deformation angle of the elastic constraint edges, and the stiffness adjustment value. During the calculation process, the dynamic stiffness protocol must be strictly followed to ensure that the deformation involved in the proposal conforms to the physical properties of the clothing fabric, avoiding unreasonable deformations exceeding the stiffness limit. After generating a local deformation proposal, the proposal is simulated and executed using simulation technology. During the simulation, the coupling between the field node detection signal and image features is monitored in real time to obtain the new coupling strength after the proposal is executed. The new coupling strength is compared with the initial coupling strength to determine the change in the matching degree after the proposal is executed. If the new coupling strength is higher than the initial coupling strength, it means that the proposal is executed effectively, and a positive feedback signal is generated. If the new coupling strength is lower than the initial coupling strength, it means that the proposal is executed ineffective or even aggravates the deviation, and a reverse feedback signal is generated. The reverse feedback signal will clearly point out the unreasonable aspects of the proposal and the direction of adjustment, forming an image correction feedback signal. Each local deformation proposal and its corresponding feedback signal form a proposal and feedback pair. Through the continuous process of generating local deformation proposals, simulating execution, comparing coupling strength, and generating feedback signals, a sequence of proposal and feedback pairs is gradually accumulated. This sequence records each deformation attempt and its corresponding effect.
[0033] Step 23: Based on the proposal and feedback sequence, find a stable consensus proposal set, encode the consensus proposal set into a deformation protocol, replay and execute the deformation protocol to drive the topology skeleton deformation, and obtain the skeleton instance and deformation parameters. The specific operations are as follows: The proposal and feedback sequences are analyzed to screen out proposals that significantly improve coupling strength and have no obvious negative feedback after execution; these proposals are considered valid. Invalid proposals that lead to decreased coupling strength or have obvious irrationalities are removed. Based on this, a stable consensus proposal set is sought. A stable consensus proposal set refers to a group of mutually compatible, synergistic, and effective proposals that, when executed together, achieve a high degree of matching in the overall topology skeleton. The screening process ensures that there are no conflicts between proposals, and that the deformations of each region after execution are mutually coordinated, avoiding situations where local deformations are reasonable but the overall outline deviates. Once the stable consensus proposal set is determined, it is encoded into a deformation protocol. The encoding process converts the specific parameters, execution order, and execution conditions of each proposal in the proposal set into a computer-recognizable and executable instruction sequence, ensuring the standardization and executability of the deformation protocol. Subsequently, the deformation protocol is re-executed according to the instruction sequence of the deformation protocol. During the process, the instructions are strictly followed in sequence, and the deformation operations in each proposal are executed step by step. The deformation process and coupling strength of the topology skeleton are monitored in real time to ensure that the deformation process is smooth and accurate until all deformation protocols are executed. After the deformation protocols are executed, the outline of the topology skeleton will be accurately matched with the outline of the input clothing image. At this time, the topology skeleton is the skeleton instance adapted to the input clothing image. At the same time, all deformation-related parameters generated during the execution of the deformation protocol, including the final displacement of each field node, the deformation parameters of the elastic constraint edges, stiffness adjustment values, etc., are integrated, quantified and recorded to form complete deformation parameters. These deformation parameters will comprehensively reflect the deformation of the skeleton instance relative to the initial topology skeleton.
[0034] In a preferred embodiment of the present invention, step 3 is further included: performing graph similarity matching between the skeleton instance and skeletons in the skeleton library to determine the matching category and output the deformation parameter vector. The specific operation is as follows: The skeleton instance is matched with the two-dimensional graph structure of the candidate skeleton in the garment pattern topology skeleton library for similarity. Through multi-stage verification and screening, the matching category corresponding to the input garment image is determined, and the calibrated deformation parameter vector is output. Before the matching begins, the skeleton instance and corresponding deformation parameters obtained in step 2 are called. At the same time, the candidate skeletons that initially match the input garment are extracted from the parameterized garment pattern topology skeleton library constructed in step 1. The screening of candidate skeletons is based on the features of the garment category, which can quickly narrow down the matching range and improve the matching efficiency. The matching process is not a simple contour comparison, but through the processing in steps 31 to 34, combined with physical deformation simulation, feature association analysis, semantic constraint verification and parameter fine-tuning, accurate matching at the skeleton level is achieved, thereby determining the garment category and ensuring the accuracy of the matching results. At the same time, the deformation parameters are optimized and calibrated so that the output deformation parameter vector is more in line with the design intent of the candidate category and the actual shape of the input garment. The entire matching process strictly follows the previously defined dynamic stiffness protocol, design semantics and other constraints to ensure that each operation conforms to the physical characteristics of the garment and the category definition specifications.
[0035] The process of determining the matching product category and outputting the deformation parameter vector also includes the following steps: Step 31: Apply the deformation parameter vector to each candidate skeleton in the garment pattern topology skeleton library, perform physical deformation and relaxation simulation, and obtain a set of relaxed candidate skeletons. The specific operation is as follows: From the parametric garment pattern topology skeleton library, all candidate skeletons that initially match the input garment are extracted. These candidate skeletons cover various subcategories consistent with the major category of the input garment, ensuring comprehensive matching. Subsequently, the deformation parameter vector obtained in step 2 is applied to each candidate skeleton one by one. During the application process, the displacement, stiffness adjustment value, and other parameters contained in the deformation parameter vector must be strictly followed to drive the candidate skeleton to perform the corresponding physical deformation. The deformation process follows the dynamic stiffness protocol defined in step 15 to simulate the deformation law of garment fabric under actual stress, ensuring that the deformed candidate skeleton can initially conform to the shape characteristics of the input garment. After the physical deformation is completed, relaxation simulation is performed on each deformed candidate skeleton. The purpose of relaxation simulation is to eliminate the internal stress generated during deformation, so that the candidate skeleton reaches a stable state and avoids the distortion of skeleton shape due to residual stress, which would affect the subsequent matching accuracy. During relaxation simulation, the deformation state of the candidate skeleton is monitored in real time, and the internal stress is gradually released through iterative calculation until the shape of the candidate skeleton no longer changes significantly and reaches a dynamic equilibrium state. At this time, the candidate skeleton in a stable equilibrium state is the relaxed candidate skeleton. Each candidate skeleton will generate a corresponding relaxed skeleton after the above process, and finally form a complete set of relaxed candidate skeletons. This set of skeletons retains the core structural features of the candidate category and incorporates the deformation characteristics of the input clothing, which is convenient for accurate comparison with the skeleton instance.
[0036] Step 32: Extract key response patterns from the relaxed candidate skeleton, analyze feature signs from the input clothing image, compare key response patterns with feature signs, and calculate causal correlation scores. The specific operations are as follows: For each relaxed candidate skeleton obtained in step 31, key response patterns are extracted. Key response patterns refer to skeleton response information that reflects the structural characteristics of the candidate category. These mainly include the distribution pattern of key points in the relaxed candidate skeleton, the deformation response characteristics of elastic constraint edges, and the coupling state of each structural unit. This information accurately reflects the pattern topology of the candidate skeleton and is also the basis for distinguishing different subcategories. During the extraction process, priority is given to screening response features closely related to the clothing category definition, while irrelevant minor interference features are eliminated to ensure the relevance of the key response patterns. Feature trace analysis is performed on the input clothing image to extract visual features corresponding to the key response patterns, including the overall outline of the clothing, the shape of key parts, texture distribution patterns, and local structural features. During the analysis, background features must be excluded. To mitigate interference from complex scenes and partial occlusion, the system focuses on the core visual features of the clothing itself, ensuring that the extracted features accurately reflect the actual shape of the input garment. Subsequently, the key response patterns of each relaxed candidate skeleton are systematically compared with the feature features of the input garment image. This comparison is not a simple feature overlay but rather an analysis based on causal logic, examining the correspondence between key response patterns and feature features to determine their degree of fit in structure, form, and deformation trends. A pre-defined quantization algorithm converts this fit into a specific numerical value, known as the causal correlation score. A higher score indicates a higher degree of fit between the key response patterns of the relaxed candidate skeleton and the feature features of the input garment, signifying a better match. Conversely, a lower score indicates a lower degree of fit and a worse match.
[0037] Step 33: Project the relaxed candidate skeleton onto the preset design semantic constraint space for verification, and check the deformation parameter vector and the design intent of the candidate category. Figure One To determine consistency and obtain a semantic constraint conflict score, the specific steps are as follows: The design semantics obtained in step 13 are invoked to construct a preset design semantic constraint space. This constraint space contains core semantic information such as design intent, pattern specifications, and structural constraints for various clothing categories. Each candidate category has clear semantic boundaries and constraints in the constraint space, used to standardize the shape and parameter range of the candidate skeleton. Subsequently, each relaxed candidate skeleton obtained in step 31 is projected one by one into this design semantic constraint space. The projection process involves converting the topological parameters and deformation features of the relaxed candidate skeleton into parameters recognizable by the semantic constraint space, and comparing and verifying them with the semantic constraints corresponding to the candidate category. This checks whether the shape and structure of the relaxed candidate skeleton conform to the design semantic requirements of the category and whether there are any unreasonable features that exceed the semantic constraint boundaries. Simultaneously, a consistency check is performed on the deformation parameter vector and the design intent of the candidate product category. The displacement, stiffness adjustment, and other parameters contained in the deformation parameter vector are analyzed to determine whether they match the design concept and pattern requirements of the candidate product category. For example, the design intent of casual products is often loose and comfortable; if the deformation parameter vector shows excessive waist-cinching or tightness, it indicates a conflict with the design intent. Through the above verification and checks, the degree of conflict between the relaxed candidate skeleton and the deformation parameter vector and the design semantic constraints is quantified, forming a semantic constraint conflict score. The higher the score, the more obvious the conflict and the lower the matching rationality of the candidate skeleton; the lower the score, the smaller the conflict and the higher the matching rationality of the candidate skeleton. Together with the causal correlation score, this constitutes a dual quantitative indicator for candidate selection.
[0038] Step 34: Based on the causal correlation score and semantic constraint conflict score, candidates are screened, and the deformation parameters of the screened candidate skeletons are fine-tuned and calibrated. The best matching category is determined and the calibrated deformation parameter vector is output. The specific operations are as follows: A reasonable screening threshold is set, which is determined based on the matching accuracy requirements in actual applications and calibrated with reference to a large amount of sample data. This threshold is used for preliminary screening of candidate skeletons. During the screening process, both the causal correlation score and the semantic constraint conflict score are considered. Candidate skeletons with causal correlation scores higher than the preset threshold and semantic constraint conflict scores lower than the preset threshold are retained, while candidate skeletons that do not meet the threshold requirements are removed. This ensures that the screened candidate skeletons have both high morphological matching and meet the design semantic constraints, reducing the impact of unreasonable candidates on the final matching results. After screening, the deformation parameters corresponding to the remaining candidate skeletons are fine-tuned and calibrated. The core purpose of fine-tuning and calibration is to further improve the matching accuracy, eliminate minor semantic conflicts, and make the deformation parameter vector more consistent with the design intent of the candidate category and the actual shape of the input garment.
[0039] During the fine-tuning process, the goal is to maximize the causal correlation score and minimize the semantic constraint conflict score. Combined with the designed semantic constraints, specific parameters such as displacement and stiffness adjustment values in the deformation parameter vector are slightly optimized. The adjustment process strictly adheres to the dynamic stiffness protocol and the physical properties of the garment fabric to avoid unreasonable deformation. After fine-tuning and calibration, for each optimized candidate skeleton, a comprehensive decision is made based on the adjusted causal correlation score and semantic constraint conflict score. The candidate skeleton with the highest causal correlation score and the lowest semantic constraint conflict score is selected as the optimal matching skeleton. The category corresponding to this optimal matching skeleton is the matching category of the input garment. Finally, the fine-tuned and calibrated deformation parameter vector is integrated and quantified to form a complete calibrated deformation parameter vector. This vector accurately reflects the deformation correspondence between the input garment and the matching category skeleton, while also conforming to the design intent of the matching category. The matching category and the calibrated deformation parameter vector are output together.
[0040] In a preferred embodiment of the present invention, step 4 is further included: deforming the skeleton of the matching category according to the deformation parameter vector, and combining it with the texture information extracted from the input image to generate a virtual try-on image. The specific operation is as follows: The topological skeleton corresponding to the matching category determined in step 3 is invoked. The calibrated deformation parameter vector output in step 3 is applied to the skeleton to drive it to deform precisely. The deformation process strictly follows the dynamic stiffness protocol defined in step 15 to simulate the actual deformation law of the garment fabric, ensuring that the deformed skeleton can perfectly fit the contour and pattern features of the input garment. At the same time, texture information is extracted from the input garment image. During the extraction process, interference factors such as complex background and local occlusion are eliminated, focusing on the texture details of the garment itself, including texture patterns, color distribution, and fabric texture. The deformed skeleton and the extracted texture information are deeply fused. Combined with lighting simulation technology, the light and shadow effects and texture presentation of the garment are restored, and finally a virtual try-on image is generated. This image not only retains the core pattern features of the matching category, but also restores the real texture details of the input garment, and can clearly present the overall shape and local features of the garment.
[0041] The process of generating virtual try-on images also includes the following steps: Step 41: Establish a continuous texture field for the input clothing area, analyze the texture field to assign it physical properties, and record the original lighting reference. The specific operations are as follows: The input clothing image is segmented into clothing regions. Image segmentation algorithms accurately separate the clothing and background regions, eliminating interference factors such as complex backgrounds and partial occlusions to ensure that subsequent texture extraction focuses solely on the clothing itself. After segmentation, global texture sampling is performed on the clothing region. A combination of uniform and targeted sampling is used to cover the entire clothing area while densely sampling areas with rich texture details, ensuring that the sampled data fully reflects the texture features of the clothing. Based on the sampled texture data, a continuous texture field is constructed using an interpolation algorithm. The continuous texture field eliminates texture breakage caused by sampling discreteness, achieving a smooth transition of clothing textures and ensuring the continuity and integrity of texture morphology during subsequent texture evolution. The established continuous texture field is then analyzed for physical properties. Analysis and assignment: Combining the actual characteristics of clothing fabrics, corresponding physical parameters are assigned to the texture field, including tensile strength, elastic modulus, wrinkle recovery coefficient, and texture density. These physical properties directly determine the response law of the texture field under the action of subsequent geometric deformation field. For example, denim texture has high tensile strength and is assigned a large elastic modulus parameter, while silk texture has low tensile strength and a small elastic modulus parameter. While assigning physical properties, the original lighting reference of the input clothing image is recorded simultaneously. Through image lighting analysis algorithms, lighting parameters such as lighting direction, lighting intensity, color temperature, and light and shadow distribution in the image are extracted. These parameters can truly reflect the original lighting environment of the input clothing image, ensuring that the lighting effect of the generated virtual try-on image is consistent with the original image, thus improving visual realism.
[0042] Step 42: Generate a geometric deformation field based on the calibrated deformation parameter vector. Calculate the non-rigid response of the texture field under the geometric deformation field according to its physical properties, thus obtaining the evolved texture field. The specific operations are as follows: The calibrated deformation parameter vector output in step 3 is used to generate a geometric deformation field. This field is a spatial field that drives the texture field to undergo corresponding deformations. It contains information such as the deformation direction, amplitude, and rate of deformation in various areas of the garment, accurately mapping the deformation requirements corresponding to the calibrated deformation parameter vector. This ensures that the evolution of the texture field is synchronized with the deformation of the skeleton. The generation process of the geometric deformation field strictly follows the dynamic stiffness protocol and the physical properties of the texture field, avoiding unreasonable deformation driving beyond the physical limits of the texture. After generating the geometric deformation field, the non-rigid response of the texture field under the action of the geometric deformation field is calculated according to the physical properties assigned to the continuous texture field in step 41. The non-rigid response mainly manifests as the texture field changing with geometric deformation. The stretching, compression, wrinkling, and twisting effects generated by the field require careful consideration of the differences in the physical properties of the texture during calculation. Textures with different physical properties will produce different responses under the same geometric deformation field. For example, denim textures with a high elastic modulus will exhibit smaller deformation amplitudes and fewer wrinkles under stretching deformation, while cotton textures with a low elastic modulus will exhibit larger deformation amplitudes and are more prone to wrinkling under the same stretching deformation. Through iterative calculation, the dynamic evolution process of the texture field under the action of the geometric deformation field is simulated, and the shape and distribution of the texture are adjusted in real time until the deformation of the texture field perfectly matches the driving requirements of the geometric deformation field, and the texture shape conforms to its physical property laws. The resulting texture field is the evolved texture field. The evolved texture field retains the original texture details of the input garment, and its shape is perfectly adapted to the deformed skeleton.
[0043] Step 43: Establish a virtual lighting environment based on the original lighting reference, calculate the lighting information of the deformed skeleton geometry, and synthesize the color information and lighting information of the evolved texture field to generate a virtual try-on image. The specific operations are as follows: The original lighting reference recorded in step 41 is used to establish a virtual lighting environment. All parameters of the virtual lighting environment are consistent with the original lighting reference, including lighting direction, intensity, and color temperature, ensuring that the virtual lighting environment accurately replicates the lighting effects of the original image. This ensures that the generated virtual try-on image maintains consistency with the input clothing image in terms of lighting and shadow, enhancing visual realism. After the virtual lighting environment is established, geometric analysis is performed on the deformed skeleton from step 4 to extract its geometric features, including the skeleton's outline, key point positions, and the concave and convex shapes of various parts. Based on these geometric features and the parameters of the virtual lighting environment, the lighting and shadow information of the deformed skeleton geometry is calculated. This information mainly includes highlight areas, shadow areas, and lighting transition gradients for various parts of the skeleton. During the calculation, the reflection and refraction laws of light and skeleton geometry must be simulated to ensure the realism and rationality of the lighting and shadow information. For example, protruding parts of the skeleton will form highlights, and concave parts will form shadows; the lighting transition gradient conforms to the laws of light propagation.
[0044] Extract the color information of the evolved texture field obtained in step 42, including the color distribution, color gradation, and color saturation of the texture. Deeply fuse and synthesize this color information with the calculated light and shadow information. During the synthesis process, the light and shadow information needs to be accurately superimposed on the corresponding texture area so that the texture can reflect the corresponding highlight and shadow effects while presenting the original color and details, simulating the visual presentation of clothing under actual lighting. After the synthesis is completed, the image is smoothed and optimized to eliminate problems such as texture breakage and harsh lighting caused during the synthesis process, ensuring the visual smoothness and realism of the virtual try-on image. Finally, a complete virtual try-on image is output. This image has both the pattern features of the deformed skeleton and the texture details of the input clothing, while also having realistic lighting effects, which can clearly and accurately present the overall shape and local features of the clothing.
[0045] In a preferred embodiment of the present invention, step 5 is further included: applying a predetermined offset to the key points of the deformed skeleton to generate a perturbed skeleton, and generating a set of perturbed comparison images by combining texture information. The specific operation is as follows: A predetermined offset is applied to the key points of the deformed skeleton in step 4 to generate a perturbed skeleton. Combined with the texture information of the input clothing, a set of perturbed comparison images is synthesized through multiple optimization steps. This provides diverse comparison samples for the accurate classification decision of the dual-channel decision network in step 6, helping to solve the semantic confusion problem in existing technologies. The entire process must strictly connect the core technical elements mentioned earlier, such as virtual try-on images, evolved texture fields, semantic constraint spaces, and dynamic stiffness protocols, to ensure the rationality and relevance of the perturbation generation. Through gradient analysis and semantic vulnerability identification of the virtual try-on images, a directional perturbation scheme is determined, generating a perturbed skeleton and a preliminary family of perturbed comparison images. Subsequently, texture strain analysis and texture response prediction are performed on the geometric deformation of the perturbed skeleton, and collaborative optimization is carried out through the mid-term feedback of the dual-channel decision network. Finally, lighting parameters are adjusted and detail representation is enhanced to synthesize high-quality perturbed comparison images. The generated set of perturbed comparison images is essentially a structured perturbation of the original virtual try-on images, highlighting the semantic differences between different product categories.
[0046] The process of applying a predetermined offset to the key points of the deformed skeleton to generate a disturbed skeleton also includes the following steps: Step 51: Calculate the classification gradient of the dual-channel decision network based on the virtual try-on image, and inversely map the gradient into a keypoint displacement proposal field for the deformed skeleton. The specific operation is as follows: The virtual try-on image generated in step 4 is input into the dual-channel decision network to initiate the network's classification prediction process. The classification gradient is calculated in real-time, reflecting the influence of visual features in each region of the virtual try-on image on the network's classification results. Regions with higher gradient values indicate stronger determinative influence of visual features on category classification and are key regions for distinguishing similar categories; regions with lower gradient values have less impact on classification results and can be considered non-critical perturbation regions. After calculating the classification gradient, it needs to undergo inverse mapping. Since the classification gradient is calculated at the image pixel level, while perturbation needs to act on skeleton keypoints, inverse mapping establishes a correspondence between image pixel gradients and skeleton keypoint displacements, converting pixel-level gradient information into skeleton keypoint displacement suggestions. Through a preset mapping algorithm, the classification gradient of each region is converted into suggestions for the offset direction and magnitude of the corresponding skeleton keypoint. All keypoint offset suggestions are integrated to form a keypoint displacement suggestion field. This field clarifies the potential perturbation direction and magnitude for each skeleton keypoint and prioritizes key regions with higher classification gradients, improving the targeting and effectiveness of the perturbation.
[0047] Step 52: Identify the semantic structural vulnerabilities of the current category from the category design semantic constraint space, and filter and strengthen the displacement proposal field in conjunction with the dynamic stiffness protocol to generate a directional perturbation scheme. The specific operations are as follows: The semantic constraint space for category design, constructed in step 13 and used in step 33, is invoked. Combined with the matching categories determined in step 3, semantic structural vulnerabilities of the current category are identified within this constraint space. Semantic structural vulnerabilities refer to key skeletal structural points in the current category that are prone to semantic confusion with other similar categories and play a crucial role in the category definition. For example, the semantic structural vulnerability of the hood connection point of a hoodie; even a slight deformation of this point could lead to semantic confusion with the hoodie. These vulnerabilities are the targets of directional perturbations. After identifying semantic structural vulnerabilities, the dynamic stiffness protocol defined in step 15 is invoked. This protocol is then used to filter the displacement suggestion field obtained in step 51, eliminating offset suggestions that do not meet the requirements of the dynamic stiffness protocol. For example... For displacement suggestions that exceed the stiffness limit of key skeleton points and may lead to skeleton deformation, ensure that the selected displacement suggestions conform to the physical properties of the garment fabric and the deformation law of the skeleton; strengthen the displacement suggestions corresponding to semantic structural vulnerabilities by appropriately increasing the displacement amplitude of these vulnerabilities so that the perturbation can more clearly highlight the semantic feature differences of this part and improve the perturbation's ability to distinguish similar categories; appropriately weaken the displacement suggestions for non-semantic structural vulnerabilities to avoid unnecessary ineffective perturbations; after screening and strengthening, integrate all reasonable displacement suggestions and combine them with the perturbation priority of semantic structural vulnerabilities to formulate a complete directional perturbation scheme. This scheme clarifies the perturbation sequence, displacement direction, displacement amplitude, and constraint conditions for each key skeleton point.
[0048] Step 53: Design a structured perturbation sequence based on the directional perturbation scheme, generate a perturbation skeleton and corresponding explanatory labels, and synthesize a family of perturbation contrast images based on the evolved texture field. The specific operations are as follows: The directional perturbation scheme obtained in step 52 is decomposed and sorted to design a structured perturbation sequence. This sequence arranges the perturbation tasks in the directional perturbation scheme according to perturbation priority and the correlation between perturbation locations, forming progressively executable perturbation steps. This avoids simultaneous perturbation of multiple key points, which could lead to skeleton morphology distortion and excessive internal stress, ensuring a smooth and controllable perturbation process. Each perturbation step corresponds to an offset operation of one or a group of associated key points, clearly defining the perturbation parameters and execution conditions for that step. Following the steps of the structured perturbation sequence, predetermined offsets are applied progressively to the key points of the deformed skeleton in step 4, driving the skeleton to undergo targeted deformation. Each completed perturbation step generates a corresponding perturbation skeleton. This process is repeated after the complete sequence... The execution process ultimately generates multiple perturbation skeletons with different perturbation levels and locations, forming a set of perturbation skeletons. Simultaneously, each perturbation skeleton is generated with a corresponding explanatory label. These labels detail the perturbation location, keypoint offset magnitude, perturbation basis, and corresponding semantic impact. For example, a label indicating a 3mm upward shift of the hat connection keypoint simulates the hoodless structure of a sweatshirt, distinguishing it from a sweatshirt. Finally, the evolved texture field obtained in step 42 is called to initially fuse each perturbation skeleton with the evolved texture field, simulating the texture's attachment effect on the perturbation skeleton, and synthesizing preliminary perturbation comparison images. Multiple perturbation skeletons correspond to multiple preliminary perturbation comparison images, collectively forming a perturbation comparison image family.
[0049] The process of generating a set of perturbation contrast images by combining texture information also includes the following steps: Step 54: Deconstruct the geometric deformation of the perturbed skeleton into a local strain field applied to the texture field. Based on the physical properties of the texture field, predict its state changes using a microscopic material response model to obtain the texture response prediction field. The specific operations are as follows: Perform geometric deformation analysis on each perturbation skeleton generated in step 53. Deconstruct the geometric deformation of the perturbation skeleton relative to the deformed skeleton in step 4 into a local strain field applied to the texture field after evolution in step 42. The local strain field is a spatial distribution field that can accurately reflect the tensile, compressive, torsional and other strain effects on each region in the texture field. Its strain distribution corresponds completely to the geometric deformation distribution of the perturbation skeleton. For example, the cuff deformation caused by the offset of the cuff key point of the perturbation skeleton will form a corresponding tensile or compressive strain in the corresponding region of the texture field.
[0050] Step 41 is called to assign physical properties to the evolved texture field, including core parameters such as tensile strength, elastic modulus, wrinkle recovery coefficient, and texture density. These physical properties are then input into a preset micromaterial response model. The micromaterial response model is used to simulate the microscopic state changes of the texture under different strains. It can accurately predict the state changes of the texture, such as morphology, density, and wrinkles, based on the strain parameters of the local strain field and the physical properties of the texture field. For example, cotton texture with a small elastic modulus will be predicted to produce obvious wrinkles and texture stretching deformation when subjected to large tensile strain; while denim texture with a large elastic modulus will be predicted to produce smaller deformations under the same tensile strain. Through iterative calculation of the micromaterial response model, the dynamic change process of the texture under the action of the local strain field is simulated. Finally, the texture state prediction result that can accurately adapt to the geometric deformation of the perturbed skeleton is output, namely the texture response prediction field. This prediction field retains the original texture details of the evolved texture field and incorporates the texture changes brought about by the deformation of the perturbed skeleton, ensuring the coordination and consistency between the texture and the perturbed skeleton.
[0051] Step 55: The preliminary results of the texture response prediction field and the perturbation skeleton are combined and input into the dual-channel decision network to obtain intermediate perceptual features. Based on the comparison with the reference features, the texture response prediction field and the perturbation skeleton are iteratively adjusted collaboratively to obtain the optimized pairing. The specific operations are as follows: The texture response prediction field obtained in step 54 is initially combined with the corresponding perturbation skeleton to form an initial pairing of the perturbation skeleton and the texture response prediction field. This initial pairing may have problems such as inaccurate adaptation between texture and skeleton and unreasonable texture deformation, which need to be further optimized. The initial pairing is input into a dual-channel decision network to start the network's intermediate perception feature extraction process to obtain the intermediate perception features of the initial pairing. The intermediate perception features are used to reflect the degree of cooperative adaptation between the perturbation skeleton and the texture response prediction field, as well as the ability of the paired image to support category classification. The higher the feature value, the better the pairing adaptation and the better it can highlight the semantic differences of the categories. The lower the feature value, the more adaptation defects there are in the pairing, which need to be adjusted. At the same time, reference features are obtained. The reference features are the perception features corresponding to the pairing of the deformed skeleton and the evolved texture field in step 4. As a benchmark for pairing optimization, the intermediate-term perceptual features are systematically compared with the reference features to analyze the deviation between them and identify problems in the initial pairing, such as mismatch between the wrinkle distribution of the texture response prediction field and the deformation of the perturbed skeleton, or the texture stretching degree exceeding the physical property limit. Based on the deviation information obtained from the comparison, the texture response prediction field and the perturbed skeleton are adjusted collaboratively and iteratively. For example, the wrinkle distribution of the texture response prediction field is optimized by adjusting the local strain field parameters, and the offset amplitude of the key points of the perturbed skeleton is fine-tuned to correct the skeleton deformation. After each adjustment, the pairing input is re-inputted into the dual-channel decision network to extract the intermediate-term perceptual features and compared with the reference features until the deviation between the intermediate-term perceptual features and the reference features is reduced to below the preset threshold, and the pairing adaptability reaches the optimal level. Finally, the optimized perturbed skeleton and texture response prediction field pairing is obtained.
[0052] Step 56: Adjust the parameters of the virtual lighting environment according to the interpretive labels. In areas where the pattern density changes in the texture response prediction field, increase the representation of the micro-geometric details of the surface during the rendering process, and synthesize a perturbation contrast image. The specific operations are as follows: The corresponding explanatory labels generated in step 53 are invoked, and the information recorded by these labels, such as the location of the disturbance and its semantic impact, is analyzed. Based on this information, the virtual lighting environment parameters established in step 43 are adjusted. The core purpose of the adjustment is to highlight the semantic differences brought about by the disturbance. For example, if the explanatory label indicates waist key point contraction, which is used to distinguish between a straight skirt and an A-line skirt, the lighting direction and intensity are adjusted to enhance the light and shadow contrast in the waist area, making the disturbance deformation and texture changes in the waist area more clearly visible. For areas without key disturbances, the lighting parameters are kept consistent with the original virtual lighting environment to avoid invalid lighting interference. After the lighting parameters are adjusted, the focus is on the areas with changes in pattern density in the texture response prediction field. These areas are where the texture changes brought about by the disturbance are most obvious and are also key texture areas for distinguishing similar categories. During the image rendering process, these areas are enhanced through a preset detail enhancement algorithm. The representation of micro-geometric details on the domain surface, such as enhanced texture wrinkles, texture graininess, and pattern edge clarity, makes the texture differences caused by perturbations easier to identify, improving the distinguishability of perturbation comparison images. Subsequently, the color information and micro-details of the optimized texture response prediction field are deeply fused and synthesized with the lighting information under the adjusted virtual lighting environment. During the synthesis process, the adaptation relationship between the texture and the perturbation skeleton is strictly followed to ensure natural texture attachment and smooth lighting transition. After synthesis, the image is smoothed and optimized to eliminate problems such as texture breaks, harsh lighting, and jagged details that occur during the synthesis process, ensuring the visual smoothness and realism of the perturbation comparison images. The above process is repeated to synthesize and optimize each optimized pair, finally generating a complete set of perturbation comparison images. This set of images can clearly present the category semantic differences caused by different perturbations.
[0053] In a preferred embodiment of the present invention, step 6 is further included: inputting the topological features, texture features, and perturbation contrast image of the skeleton instance into a dual-channel decision network, and outputting category labels and decoupled confidence vectors. The specific operation is as follows: The topological features, texture features, and perturbation comparison images of the skeleton instances obtained in the previous section are input into a dual-channel decision network. Through sub-steps such as counterfactual feature reconstruction, causal intervention calculation, dynamic logic arbitration, and decision evidence encapsulation, the network ultimately outputs accurate category labels and decoupled confidence vectors. This completely solves the semantic confusion problem that easily occurs in existing technologies, ensuring the accuracy and interpretability of label generation. The entire process strictly connects the technical results of each step in the previous section, including the topological features of the skeleton instances obtained in step 2, the clothing texture features extracted in step 4, the perturbation comparison images generated in step 5, the explanatory labels in step 53, and the semantic constraint conflict scores in step 33, forming a complete inference chain. The dual-channel decision network is responsible for receiving multi-dimensional inputs and performing preliminary classification, then optimizing, calibrating, and verifying the classification results. At the same time, it encapsulates the decision evidence chain to ensure that the output results are not only accurate but also traceable and interpretable, meeting the needs of e-commerce retail and digital management of clothing for the accuracy and verifiability of clothing category labels.
[0054] Step 6 also includes the following sub-steps: Step 61: Analyze the explanatory labels corresponding to the perturbation skeleton and select a subset of relevant perturbation images. Reconstruct the counterfactual feature representation of the competing categories through feature inverse mapping. The specific operations are as follows: The explanatory labels corresponding to each perturbation skeleton generated in step 53 are invoked, and a systematic analysis is performed on all explanatory labels to clarify the perturbation location, perturbation parameters, semantic impact, and corresponding perturbation comparison images recorded by each label. The focus is on filtering out perturbation images related to the current matching category (i.e., the decision result in step 3) and competing categories with similar semantics that are easily confused. The filtering logic is based on the semantic impact description in the explanatory labels, prioritizing perturbation images that highlight the semantic differences between the current category and competing categories. For example, when the current category is hoodies, images labeled with simulated hoodless hoodies and adjusted hood connection points are filtered out. Perturbation images related to hoodies, i.e., competing categories, are used to identify key points and other perturbation images. These images together constitute a subset of perturbation images, ensuring that the subset can accurately distinguish similar categories. After filtering, feature extraction is performed on the subset of perturbation images to obtain their visual features. Then, through a feature inverse mapping algorithm, these visual features are inversely mapped to the category feature space to reconstruct the counterfactual feature representation of the competing categories. The counterfactual feature representation essentially simulates the combination of topological and texture features that the current input clothing should possess if it belongs to a certain competing category, which can clearly present the semantic differences between the current category and competing categories.
[0055] Step 62: Establish a causal intervention calculation layer, which receives the topological features, texture features, counterfactual feature representations, and semantic constraint conflict scores of the skeleton instance, performs feature replacement intervention, and calculates the causal effect matrix. The specific operations are as follows: A causal intervention computation layer is constructed. This layer serves as an enhancement module for the dual-channel decision network, specifically designed to analyze the causal relationship between each input feature and the classification result, rather than simply statistical correlation. This avoids the semantic confusion problem caused by the over-reliance on local visual features in existing technologies. The input to the causal intervention computation layer includes four types of key data: the skeleton instance topological features obtained in step 2, the clothing texture features extracted in step 4, the counterfactual feature representation of competing categories reconstructed in step 61, and the semantic constraint conflict score obtained in step 33. These four types of data comprehensively cover the structural features, visual features, competitive contrast features, and semantic constraint features of clothing, ensuring the comprehensiveness and accuracy of causal analysis. After the input data is prepared, the causal... The intervention computation layer performs feature replacement intervention operations, and then replaces the topological features and texture features of the skeleton instance one by one with the corresponding counterfactual features of the competing categories. After each replacement, the changes in the classification results of the dual-channel decision network are recorded. Through this targeted replacement, the intervention effect of different features on the classification results is simulated. Subsequently, based on the results of each feature replacement intervention, a causal effect matrix is calculated through a preset quantization algorithm. Each element in the matrix corresponds to the degree to which the classification result shifts to a certain category after a certain feature replacement, i.e., the causal effect strength. The causal effect matrix can clearly quantify the impact of topological features, texture features, and counterfactual features on the classification results of each category, and clarify which features are the key to distinguishing the current category from competing categories.
[0056] Step 63: Establish a dynamic arbitrator to check logical conflicts based on the causal effect matrix and semantic constraint conflict scores, dynamically calibrate the original classification probabilities, and output the calibrated category labels and decoupling confidence vectors based on the strength of causal effects. The specific operations are as follows: A dynamic arbitrator is established, which has three main functions: logical conflict detection, classification probability calibration, and result output. Its role is to coordinate the results of causal effect analysis and semantic constraint verification, resolve potential biases in a single analysis dimension, and ensure the rationality and accuracy of the classification results. The dynamic arbitrator first receives the causal effect matrix obtained in step 62 and the semantic constraint conflict score obtained in step 33, and performs collaborative analysis on the two to check for logical conflicts. Logical conflicts are mainly manifested when the causal effect matrix shows that a certain category has the highest causal effect strength, tending to classify the input clothing into that category, but the semantic constraint conflict score shows that the category has a significant conflict with the design semantics of the input clothing, or the causal effect is inconsistent with the optimal category pointed to by the semantic constraints. For the detected logical conflicts, the dynamic arbitrator dynamically adjusts the influence weights of the two through a weight allocation algorithm, combining the causal effect strength and the degree of semantic constraint conflict. For example, if the semantic constraint conflict score is low, it means that the category conforms to the design semantics, so the weight of the corresponding element in the causal effect matrix is increased; if the semantic constraint conflict score is high, the corresponding weight is decreased, thereby coordinating the logical conflicts.
[0057] The dynamic arbitrator invokes the original classification probabilities output by the dual-channel decision network. These original probabilities are preliminary category matching probabilities calculated by the network based on input features, without considering the synergistic effects of causal effects and semantic constraints. Based on the coordinated causal effect matrix and semantic constraint conflict scores, the original classification probabilities are dynamically calibrated. During calibration, the probabilities of categories with high causal effect strength and low semantic constraint conflict are emphasized, while the probabilities of categories with low causal effect strength and high semantic constraint conflict are weakened. This ensures that the calibrated classification probabilities can simultaneously consider causal logic and semantic rationality. After calibration, the calibrated category labels are output, which are the final category labels for the input clothing. Simultaneously, a decoupled confidence vector based on the causal effect strength is output. Each dimension of the decoupled confidence vector corresponds to a clothing category, and the vector value represents the confidence level for that category. The calculation of confidence level is directly related to the causal effect strength, achieving decoupling of confidence level from feature causal effects. This clearly presents the source of confidence level for each category, further improving the interpretability of the classification results and avoiding judgment bias caused by confidence level confusion.
[0058] Step 64: Encapsulate a structured chain of decision-making evidence containing key intermediate data and the reasoning process. The specific operations are as follows: The encapsulation of a structured decision-making evidence chain needs to comprehensively cover the entire process from input clothing images to output category labels. First, key intermediate data for each step should be compiled and collected. This data includes, but is not limited to, parameters of the clothing pattern topology skeleton library in step 1, skeleton instances and deformation parameters in step 2, causal correlation scores and semantic constraint conflict scores in step 3, evolved texture fields and virtual lighting parameters in step 4, perturbed skeletons, explanatory labels and perturbed comparison images in step 5, counterfactual feature representations in step 61, causal effect matrices in step 62, and the original and calibrated classification probabilities in step 63. Simultaneously, the reasoning logic and operational steps of the entire process should be recorded in detail, including the execution order of each step, the basis for parameter calculation, feature extraction logic, intervention operation procedures, and arbitration decision-making processes. To ensure the integrity and clarity of the reasoning process, key intermediate data is then integrated and encapsulated according to a pre-defined structured format to form a structured decision evidence chain. During encapsulation, a one-to-one correspondence between data and reasoning must be maintained, each decision result must be supported by clear evidence, and each intermediate data point must be labeled with its source and generation logic to avoid data confusion or reasoning gaps. The encapsulated structured decision evidence chain can completely reconstruct the entire category label generation process. Those skilled in the art can use this evidence chain to verify the operational rationality and data accuracy of each step, trace the basis for the classification results, and provide clear data support and logical reference for subsequent objection verification and technical optimization of label generation results, further improving the practicality and reliability of the entire technical solution.
[0059] In a preferred embodiment of the present invention, the invention further includes the construction of a microscopic material response model, the specific operations of which are as follows: Model construction: S1. Determine the input and output dimensions of the model. The input layer is set as a feature parameter set that matches this scheme, including parameters of the local strain field, such as tensile strain coefficient, compressive strain coefficient, torsional strain angle, spatial coordinates of the strain area, and physical property parameters assigned to the texture field in step 41, such as tensile strength, elastic modulus, wrinkle recovery coefficient, texture density, and fabric material category scalar. The number of neurons in the input layer corresponds one-to-one with the above parameter dimensions to ensure that no parameters are omitted from the input. The output layer consists of micro-state parameters of the texture response prediction field, including wrinkle distribution probability, texture stretching rate, pattern density change value, and displacement offset of texture pixels. The number of neurons in the output layer matches the sampling resolution of the texture field to ensure that the output can be mapped to the global space of the texture field.
[0060] S2. The hidden layer structure of the design model adopts a hybrid structure of three fully connected hidden layers combined with two-dimensional convolutional layers. The first fully connected layer realizes the linear feature mapping of the input parameters, transforming discrete strain and physical property parameters into continuous feature vectors. The two-dimensional convolutional layer uses a 3×3 convolution kernel to perform spatial dimension convolution operations on the feature vectors, simulating the spatial transmission law of strain in the texture field, which fits the spatial continuity of clothing texture deformation. The last two fully connected layers complete the nonlinear fitting of the convolutional features, establishing a nonlinear relationship between strain parameters, physical properties and texture microstate. The number of neurons in the hidden layer is set by decreasing layer by layer to achieve feature dimensionality reduction and extraction of core information.
[0061] S3. Embed the constraint rules of this scheme, add hard constraints to the calculation of each layer of the model, transform the dynamic stiffness protocol defined in step 15 into mathematical constraint terms, and embed them into the calculation process of convolutional layers and fully connected layers to ensure that the texture deformation predicted by the model does not exceed the physical stiffness limit of the fabric; at the same time, add texture continuity constraint terms to limit the displacement offset difference of adjacent texture pixels, avoid unreasonable results such as texture breakage and distortion, and make the texture response prediction field output by the model conform to the actual deformation law of clothing fabric.
[0062] S4. Define the core computational mapping relationship of the model. Achieve non-linear transformation of parameters through non-linear activation functions. The convolutional layer uses the ReLU activation function to solve the gradient vanishing problem in the feature mapping process. The fully connected layer uses the Tanh activation function to map the output micro-state parameters to a reasonable numerical range, such as mapping the texture stretching rate to the range of 0 to 2, which conforms to the actual stretching range of the fabric. Finally, a complete computational mapping relationship from the input parameters to the texture micro-state parameters is formed, and the overall framework of the model is completed.
[0063] Training of microscopic material response models: S01. Construct a dedicated training dataset. The samples in the dataset are all generated based on the technical system of this solution. First, select common clothing fabric material samples in the e-commerce retail field, such as denim, cotton, silk, wool, lace, etc., and assign the standard physical property parameters defined in step 41 to each material. Then, simulate different local strain field parameters to cover the tensile, compressive, and torsional strain ranges that may be generated by the perturbed skeleton in this solution. Obtain the actual measured values of the texture microstate of each material under different strain fields through fabric physics simulation software, such as wrinkle distribution and texture stretching rate. Finally, use the local strain field parameters and fabric physical property parameters as sample inputs, and the corresponding actual measured values of texture microstate as sample labels to construct a fully labeled training dataset. At the same time, divide the dataset into training set, validation set, and test set in a ratio of 7:2:1 to meet the needs of model training, optimization, and performance verification.
[0064] S02. Model initialization and hyperparameter setting: Randomly initialize the convolutional kernel weights, fully connected layer weights, and biases of the model, ensuring that the initial values follow a standard normal distribution; set the model's training hyperparameters, selecting the Adam optimizer, with an initial learning rate of 0.001 and a step-decay strategy, reducing the learning rate to 0.5 every 100 epochs; set the batch size to match the texture field sampling resolution, and set the number of training epochs to 500 to ensure full model convergence.
[0065] S03. Forward and Backward Propagation Training of the Model: Samples from the training set are input into the initialized model. Forward propagation calculations are performed through the input layer, hidden layer, and output layer of the model to obtain the predicted values of the texture microstates. The mean square error between the predicted value and the sample label is used as the core loss function to calculate the error value between the predicted value and the actual measured value. The calculation process of the loss function is in line with the texture deformation requirements of this scheme, and higher error weights are assigned to core parameters such as wrinkle distribution and texture stretching rate to improve the model's prediction accuracy of key microstates. Through the backpropagation algorithm, the loss value is propagated back from the output layer to the input layer layer by layer. The gradient descent method is used to update the weights and biases of each layer of the model. At the same time, during the backpropagation process, gradient clipping is performed on parameters that exceed the dynamic stiffness protocol constraints to ensure that the model always meets the physical constraints of this scheme.
[0066] S04. Model Validation and Tuning: Every 50 epochs of training, input validation set samples into the model and calculate the model's loss and prediction accuracy on the validation set. If the validation set loss does not decrease for 100 consecutive epochs, stop training to avoid overfitting. If the validation set loss fluctuates, adjust the number of hidden layer neurons or the learning rate decay strategy and retrain. After training, input test set samples into the model to verify the model's prediction accuracy on unseen samples and ensure the model's generalization ability. The prediction accuracy on the test set must reach above 95% before it can be applied to step 54 of the solution.
[0067] S05. Deployment and adaptation of the model: The trained model is deployed in a lightweight manner and transformed into a calculation module that can be embedded in the image generation process of this scheme. This enables seamless integration with the local strain field parameters and texture field physical property parameters in step 54. The input and output interfaces of the model adopt the parameter format of this scheme, without the need for additional parameter conversion. This ensures that the model can be directly called to complete the generation of the texture response prediction field during the calculation process in step 54.
[0068] After the model is built and trained, the calculation process of step 54 is directly embedded. The local strain field parameters obtained by deconstruction in step 54 and the physical property parameters of the evolved texture field are input into the model. Through the forward propagation calculation of the model, the micro-state parameters of the texture response prediction field are directly output. Then, the parameters are mapped to the sampling space of the texture field to obtain the complete texture response prediction field, thus achieving the adaptation with the geometric deformation of the perturbed skeleton.
[0069] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concept, should be covered within the scope of protection of the present invention.
Claims
1. A method for automatically generating clothing category tags based on deep learning, characterized in that, include: Define a parametric clothing pattern topology skeleton library that covers the target product category. Each skeleton is a two-dimensional graph structure composed of key points and elastic constraint edges. The input clothing image is fitted with the topological skeleton of the clothing pattern, and the skeleton outline is matched with the clothing outline in the image by deformation to obtain the skeleton instance and deformation parameters. Perform graph similarity matching between the skeleton instance and the skeleton in the skeleton library to determine the matching category and output the deformation parameter vector; The skeleton of the matching product category is deformed based on the deformation parameter vector, and combined with the texture information extracted from the input image to generate a virtual try-on image; A perturbation skeleton is generated by applying a predetermined offset to the key points of the deformed skeleton and then combining the texture information to generate a set of perturbation comparison images. The topological features, texture features, and perturbation contrast images of the skeleton instance are input into a dual-channel decision network, which outputs category labels and decoupled confidence vectors.
2. The method for automatically generating clothing category tags based on deep learning according to claim 1, characterized in that, Define a parametric garment pattern topology skeleton library that covers the target product category, including: Define a set of clothing structural primitives, each primitive encapsulating the biomechanical interaction logic associated with a specific human body part and the dynamic response rules of the fabric; Based on the target product category, relevant primitives are selected from the primitives. A dynamic topology network is formed through physical coupling negotiation between the primitives. The stiffness distribution function of the connection points and the boundary deformation transfer coefficient are calculated to obtain the set of physical parameters of the target product category. The design semantics are obtained, and a modulation instruction sequence is generated based on the design semantics. The physical parameter set is then subjected to gradient modulation to generate the pattern topology skeleton of the target product category.
3. The method for automatically generating clothing category tags based on deep learning according to claim 2, characterized in that, The construction of the pattern topology skeleton includes: Define a field node, which has a perceptual function that scans the surrounding image space to output a set of orientation fit vectors; Compare the directional consistency vector sets of two field nodes to find the consensus direction, generate virtual connection paths based on the consensus direction, and define a dynamic stiffness protocol for the path. A two-dimensional external potential field is generated based on the input image. Field nodes and virtual connection paths are placed in the two-dimensional external potential field. Dynamic equilibrium is achieved through iterative calculation, and the equilibrium two-dimensional graph structure is output.
4. The method for automatically generating clothing category tags based on deep learning according to claim 3, characterized in that, Fitting the input clothing image to the topological skeleton of the clothing pattern includes: The orientation sensing function of the field node is activated to generate a detection signal, which generates a multi-level image response field of the input clothing image. The detection signal is coupled with the image response field to obtain an initial coupling strength vector set. Deformation intention is generated based on the initial coupling strength vector group. Local deformation proposal is calculated according to the deformation intention and dynamic stiffness protocol. After simulating the execution of the proposal, a new coupling strength is obtained. The image correction feedback signal is generated by comparison to obtain the proposal and feedback pair sequence. Based on the sequence of proposals and feedback, a stable consensus proposal set is found, the consensus proposal set is encoded into a deformation protocol, and the deformation protocol is replayed to drive the deformation of the topology skeleton, thus obtaining the skeleton instance and deformation parameters.
5. The method for automatically generating clothing category tags based on deep learning according to claim 4, characterized in that, Determine the matching product category and output the deformation parameter vector, including: The deformation parameter vector is applied to each candidate skeleton in the clothing pattern topology skeleton library, and physical deformation and relaxation simulation is performed to obtain a set of relaxed candidate skeletons. Key response patterns are extracted from relaxed candidate skeletons, feature signs are analyzed from input clothing images, key response patterns are compared with feature signs, and causal correlation scores are calculated. The relaxed candidate skeleton is projected onto the preset design semantic constraint space for verification, and the consistency between the deformation parameter vector and the design intent of the candidate category is checked to obtain the semantic constraint conflict score. Candidates are selected based on causal correlation score and semantic constraint conflict score. The deformation parameters of the selected candidate skeletons are fine-tuned and calibrated to determine the best matching category and output the calibrated deformation parameter vector.
6. The method for automatically generating clothing category tags based on deep learning according to claim 5, characterized in that, Generate virtual try-on images, including: Establish a continuous texture field for the input clothing area, analyze the texture field to assign it physical properties and record the original lighting reference; A geometric deformation field is generated based on the calibrated deformation parameter vector. The non-rigid response of the texture field under the action of the geometric deformation field is calculated based on the physical properties of the texture field, and the evolved texture field is obtained. A virtual lighting environment is established based on the original lighting reference. The light and shadow information of the deformed skeleton geometry is calculated. The color information of the evolved texture field is combined with the light and shadow information to generate a virtual try-on image.
7. The method for automatically generating clothing category tags based on deep learning according to claim 6, characterized in that, A perturbed skeleton is generated by applying a predetermined offset to the key points of the deformed skeleton, including: The classification gradient of the dual-channel decision network is calculated based on virtual try-on images, and the gradient is inversely mapped into a key point displacement proposal field for the deformed skeleton. Identify the semantic structural vulnerabilities of the current category from the semantic constraint space of the category design, and combine the dynamic stiffness protocol to filter and strengthen the displacement suggestion field to generate a directional perturbation scheme. A structured perturbation sequence is designed based on the directional perturbation scheme, generating a perturbation skeleton and corresponding explanatory labels. A family of perturbation contrast images is then synthesized based on the evolved texture field.
8. The method for automatically generating clothing category tags based on deep learning according to claim 7, characterized in that, A set of perturbation contrast images is generated by combining texture information, including: The geometric deformation of the perturbation skeleton is deconstructed into a local strain field applied to the texture field. Based on the physical properties of the texture field, its state change is predicted by a micromaterial response model to obtain the texture response prediction field. The preliminary results of the texture response prediction field and the perturbation skeleton are combined and input into a dual-channel decision network to obtain intermediate perceptual features. Based on the comparison with the reference features, the texture response prediction field and the perturbation skeleton are adjusted collaboratively and iteratively to obtain the optimized pairing. Based on the explanatory labels, the parameters of the virtual lighting environment are adjusted, and the representation of the micro-geometric details of the surface in the region with pattern density variation in the texture response prediction field is increased during the rendering process, and a perturbation contrast image is synthesized.
9. The method for automatically generating clothing category tags based on deep learning according to claim 8, characterized in that, Output category labels and decoupled confidence vectors, including: Analyze the explanatory labels corresponding to the perturbation skeleton and select the relevant perturbation image subsets. Reconstruct the counterfactual feature representation of the competing categories through feature inverse mapping. Establish a causal intervention calculation layer, which receives the topological features, texture features, counterfactual feature representations and semantic constraint conflict scores of skeleton instances, performs feature replacement interventions and calculates the causal effect matrix.
10. The method for automatically generating clothing category tags based on deep learning according to claim 9, characterized in that, The output includes category labels and decoupled confidence vectors, and also: A dynamic arbitrator is set up to check logical conflicts based on the causal effect matrix and semantic constraint conflict score, dynamically calibrate the original classification probability, and output the calibrated category label and the decoupling confidence vector based on the causal effect strength. Encapsulate a structured chain of decision-making evidence that includes key intermediate data and reasoning processes.