Computer vision-based intelligent garment personalization customization recommendation method and system
By using deep learning networks and 3D virtual try-on technology, a representation of user body features is constructed and a virtual try-on effect image is generated. This solves the shortcomings of existing personalized clothing recommendations in capturing detailed body features and understanding semantic needs, and achieves accurate matching and intuitive display of clothing with users, thereby improving the accuracy of recommendations and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINSILK COM
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-29
Smart Images

Figure CN122115061A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computer vision technology, and more particularly to a method and system for intelligent clothing personalized customization recommendation based on computer vision. Background Technology
[0002] With rising living standards and growing demand for personalized consumption, consumers' need for customized clothing is increasing. Traditional clothing retail models struggle to meet this demand. In recent years, with the rapid development of artificial intelligence technologies such as computer vision and deep learning, personalized clothing customization services based on intelligent recommendations have gained increasing attention. Currently, personalized clothing recommendations primarily analyze users' body characteristics, clothing preferences, and garment style features to provide the most suitable clothing recommendations. This technology not only enhances the user's shopping experience but also helps clothing companies accurately grasp market demand and reduce inventory pressure.
[0003] Existing personalized clothing recommendation technologies have the following shortcomings: First, existing technologies mainly rely on users' historical purchase records and simple body shape data for recommendations, making it difficult to accurately capture the detailed characteristics of users' body shapes, resulting in a low matching degree between recommended clothing and users' actual body shapes. Second, traditional recommendation systems lack a deep understanding of the semantics of users' clothing needs, failing to accurately establish an effective association between users' abstract needs and specific clothing features, causing a gap between the recommendation results and users' expectations. Third, existing technologies cannot intuitively demonstrate the actual effect of recommended clothing on users, making it impossible for users to predict the fit and coordination of clothing before purchasing, increasing the uncertainty of purchasing decisions and reducing user experience and conversion rates. Summary of the Invention
[0004] This invention provides a method and system for intelligent clothing personalized customization recommendation based on computer vision, which can solve the problems in the prior art.
[0005] A first aspect of this invention provides a computer vision-based intelligent clothing personalized customization recommendation method, comprising: The system acquires user body posture image data and clothing demand semantic description, decomposes the detailed dimensional information in the body posture image data through a deep learning network, and constructs a user body posture feature representation; the user body posture feature representation is combined with the clothing demand semantic description and transformed into a unified semantic space to construct a fused feature vector. Based on the fused feature vector, a semantic matching association is established in the pre-constructed clothing style feature library to generate a set of candidate clothing styles that meet the fit constraint. For each clothing style in the candidate clothing style set, the user's body shape features are reconstructed into a three-dimensional human body model through a three-dimensional virtual try-on module, and the clothing style is physically simulated and deformed to generate a virtual try-on effect image. The clothing fit and visual coordination features contained in the virtual try-on effect image are analyzed. Based on the fit characteristics and visual harmony characteristics, combined with the personalized preference patterns formed by the user's historical interaction behavior, the candidate clothing style set is comprehensively scored and ranked, and a personalized recommendation list is output. The personalized recommendation list is presented to the user interface, and feedback signals are generated based on the user's clicks, dwell time and purchase conversion behavior on the recommendation results. The personalized preference mode is then optimized based on the feedback signals.
[0006] The detailed dimensional information in the body posture image data is decomposed using a deep learning network to construct a user body posture feature representation; the user body posture feature representation is then combined with the semantic description of clothing requirements and transformed into a unified semantic space to construct a fused feature vector, including: The body image data is subjected to hierarchical feature decomposition through a multi-level deep learning network. The shallow network extracts the detailed dimension information of local texture and edge, and the deep network extracts the semantic dimension information of overall body shape and posture. The spatial correspondence between the detailed dimension information and the semantic dimension information is established, and the detailed dimension information and semantic dimension information of the same body area are concatenated at the channel level through the feature fusion module to generate a user body feature representation. For each body region feature in the user's body shape feature representation, perform region-level semantic matching with the corresponding clothing part requirements in the clothing requirement semantic description, and calculate the semantic similarity score between each body region feature and the corresponding clothing part requirements. The user's body shape feature representation is weighted and recombined based on the semantic similarity score, and the body shape feature with the highest similarity score is assigned a weight. The weighted and recombined user body shape feature representation and the clothing demand semantic description are projected onto a unified semantic space through a nonlinear mapping function to finally generate a fused feature vector.
[0007] Establishing a spatial correspondence between the detailed dimension information and the semantic dimension information, and then using a feature fusion module to perform channel-level concatenation of the detailed dimension information and semantic dimension information for the same body region to generate a user posture feature representation, including: A body region segmentation mask is constructed based on the body position in the body image data. The body image data is divided into multiple non-overlapping body regions. The detail dimension information and semantic dimension information are mapped to the corresponding body regions according to the body region segmentation mask. For each body region, extract the spatial coordinates of the detailed dimension information and the spatial coordinates of the semantic dimension information within the corresponding body region; The spatial coordinates of the detailed dimension information and the spatial coordinates of the semantic dimension information within the same body area are aligned at the pixel level to establish a one-to-one spatial correspondence. Based on the spatial correspondence, for each body region, the feature channels of the detail dimension information and the feature channels of the semantic dimension information are concatenated along the channel dimension to form a fused feature representation; the fused feature representation is then concatenated and recombined according to the spatial layout of the body region segmentation mask to generate a user body posture feature representation.
[0008] Based on the fused feature vectors, semantic matching associations are established in a pre-constructed clothing style feature library to generate a set of candidate clothing styles that meet the fit constraints, including: The fused feature vector is input into the multi-head semantic matching module. Multiple parallel semantic matching heads are used to calculate the similarity between the fused feature vector and each clothing style feature vector in the clothing style feature library in different semantic subspaces, thereby obtaining multiple semantic subspace similarity components. Based on the weighted fusion of the multiple semantic subspace similarity components, and according to the comprehensive similarity after weighted fusion, clothing styles with a comprehensive similarity greater than a preset similarity threshold are selected from the clothing style feature library to form a preliminary matching clothing style set. For each clothing style in the preliminary matched clothing style set, the user body shape feature representation in the fused feature vector and the clothing size constraint information in the corresponding clothing style feature vector are extracted, and the geometric fit between the two is calculated. The preliminary matching set of clothing styles is filtered based on the geometric fit, and clothing styles whose geometric fit meets the fit constraint are selectively retained, thus generating a candidate set of clothing styles.
[0009] The 3D virtual try-on module reconstructs the user's body features into a 3D human body model, and performs physical simulation deformation on the clothing style to generate a virtual try-on effect image. The analysis of the clothing-human fit and visual coordination features contained in the virtual try-on effect image includes: A three-dimensional body surface mesh is constructed based on the location of key body points in the user's body posture feature representation, and its shape is optimized according to the body contour and local morphological features to form a three-dimensional human body model; based on the three-dimensional human body model, clothing mesh data of clothing styles are obtained from the candidate clothing style set, and the clothing mesh data is initially spatially registered with the three-dimensional body surface mesh; For the clothing mesh data and the three-dimensional body surface mesh after initial spatial registration, physical constraints including collision detection constraints between the two and elastic constraints of fabric material are applied, and physical simulation deformation is performed to obtain steady-state deformation results. The steady-state deformation result is combined with the three-dimensional human body model for rendering to generate a virtual try-on effect image; based on the virtual try-on effect image, the spatial distance field distribution between the steady-state deformation result and the three-dimensional body surface mesh is calculated, and the gap distance and contact pressure distribution are obtained according to the spatial distance field distribution to generate the fit characteristics between the clothing and the human body. Based on the fit characteristics, the color matching degree between the clothing area and the skin area, as well as the coordination degree of the outline proportion between the clothing and the human body, are analyzed to generate overall visual coordination characteristics.
[0010] The steady-state deformation result is combined with the three-dimensional human body model for rendering to generate a virtual try-on effect image; based on the virtual try-on effect image, the spatial distance field distribution between the steady-state deformation result and the three-dimensional body surface mesh is calculated, including: The vertex positions of the clothing mesh in the steady-state deformation result are aligned with the vertex positions of the three-dimensional body surface mesh in the three-dimensional human body model in spatial coordinates, and a lighting, material and texture mapping is applied to the steady-state deformation result, thereby generating a virtual try-on effect image that includes the superimposed view of the steady-state deformation result and the three-dimensional human body model. Based on the virtual try-on effect image, the clothing surface area corresponding to the steady-state deformation result and the human body surface area corresponding to the three-dimensional body surface mesh are extracted respectively. For each vertex of the clothing mesh in the clothing surface region, calculate the Euclidean distance between the vertex and the nearest three-dimensional body surface mesh vertex in the human body surface region, and form a distance scalar field with the clothing mesh vertex as the sampling point; The distance scalar field is spatially interpolated and extended over the surface area of the garment to ultimately generate a spatial distance field distribution that covers the entire surface area of the garment.
[0011] Based on the fit and visual harmony features, combined with the personalized preference patterns formed by the user's historical interaction behavior, the candidate clothing style set is comprehensively scored and ranked, and a personalized recommendation list is output, including: Extract the user's preference intensity for fit and visual harmony for different clothing styles from the user's historical interaction behavior data, and construct a personalized preference pattern based on the relative ratio between the preference intensity for fit and the preference intensity for visual harmony. For each garment style in the candidate garment style set, its corresponding fit characteristics and visual coordination characteristics are mapped to a unified quantization range through numerical range transformation and scale reconstruction; Based on the relative ratio between the fit preference intensity and the visual coordination preference intensity in the personalized preference pattern, the feature weight allocation of the reconstructed fit feature and visual coordination feature in the comprehensive score calculation is dynamically determined. Based on the aforementioned feature weight allocation, the fit feature and the visual coordination feature are combined and superimposed in multiple dimensions to calculate the comprehensive score for each clothing style. All clothing styles in the candidate clothing style set are sorted in descending order according to the comprehensive score, and a preset number of clothing styles with the highest comprehensive score are selected from the results of the descending sort to form a personalized recommendation list.
[0012] A second aspect of this invention provides a computer vision-based intelligent clothing personalized customization recommendation system, comprising: The transformation module is used to acquire the user's body image data and the semantic description of clothing needs, decompose the detailed dimensional information in the body image data through a deep learning network, and construct the user's body feature representation; the user's body feature representation is combined with the semantic description of clothing needs and transformed into a unified semantic space to construct a fused feature vector; The generation module is used to establish semantic matching associations in the pre-built clothing style feature library based on the fused feature vectors, and generate a set of candidate clothing styles that meet the fit constraints. The analysis module is used to reconstruct the user's body shape features into a three-dimensional human body model for each clothing style in the candidate clothing style set through the three-dimensional virtual try-on module, and to perform physical simulation deformation on the clothing style to generate a virtual try-on effect image, and to analyze the clothing-human fit and visual coordination features contained in the virtual try-on effect image. The sorting module is used to comprehensively score and sort the candidate clothing styles based on the fit feature and the visual coordination feature, combined with the personalized preference pattern formed by the user's historical interaction behavior, and output a personalized recommendation list. The optimization module is used to present the personalized recommendation list to the user interface, generate feedback signals based on the user's clicks, dwell time and purchase conversion behavior on the recommendation results, and optimize the personalized preference mode based on the feedback signals.
[0013] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0014] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0015] The beneficial effects of this application are as follows: By decomposing detailed dimensional information in body posture image data through deep learning networks, a more accurate representation of user posture features is constructed, overcoming the limitations of traditional methods that rely solely on simple size parameters, and achieving accurate capture of complex features such as human body curves and postures.
[0016] By transforming user body shape characteristics and semantic descriptions of clothing needs into a unified semantic space and constructing fused feature vectors, the organic combination of body shape data and user intent is achieved, thereby improving the accuracy of the recommendation system's understanding of user needs.
[0017] By reconstructing the user's body features into a three-dimensional human body model and performing physical simulation deformation of clothing, a virtual try-on effect is generated, enabling the system to objectively assess the actual fit between clothing and the human body, greatly improving the adaptability and reliability of the recommendation results.
[0018] By combining fit characteristics, visual harmony characteristics, and user personalized preference patterns to conduct a comprehensive scoring and ranking, the recommendation results take into account the fit, aesthetics, and personal preferences of the clothing, which significantly improves the accuracy of the recommendation results and user satisfaction. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the intelligent clothing personalized customization recommendation method based on computer vision, according to an embodiment of the present invention. Figure 2 This is a flowchart of the personalized recommendation process based on dynamic feature weight allocation of preference intensity according to an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0022] Figure 1This is a flowchart illustrating the intelligent clothing personalized customization recommendation method based on computer vision, as described in an embodiment of the present invention. Figure 1 As shown, the method includes: The system acquires user body posture image data and clothing demand semantic description, decomposes the detailed dimensional information in the body posture image data through a deep learning network, and constructs a user body posture feature representation; the user body posture feature representation is combined with the clothing demand semantic description and transformed into a unified semantic space to construct a fused feature vector. Based on the fused feature vector, a semantic matching association is established in the pre-constructed clothing style feature library to generate a set of candidate clothing styles that meet the fit constraint. For each clothing style in the candidate clothing style set, the user's body shape features are reconstructed into a three-dimensional human body model through a three-dimensional virtual try-on module, and the clothing style is physically simulated and deformed to generate a virtual try-on effect image. The clothing fit and visual coordination features contained in the virtual try-on effect image are analyzed. Based on the fit characteristics and visual harmony characteristics, combined with the personalized preference patterns formed by the user's historical interaction behavior, the candidate clothing style set is comprehensively scored and ranked, and a personalized recommendation list is output. The personalized recommendation list is presented to the user interface, and feedback signals are generated based on the user's clicks, dwell time and purchase conversion behavior on the recommendation results. The personalized preference mode is then optimized based on the feedback signals.
[0023] In one optional implementation, a deep learning network is used to decompose the detailed dimensional information in the body posture image data to construct a user body posture feature representation; the user body posture feature representation is then combined with the clothing demand semantic description and transformed into a unified semantic space to construct a fused feature vector, including: The body image data is subjected to hierarchical feature decomposition through a multi-level deep learning network. The shallow network extracts the detailed dimension information of local texture and edge, and the deep network extracts the semantic dimension information of overall body shape and posture. The spatial correspondence between the detailed dimension information and the semantic dimension information is established, and the detailed dimension information and semantic dimension information of the same body area are concatenated at the channel level through the feature fusion module to generate a user body feature representation. For each body region feature in the user's body shape feature representation, perform region-level semantic matching with the corresponding clothing part requirements in the clothing requirement semantic description, and calculate the semantic similarity score between each body region feature and the corresponding clothing part requirements. The user's body shape feature representation is weighted and recombined based on the semantic similarity score, and the body shape feature with the highest similarity score is assigned a weight. The weighted and recombined user body shape feature representation and the clothing demand semantic description are projected onto a unified semantic space through a nonlinear mapping function to finally generate a fused feature vector.
[0024] Body posture image data is processed through hierarchical feature decomposition using a convolutional neural network architecture. This network comprises two core components: a shallow feature extraction module and a deep semantic understanding module. The shallow feature extraction module employs a three-layer convolutional structure. The first layer has a kernel size of 3×3 pixels, a stride of 1 pixel, zero padding, and 32 kernels, with the modified linear unit (MRU) activation function. The second layer maintains the 3×3 pixel kernel size but increases the stride to 2 pixels and the number of kernels to 64, capturing local texture features within a larger receptive field. The third layer expands to 128 kernels, specifically responsible for extracting body contour edge information and surface texture details. Each convolutional operation is followed by a batch normalization layer, with normalization parameters including a momentum value of 0.9 and an ε value of 1×10⁻⁶. -5 This ensures the stability of the feature distribution.
[0025] The deep semantic understanding module is constructed as a residual network structure, containing four residual block units. Each residual block consists of two convolutional layers with a uniform kernel size of 3×3 pixels. The first residual block has 256 convolutional kernels, which increase sequentially to 512, 1024, and 2048 kernels in subsequent residual blocks. Residual connections employ a skip connection approach, aligning dimensions using 1×1 convolutions when input and output dimensions do not match. A global average pooling layer is placed after the last residual block, compressing the 2D feature map into a 1D semantic feature vector. This deep module is specifically designed to extract high-level semantic features of the overall body contour, body pose angles, and key body parts.
[0026] The spatial correspondence is established through feature map size alignment and position mapping. The feature map size for shallow detail dimension information is 256×256 pixels, while the size for deep semantic dimension information is 64×64 pixels after pooling. A bilinear interpolation algorithm is used to upsample the deep feature map to the same size as the shallow feature map, with interpolation weights calculated based on the distance between the four nearest neighbors. The feature fusion module receives the aligned feature maps and concatenates them along the channel dimension. The shallow feature map has 128 channels, and the deep feature map has 2048 channels; the concatenation generates a fused feature map with 2176 channels.
[0027] The body region segmentation employs a predefined anatomical segmentation scheme, dividing the human body into four main parts: the head region, the trunk region, the upper limb region, and the lower limb region. Each region corresponds to a specific spatial location in the fused feature map: the head region corresponds to the upper quarter of the feature map, the trunk region to the central half, and the upper and lower limb regions to the left and right sides and the lower part, respectively. For each body region, a feature vector corresponding to the spatial location is extracted from the fused feature map; the feature vector has a dimension of 2176. An adaptive average pooling operation is used to uniformly compress irregular region features into a fixed-length region feature representation. The pooling window size is automatically adjusted according to the region area to ensure that different region features have the same representational power.
[0028] The semantic description of clothing demand is encoded using a pre-trained word vector model, with a 300-dimensional vector representation. Clothing demand descriptions are decomposed according to clothing parts, including four categories: tops, trousers, footwear, and accessories. Each category's semantic vector representation is generated using an average word vector method, specifically by averaging the word vectors of all words within that category. For complex descriptions containing multiple adjectives and nouns, an attention weighting mechanism is used to assign importance weights to different words. These attention weights are calculated through a fully connected layer with an input dimension of 300 and an output dimension of 1, using the hyperbolic tangent function as the activation function.
[0029] Semantic similarity is calculated using the cosine similarity metric. The dot product of each body region feature vector and its corresponding clothing part requirement vector is performed, and then divided by the product of the magnitudes of the two vectors to obtain the similarity score. The similarity score ranges from -1 to 1, with values closer to 1 indicating higher similarity. Similarity scores are calculated separately for head region features and accessory requirements, torso region features and upper garment requirements, upper limb region features and upper garment requirements, and lower limb region features and trouser requirements. When a body region corresponds to multiple clothing parts, the highest similarity score is selected as the matching strength index for that region.
[0030] The weighting strategy is implemented based on a normalized softmax function of similarity scores. The similarity scores of all body regions are input into the softmax function to calculate the weight coefficient for each region. The temperature parameter of the softmax function is set to 2.0 to control the smoothness of the weight distribution. The sum of the weight coefficients equals 1.0 to ensure the numerical stability of the weighted recombination process. For the body region with the highest similarity score, an additional weight enhancement factor of 1.5 is applied to highlight the importance of this region in the user's body posture feature representation. Weighted recombination is achieved by multiplying the feature vector of each region by its corresponding weight coefficient and then summing the results to generate a weighted fused user body posture feature representation vector.
[0031] The unified semantic space mapping employs a two-layer fully connected network structure to achieve nonlinear transformation. The first fully connected layer receives the weighted and recombined user body posture feature representation vector, with an input dimension of 2176 and an output dimension of 512. The activation function used is the modified linear unit function. The second fully connected layer has an input dimension of 512 and an output dimension of 256, corresponding to the dimension of the unified semantic space. The clothing demand semantic description vector is also mapped to the 256-dimensional unified semantic space through a fully connected network with the same structure. The parameters of the two mapping networks are trained independently, and the network weights are optimized by comparative learning of the loss function, minimizing the distance between matching body posture features and clothing demands in the semantic space, and maximizing the distance between mismatched feature pairs.
[0032] The generation of the fused feature vector is accomplished through two steps: feature concatenation and attention fusion. The mapped user body shape feature vector and clothing demand semantic vector are directly concatenated along the feature dimension to form an initial 512-dimensional fused vector. A self-attention mechanism is used to calculate the correlation strength between the dimensions within the fused vector. The attention calculation includes three linear transformations: the query matrix, the key matrix, and the value matrix, each with a dimension of 512×512. The attention weights are obtained by performing a softmax normalization after the dot product of the query vector and the key vector. The final fused feature vector is generated by a weighted sum of the attention weights and the value vector.
[0033] In one optional implementation, a spatial correspondence is established between the detailed dimension information and the semantic dimension information, and the detailed dimension information and semantic dimension information of the same body region are concatenated at the channel level through a feature fusion module to generate a user posture feature representation, including: A body region segmentation mask is constructed based on the body position in the body image data. The body image data is divided into multiple non-overlapping body regions. The detail dimension information and semantic dimension information are mapped to the corresponding body regions according to the body region segmentation mask. For each body region, extract the spatial coordinates of the detailed dimension information and the spatial coordinates of the semantic dimension information within the corresponding body region; The spatial coordinates of the detailed dimension information and the spatial coordinates of the semantic dimension information within the same body area are aligned at the pixel level to establish a one-to-one spatial correspondence. Based on the spatial correspondence, for each body region, the feature channels of the detail dimension information and the feature channels of the semantic dimension information are concatenated along the channel dimension to form a fused feature representation; the fused feature representation is then concatenated and recombined according to the spatial layout of the body region segmentation mask to generate a user body posture feature representation.
[0034] The system acquires the user's body posture image data, which can be obtained through depth cameras or multi-view cameras. The body posture image data contains pixel-level information about the user's body posture, providing a foundation for subsequent body region segmentation.
[0035] A body region segmentation mask is constructed based on the body position in the body posture image data. This step uses a semantic segmentation network to process the input body posture image data, dividing the human body into multiple non-overlapping body regions, such as the head, torso, left and right arms, and left and right legs. Specifically, a pre-trained human body parsing network, based on a convolutional neural network architecture, is used. The network takes the original body posture image as input and outputs the corresponding segmentation mask. The segmentation mask is a matrix of the same size as the original image, where the value of each pixel represents the category of the body region to which that location belongs. For example, the pixel value for the head region is 1, the pixel value for the torso region is 2, and so on. These masks will be used in subsequent steps to locate the distribution of detail and semantic dimension information in each body region.
[0036] The acquired detail and semantic dimension information are mapped to their respective body regions using body region segmentation masks. Detail dimension information typically comes from shallow features of convolutional neural networks, offering higher spatial resolution and richer texture details; while semantic dimension information comes from deeper features of the network, possessing stronger semantic expressive power but lower spatial resolution. Specifically, a masking operation is performed on each body region mask: the mask is applied to both the detail and semantic dimension feature maps, extracting their respective feature information within each body region. This step ensures that subsequent feature fusion only processes information from relevant body parts, avoiding feature interference between different body regions.
[0037] For each body region, the spatial coordinates of detail dimension information and semantic dimension information within the corresponding body region are extracted. For detail dimension information, the two-dimensional coordinates (x_d, y_d) of each feature point are recorded; similarly, for semantic dimension information, the two-dimensional coordinates (x_s, y_s) of each feature point are recorded. This coordinate information will be used in the subsequent spatial alignment process.
[0038] Within the same body region, the spatial coordinates of detail dimension information and semantic dimension information are aligned pixel-level to establish a one-to-one spatial correspondence. Since the resolution of semantic dimension information is typically lower than that of detail dimension information, spatial upsampling is required. Bilinear interpolation is used to upsample the semantic dimension information to the same spatial resolution as the detail dimension information, ensuring a one-to-one spatial correspondence between the two types of information. For each pixel location within each body region, a mapping relationship is established between the detail dimension feature vector and the semantic dimension feature vector, preparing for subsequent feature fusion.
[0039] Based on the established spatial correspondence, feature fusion is performed on each body region. For each aligned pixel position, the feature channels of the detail dimension information and the feature channels of the semantic dimension information are concatenated along the channel dimension. Assuming the number of channels for the detail dimension features is C_d and the number of channels for the semantic dimension features is C_s, then the number of feature channels after fusion is C_d + C_s. This channel-level concatenation preserves the complete characteristics of both dimensions of information, which is beneficial for subsequent tasks to fully utilize multi-dimensional features.
[0040] The fused feature representations of each body region are stitched and recombined according to the spatial layout of the body region segmentation mask to generate a complete user body posture feature representation. In specific implementation, a feature map with the same size as the original image is created. Based on the body region label to which each pixel belongs, the fused features of the corresponding region are filled into the corresponding position to form a unified body posture feature representation. For pixels at region boundaries, a smooth transition strategy is adopted to ensure the continuity of the feature representation.
[0041] In practical applications, the generated user posture feature representation can be used for various downstream tasks, such as posture analysis, behavior recognition, and virtual try-on. Because it integrates texture information at the detail level and abstract representation at the semantic level, this feature representation can simultaneously capture both the local details and global structure of the user's posture, improving the performance and accuracy of related applications.
[0042] Furthermore, to improve processing efficiency, a parallel computing framework can be used to simultaneously handle feature extraction and fusion tasks for different body regions. When processing large-scale user data, the modular design of this method also facilitates system expansion and optimization.
[0043] In one optional implementation, establishing semantic matching associations in a pre-built clothing style feature library based on the fused feature vectors to generate a set of candidate clothing styles that meet the fit constraints includes: The fused feature vector is input into the multi-head semantic matching module. Multiple parallel semantic matching heads are used to calculate the similarity between the fused feature vector and each clothing style feature vector in the clothing style feature library in different semantic subspaces, thereby obtaining multiple semantic subspace similarity components. Based on the weighted fusion of the multiple semantic subspace similarity components, and according to the comprehensive similarity after weighted fusion, clothing styles with a comprehensive similarity greater than a preset similarity threshold are selected from the clothing style feature library to form a preliminary matching clothing style set. For each clothing style in the preliminary matched clothing style set, the user body shape feature representation in the fused feature vector and the clothing size constraint information in the corresponding clothing style feature vector are extracted, and the geometric fit between the two is calculated. The preliminary matching set of clothing styles is filtered based on the geometric fit, and clothing styles whose geometric fit meets the fit constraint are selectively retained, thus generating a candidate set of clothing styles.
[0044] The fused feature vector is input into a multi-head semantic matching module, which contains several parallel semantic matching heads. Each matching head is responsible for capturing the degree of matching between the fused feature vector and the clothing style feature vector in a specific semantic subspace. Specifically, eight semantic matching heads can be set, corresponding to different semantic dimensions such as clothing style, color, material, season, occasion, design elements, cut, and category. Each matching head consists of a three-layer fully connected neural network. The input layer receives the fused feature vector, the hidden layer uses the ReLU activation function for non-linear transformation, and the output layer generates the feature map of the corresponding semantic subspace. For each matching head, the cosine similarity between the mapped fused feature vector and each garment in the clothing style feature library in the corresponding semantic subspace is calculated, thus obtaining similarity components across multiple semantic subspaces.
[0045] A weighted fusion method is used based on the similarity components of multiple semantic subspaces, assigning weight coefficients to different semantic subspaces to reflect the importance of each semantic dimension in the overall matching process. Initial weights can be set based on user historical interaction data or expert knowledge, such as a weight of 0.25 for style, 0.2 for color, 0.15 for material, and 0.08 for all other dimensions. The similarities of each subspace are integrated into a comprehensive similarity score through a weighted summation method, which is the sum of the products of each dimension's similarity score and its corresponding weight. Based on the calculated comprehensive similarity score, clothing styles with a comprehensive similarity score greater than a preset similarity threshold (e.g., 0.7) are selected from the clothing style feature library to form a preliminary set of matched clothing styles.
[0046] A geometric fit assessment is performed on the initially matched clothing style set. User body shape features, including key body parameters such as height, weight, shoulder width, chest circumference, waist circumference, and hip circumference, are extracted from the fused feature vector. Simultaneously, clothing size constraint information, such as length, width, and circumference, and their allowable stretch range, is extracted from the corresponding clothing style feature vector. The geometric fit between user body shape features and clothing size constraints is calculated using a multi-dimensional matching algorithm based on fuzzy logic. This algorithm considers the matching degree between different body shape parameters and corresponding clothing sizes, as well as the relative importance of each dimension parameter. Specifically, a fit score is calculated for each body shape parameter and its corresponding clothing size, ranging from 0 to 1, where 1 represents a perfect fit and 0 represents a complete misfit. A score of 1 is awarded for body shape parameter values falling within the allowable range of clothing size; values outside the range are calculated with exponentially decreasing scores based on the degree of deviation. Finally, a weighted average is used to obtain the overall geometric fit.
[0047] The initial set of matched clothing styles is filtered based on geometric fit. Fit constraints are set, such as a geometric fit greater than 0.8. Clothing styles that meet these constraints are selectively retained, thus generating the final set of candidate clothing styles.
[0048] In a practical application scenario, taking a male user who is 175 cm tall and weighs 70 kg as an example, the above method can filter out a set of candidate clothing that matches his body shape and style from a style library containing tens of thousands of garments. First, the fused feature vector includes the user's body shape characteristics and personal style preference characteristics. During the multi-head semantic matching process, the system identifies that the user prefers simple style, dark colors, cotton fabrics, and clothing suitable for business occasions. After semantic matching and comprehensive similarity calculation, 200 style-matching garments are initially selected. Subsequently, a geometric fit evaluation is performed. The system analyzes the user's body shape data, such as shoulder width of 44 cm, chest circumference of 96 cm, and waist circumference of 82 cm, and matches them with the size information of the candidate garments. Finally, the system selects 50 garments with a geometric fit of 0.8 or higher as the final set of candidate garment styles to present to the user.
[0049] The above implementation achieves accurate matching across different semantic dimensions through a multi-head semantic matching mechanism, and combines geometric fit evaluation to ensure the matching degree between clothing and user body shape, effectively improving the accuracy of clothing recommendations and user experience. Furthermore, this method can dynamically adjust the weight coefficients of each semantic dimension based on user feedback, enabling the system to continuously learn and optimize, better meeting personalized recommendation needs.
[0050] In one optional implementation, the user's body shape features are reconstructed into a three-dimensional human body model using a three-dimensional virtual try-on module, and the clothing style is physically simulated and deformed to generate a virtual try-on effect image. The clothing-human fit characteristics and visual coordination characteristics contained in the virtual try-on effect image are analyzed, including: A three-dimensional body surface mesh is constructed based on the location of key body points in the user's body posture feature representation, and its shape is optimized according to the body contour and local morphological features to form a three-dimensional human body model; based on the three-dimensional human body model, clothing mesh data of clothing styles are obtained from the candidate clothing style set, and the clothing mesh data is initially spatially registered with the three-dimensional body surface mesh; For the clothing mesh data and the three-dimensional body surface mesh after initial spatial registration, physical constraints including collision detection constraints between the two and elastic constraints of fabric material are applied, and physical simulation deformation is performed to obtain steady-state deformation results. The steady-state deformation result is combined with the three-dimensional human body model for rendering to generate a virtual try-on effect image; based on the virtual try-on effect image, the spatial distance field distribution between the steady-state deformation result and the three-dimensional body surface mesh is calculated, and the gap distance and contact pressure distribution are obtained according to the spatial distance field distribution to generate the fit characteristics between the clothing and the human body. Based on the fit characteristics, the color matching degree between the clothing area and the skin area, as well as the coordination degree of the outline proportion between the clothing and the human body, are analyzed to generate overall visual coordination characteristics.
[0051] The system acquires the user's body posture representation data, which includes key body point location information, body size parameters, and body contour information. Key body point location information refers to the coordinates of specific locations with significant symbolic meaning in the human anatomical structure, including the three-dimensional spatial coordinates of major skeletal connections and surface feature points such as the vertex of the head, the neck connection point, and the acromion.
[0052] Based on the locations of key body points obtained from the user's body posture feature representation, an initial 3D body surface mesh is constructed. Specifically, using a pre-set human body template mesh, the key points in the user's body posture feature representation are mapped and aligned with the corresponding vertices on the template mesh. A linear blending skinning algorithm is then used to deform the initial mesh according to the mapping relationship, making it conform to the distribution of the user's body key points.
[0053] The initial 3D mesh is optimized based on the user's body contour and local morphological features. Specifically, a local weighted regression algorithm is used to locally adjust the mesh vertices based on key dimensional parameters such as the user's chest, waist, and hip circumference. Simultaneously, a Laplacian smoothing algorithm is applied to smooth the adjusted mesh surface, eliminating discontinuities or jagged edges caused by vertex adjustments, ensuring a natural and smooth surface for the generated 3D human body model. Furthermore, in local areas such as the chest and abdomen, a non-rigid deformation algorithm is applied for fine-tuning based on the user's specific body characteristics, making the final 3D human body model more closely resemble the user's realistic body shape.
[0054] After the 3D human body model is constructed, the clothing mesh data for the desired clothing style is obtained from the candidate clothing style set. The clothing mesh data contains the clothing's geometric shape information, material parameters, and initial size information. Initial spatial registration is performed between the clothing mesh data and the 3D body surface mesh, and a rigid transformation is used to place the clothing mesh in a suitable position near the human body model. Specifically, based on the skeletal structure of the human body model, the corresponding wearing areas of the clothing are identified, such as tops corresponding to the torso and arm areas, and pants corresponding to the lower limb areas. Then, a similarity transformation is used to adjust the size and orientation of the clothing mesh to initially align it with the corresponding human body areas.
[0055] Physical constraints are applied to the clothing mesh data and 3D body surface mesh after initial spatial registration to simulate deformation. First, collision detection constraints are set to prevent the clothing mesh and body mesh from penetrating each other. A layered bounding box algorithm is used to accelerate the collision detection process, and accurate collision response calculations are performed for collision-prone areas. Second, corresponding elastic constraints are set according to the material properties of different fabrics, including parameters such as tensile stiffness, bending stiffness, and shear stiffness. Explicit or implicit integration methods are used to solve the elasticity equations to simulate the natural sag and wrinkle formation process of clothing under gravity and external constraints. Through iterative calculations, when the displacement change of the clothing mesh is less than a preset threshold, the simulation is considered to have reached a steady state, yielding the final steady-state deformation result.
[0056] By combining steady-state deformation results with a 3D human body model for rendering, a virtual try-on effect image is generated. During the rendering process, lighting conditions, material reflection characteristics, and shadow effects are considered to make the generated image more realistic. Simultaneously, multiple angle views can be set, including front, side, and back views, providing a comprehensive display of the try-on effect.
[0057] Based on virtual try-on images, the spatial distance field distribution between the steady-state deformation results and the 3D body surface mesh is calculated. For each vertex on the clothing mesh, the shortest Euclidean distance to the human body surface is calculated to construct a complete spatial distance field. According to the spatial distance field distribution, the distance values are mapped to gap distance and contact pressure distribution. Specifically, a positive distance indicates a gap between the clothing and the skin; a zero distance or a small negative value indicates contact between the clothing and the skin; a large negative distance indicates pressure, with larger negative values indicating greater pressure. These values are visualized using color mapping, for example, blue represents gap areas, green represents comfortable contact areas, and red represents high-pressure areas. Combining this information, the fit characteristics between the clothing and the human body are generated.
[0058] Based on fit characteristics, the color matching degree between the clothing area and the skin area, as well as the harmony of the clothing and the human body's contour proportions, are analyzed to generate overall visual harmony features. In the color matching analysis, the clothing color and the user's skin tone are extracted, and their similarity or complementarity in the HSV color space is calculated to assess the harmony of the color combination. In the contour proportion harmony analysis, the visual proportions of the clothing on the human body are calculated, such as the ratio of top length to bottom length and clothing width to body width, and compared with aesthetic standards such as the golden ratio to assess the harmony of overall proportions. Combining these features, an overall visual harmony score is generated to guide users in choosing clothing styles that better suit them.
[0059] In one optional implementation, the steady-state deformation result is combined with the three-dimensional human body model for rendering to generate a virtual try-on effect image; based on the virtual try-on effect image, the spatial distance field distribution between the steady-state deformation result and the three-dimensional body surface mesh is calculated, including: The vertex positions of the clothing mesh in the steady-state deformation result are aligned with the vertex positions of the three-dimensional body surface mesh in the three-dimensional human body model in spatial coordinates, and a lighting, material and texture mapping is applied to the steady-state deformation result, thereby generating a virtual try-on effect image that includes the superimposed view of the steady-state deformation result and the three-dimensional human body model. Based on the virtual try-on effect image, the clothing surface area corresponding to the steady-state deformation result and the human body surface area corresponding to the three-dimensional body surface mesh are extracted respectively. For each vertex of the clothing mesh in the clothing surface region, calculate the Euclidean distance between the vertex and the nearest three-dimensional body surface mesh vertex in the human body surface region, and form a distance scalar field with the clothing mesh vertex as the sampling point; The distance scalar field is spatially interpolated and extended over the surface area of the garment to ultimately generate a spatial distance field distribution that covers the entire surface area of the garment.
[0060] The vertex positions of the clothing mesh in the steady-state deformation results are aligned with the vertex positions of the 3D body surface mesh in the 3D human body model through coordinate transformation. The coordinate transformation uses a 4×4 transformation matrix, including rotation, translation, and scaling components. The rotation component is represented by Euler angles, with rotation angles around the x, y, and z axes ranging from -180 to 180 degrees. The translation component includes offsets in the x, y, and z directions, with values determined according to the model size, typically within the range of -500 to 500 mm. The scaling component uses uniform scaling, with scaling factors ranging from 0.8 to 1.2. During the coordinate alignment process, the coordinates of each vertex of the clothing mesh are transformed to the same coordinate system as the human body mesh through transformation matrix operations, maintaining an alignment accuracy of 0.1 mm.
[0061] The application of lighting, material, and texture mapping consists of two parts: lighting settings and material property definitions. Lighting settings include a main light source, an auxiliary light source, and an ambient light source. The main light source uses a directional light, with a set direction vector and an intensity of 1200 lumens. The auxiliary light source uses a point light, with its position coordinates adjusted according to the scene requirements, and an intensity of 800 lumens. The ambient light intensity is set to 200 lumens. Material property definitions include diffuse reflection coefficient, specular reflection coefficient, and roughness parameters. The diffuse reflection coefficient for clothing materials is set to 0.8, the specular reflection coefficient to 0.1, and the roughness is adjusted within the range of 0.3 to 0.9 depending on the fabric type. Texture mapping maps the 3D surface to a 2D texture space using UV coordinates, with UV coordinate values ranging from 0 to 1.
[0062] The virtual try-on effect image generated by the overlay view is achieved through depth testing and color blending. The depth buffer resolution is set to 2048×2048 pixels, storing depth values ranging from 0.1 to 1000 mm. During rendering, the human body mesh and clothing mesh are sorted by depth according to their distance from the camera, with closer pixels covering farther pixels. Color blending is processed based on material transparency; the transparency value for opaque fabric is set to 1.0, and the transparency value for semi-transparent fabric is in the range of 0.7 to 0.9. The rendered output is a virtual try-on effect image in RGB format with a resolution of 1024×1024 pixels and 8-bit precision for each color channel.
[0063] In virtual try-on images, clothing surface area extraction is achieved through pixel classification and region recognition. Pixel classification distinguishes clothing materials from human skin based on color and depth information. A color threshold range is set in the RGB color space to identify clothing areas, with a color tolerance of 20 gray levels. Depth information assists in segmentation, as the clothing surface has a depth difference of 0.5 to 5 millimeters relative to the human body surface. Region recognition uses connectivity analysis to label continuous clothing pixel areas, filtering out noise areas with an area smaller than 100 pixels. Human body surface area extraction identifies human body surface pixels through skin color detection and depth analysis. Skin color range is defined in the HSV color space, with the hue component ranging from 20 to 35 degrees and the saturation component ranging from 30 to 80.
[0064] The Euclidean distance between a vertex of the clothing mesh and the nearest 3D body surface mesh vertex in the human body surface region is calculated using a nearest neighbor search. For each vertex of the clothing mesh, all 3D body surface mesh vertices in the human body surface region are traversed, and the Euclidean distance between the two points is calculated. The Euclidean distance is calculated by taking the square root of the sum of the squares of the differences in the coordinates of the two points, maintaining a calculation accuracy of 0.01 mm. During the search process, the minimum distance value corresponding to each clothing mesh vertex and the coordinates of the nearest neighbor point are recorded. The distance calculation results form a distance scalar field indexed by the clothing mesh vertices, with each sampling point in the scalar field storing the corresponding distance value.
[0065] Spatial interpolation extension of the distance scalar field on the garment surface is achieved through an interpolation method. Weighted average interpolation is used, with weights inversely proportional to distance. During the interpolation extension process, interpolation query points are generated within the garment surface area, with a density of 4 points per square millimeter. The distance value of each query point is calculated by weighted averaging of surrounding sampling points. The weights are calculated based on the spatial distance from the query point to the sampling point; closer points have larger weights. Weight normalization ensures that the sum of all weights equals 1. Interpolation boundary handling ensures the continuity of the distance field at the region edges through boundary condition constraints.
[0066] Complete coverage of the spatial distance field distribution is achieved through sampling point layout and interpolation expansion. The garment surface area is divided into regions with different sampling densities according to geometric complexity: smooth regions use a lower sampling density (2 sampling points per square millimeter), while complex regions use a higher sampling density (8 sampling points per square millimeter). The sampling point distribution ensures coverage of all locations on the garment surface area, avoiding unsampled areas. Interpolation expansion extends the discrete sampling point distance values into a continuous distance field distribution, and the interpolation results are stored in a grid structure for easy subsequent querying and processing.
[0067] In one optional implementation, based on the fit characteristics and visual harmony characteristics, combined with the personalized preference pattern formed by the user's historical interaction behavior, the candidate clothing style set is comprehensively scored and ranked, and a personalized recommendation list is output, including: Extract the user's preference intensity for fit and visual harmony for different clothing styles from the user's historical interaction behavior data, and construct a personalized preference pattern based on the relative ratio between the preference intensity for fit and the preference intensity for visual harmony. For each garment style in the candidate garment style set, its corresponding fit characteristics and visual coordination characteristics are mapped to a unified quantization range through numerical range transformation and scale reconstruction; Based on the relative ratio between the fit preference intensity and the visual coordination preference intensity in the personalized preference pattern, the feature weight allocation of the reconstructed fit feature and visual coordination feature in the comprehensive score calculation is dynamically determined. Based on the aforementioned feature weight allocation, the fit feature and the visual coordination feature are combined and superimposed in multiple dimensions to calculate the comprehensive score for each clothing style. All clothing styles in the candidate clothing style set are sorted in descending order according to the comprehensive score, and a preset number of clothing styles with the highest comprehensive score are selected from the results of the descending sort to form a personalized recommendation list.
[0068] like Figure 2 As shown, the method includes: User historical interaction data is collected and stored through a behavior recording module. This module records user interactions such as clicking, browsing, adding to favorites, purchasing, and rating different clothing styles. The interaction behavior data structure includes fields such as user identifier, clothing style identifier, interaction type, interaction timestamp, dwell time, and interaction intensity. The user identifier uses a 32-bit integer encoding, the clothing style identifier uses a 64-bit string encoding, and the interaction type is represented by an enumeration value, including browsing, adding to favorites, purchasing, and rating interactions. The interaction intensity is weighted according to the interaction type: browsing interaction has a weight of 1, adding to favorites has a weight of 3, purchasing interaction has a weight of 5, and the weight of rating interaction is dynamically adjusted within the range of 1 to 5 based on the rating score.
[0069] The extraction of fit preference intensity is achieved by statistically analyzing the frequency and intensity of user interactions with clothing of different fit characteristics. Fit characteristics are categorized into three types based on the closeness of the clothing to the body: loose, moderate, and tight. The fit characteristic value for loose clothing ranges from 0 to 0.3, for moderate clothing from 0.3 to 0.7, and for tight clothing from 0.7 to 1.0. The user's preference intensity for each fit category is calculated by weighted summation, with the weight being the interaction intensity. The summation result is divided by the total number of interactions for that category to obtain the average preference intensity. Preference intensity normalization ensures that the values are within the range of 0 to 1. The normalization method uses maximum value normalization, dividing the preference intensity of each category by the maximum preference intensity among all categories.
[0070] Visual harmony preference intensity extraction is achieved by analyzing user interaction patterns with clothing of different visual harmony characteristics. Visual harmony characteristics are categorized into three levels—low harmony, medium harmony, and high harmony—based on color matching and proportion coordination. The visual harmony characteristic value for low harmony clothing ranges from 0 to 0.33, for medium harmony clothing from 0.33 to 0.66, and for high harmony clothing from 0.66 to 1.0. The calculation method for visual harmony preference intensity is the same as for fit preference intensity, using weighted summation and normalization to obtain the preference intensity value for each harmony level.
[0071] Personalized preference patterns are constructed based on the relative ratio between the strength of fit preference and the strength of visual harmony preference. This ratio is determined by calculating the ratio of the two preference strengths, ranging from 0.1 to 10. When the fit preference strength is greater than the visual harmony preference strength, the ratio is greater than 1, indicating that the user prioritizes the fit of the clothing. When the visual harmony preference strength is greater than the fit preference strength, the ratio is less than 1, indicating that the user prioritizes the visual effect of the clothing. The ratio is stored in a preference pattern data structure, which includes fields such as user identifier, fit preference weight, visual harmony preference weight, and weight update timestamp. Weight updates employ a sliding window mechanism with a window length of 30 days to ensure that the preference pattern reflects the user's latest preference changes.
[0072] The numerical range transformation between fit and visual coordination features is achieved through linear mapping. The original value range of the fit feature is determined according to the specific calculation method, typically between 0 and 100. The original value range of the visual coordination feature is also between 0 and 100. The numerical range transformation maps the value ranges of both features to a standardized range of 0 to 1. The transformation formula is: the new value equals the original value minus the minimum value, divided by the difference between the maximum and minimum values. During the transformation process, the original minimum and maximum values of each feature are recorded for subsequent reverse transformation and numerical interpretation.
[0073] Scale reconstruction achieves scale uniformity for feature values through Z-score standardization. Z-score standardization calculates the difference between each feature value and its mean, then divides it by the feature's standard deviation. The standardized feature values have a mean of 0 and a standard deviation of 1, eliminating dimensional differences between features. The reconstruction process requires pre-calculating the mean and standard deviation of the fit and visual harmony features across the candidate clothing style set. The mean is calculated by summing the feature values of all clothing styles and dividing by the number of clothing styles. The standard deviation is calculated by taking the square root of the sum of the squares of the differences between feature values and the mean, divided by the number of clothing styles.
[0074] Feature weights are dynamically determined based on the relative proportions within personalized preference patterns. A normalized allocation strategy is employed to ensure that the sum of the fit feature weight and the visual coordination feature weight equals 1. When the ratio of fit preference intensity to visual coordination preference intensity is R, the fit feature weight is calculated as R divided by R plus 1, and the visual coordination feature weight is calculated as 1 divided by R plus 1. The weight calculation results are preserved to three decimal places. A minimum weight threshold of 0.1 is set during the weight allocation process to prevent feature information loss due to excessively small feature weights.
[0075] Multi-dimensional feature combination is achieved through a weighted linear combination method. The combination process multiplies the reconstructed fit feature value by its corresponding weight, and the reconstructed visual coordination feature value by its corresponding weight. The two weighted results are then added together to obtain the combined feature value. Intensity superposition applies an intensity adjustment factor to the feature combination. This adjustment factor is determined based on the activity level of user interaction. Users with active interaction are given an intensity adjustment factor of 1.2, while users with less interaction are given an intensity adjustment factor of 0.8, ensuring that the recommendation results match the user's level of engagement.
[0076] The comprehensive score calculation multiplies the result of combining multi-dimensional features with the result of intensity superposition to obtain the final score. The score ranges from 0 to 10, and the score precision is retained to two decimal places. A smoothing process is applied during the score calculation to avoid the score distribution being too concentrated or dispersed. The smoothing process uses the Sigmoid function, which maps the linear combination result to the interval between 0 and 1, and then multiplies it by 10 to obtain the final score. The parameters of the Sigmoid function include a gain coefficient of 2 and an offset of 0 to ensure the reasonableness of the score distribution.
[0077] The descending sort is implemented using the quicksort algorithm, with the sorting key being the overall score. Quicksort has a time complexity of O(nlogn), making it suitable for sorting large sets of candidate clothing styles. During the sorting process, clothing styles with the same score are handled by using the lexicographical order of their identifiers as a secondary sorting key to ensure the stability and reproducibility of the sorting results. The sorted results are stored in an ordered list data structure, supporting efficient sequential access and range query operations.
[0078] The personalized recommendation list is constructed by selecting a preset number of clothing styles from the descendingly sorted results. This preset number is determined based on the application scenario and user preferences, with a default value of 10 styles and an adjustable range of 5 to 50. The selection process starts from the first style in the sorted results and sequentially selects the clothing styles with the highest overall scores until the preset number is reached. The selected results include clothing style identifiers, overall scores, fit feature values, visual harmony feature values, and recommendation reasons. The recommendation reasons are automatically generated based on feature weight allocation and score contribution, providing explanatory information for the recommendation decision.
[0079] A second aspect of this invention provides a computer vision-based intelligent clothing personalized customization recommendation system, comprising: The transformation module is used to acquire the user's body image data and the semantic description of clothing needs, decompose the detailed dimensional information in the body image data through a deep learning network, and construct the user's body feature representation; the user's body feature representation is combined with the semantic description of clothing needs and transformed into a unified semantic space to construct a fused feature vector; The generation module is used to establish semantic matching associations in the pre-built clothing style feature library based on the fused feature vectors, and generate a set of candidate clothing styles that meet the fit constraints. The analysis module is used to reconstruct the user's body shape features into a three-dimensional human body model for each clothing style in the candidate clothing style set through the three-dimensional virtual try-on module, and to perform physical simulation deformation on the clothing style to generate a virtual try-on effect image, and to analyze the clothing-human fit and visual coordination features contained in the virtual try-on effect image. The sorting module is used to comprehensively score and sort the candidate clothing styles based on the fit feature and the visual coordination feature, combined with the personalized preference pattern formed by the user's historical interaction behavior, and output a personalized recommendation list. The optimization module is used to present the personalized recommendation list to the user interface, generate feedback signals based on the user's clicks, dwell time and purchase conversion behavior on the recommendation results, and optimize the personalized preference mode based on the feedback signals.
[0080] A third aspect of the present invention provides an electronic device, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0081] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0082] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A computer vision-based intelligent clothing personalized customization recommendation method, characterized in that, include: The system acquires user body posture image data and clothing demand semantic description, decomposes the detailed dimensional information in the body posture image data through a deep learning network, and constructs a user body posture feature representation; the user body posture feature representation is combined with the clothing demand semantic description and transformed into a unified semantic space to construct a fused feature vector. Based on the fused feature vector, a semantic matching association is established in the pre-constructed clothing style feature library to generate a set of candidate clothing styles that meet the fit constraint. For each clothing style in the candidate clothing style set, the user's body shape features are reconstructed into a three-dimensional human body model through a three-dimensional virtual try-on module, and the clothing style is physically simulated and deformed to generate a virtual try-on effect image. The clothing fit and visual coordination features contained in the virtual try-on effect image are analyzed. Based on the fit characteristics and visual harmony characteristics, combined with the personalized preference patterns formed by the user's historical interaction behavior, the candidate clothing style set is comprehensively scored and ranked, and a personalized recommendation list is output. The personalized recommendation list is presented to the user interface, and feedback signals are generated based on the user's clicks, dwell time and purchase conversion behavior on the recommendation results. The personalized preference mode is then optimized based on the feedback signals.
2. The method according to claim 1, characterized in that, The detailed dimensional information in the body posture image data is decomposed using a deep learning network to construct a user body posture feature representation; the user body posture feature representation is then combined with the semantic description of clothing requirements and transformed into a unified semantic space to construct a fused feature vector, including: The body image data is subjected to hierarchical feature decomposition through a multi-level deep learning network. The shallow network extracts the detailed dimension information of local texture and edge, and the deep network extracts the semantic dimension information of overall body shape and posture. The spatial correspondence between the detailed dimension information and the semantic dimension information is established, and the detailed dimension information and semantic dimension information of the same body area are concatenated at the channel level through the feature fusion module to generate a user body feature representation. For each body region feature in the user's body shape feature representation, perform region-level semantic matching with the corresponding clothing part requirements in the clothing requirement semantic description, and calculate the semantic similarity score between each body region feature and the corresponding clothing part requirements. The user's body shape feature representation is weighted and recombined based on the semantic similarity score, and the body shape feature with the highest similarity score is assigned a weight. The weighted and recombined user body shape feature representation and the clothing demand semantic description are projected onto a unified semantic space through a nonlinear mapping function to finally generate a fused feature vector.
3. The method according to claim 2, characterized in that, Establishing a spatial correspondence between the detailed dimension information and the semantic dimension information, and then using a feature fusion module to perform channel-level concatenation of the detailed dimension information and semantic dimension information for the same body region to generate a user posture feature representation, including: A body region segmentation mask is constructed based on the body position in the body image data. The body image data is divided into multiple non-overlapping body regions. The detail dimension information and semantic dimension information are mapped to the corresponding body regions according to the body region segmentation mask. For each body region, extract the spatial coordinates of the detailed dimension information and the spatial coordinates of the semantic dimension information within the corresponding body region; The spatial coordinates of the detailed dimension information and the spatial coordinates of the semantic dimension information within the same body area are aligned at the pixel level to establish a one-to-one spatial correspondence. Based on the spatial correspondence, for each body region, the feature channels of the detail dimension information and the feature channels of the semantic dimension information are concatenated along the channel dimension to form a fused feature representation; the fused feature representation is then concatenated and recombined according to the spatial layout of the body region segmentation mask to generate a user body posture feature representation.
4. The method according to claim 1, characterized in that, Based on the fused feature vectors, semantic matching associations are established in a pre-constructed clothing style feature library to generate a set of candidate clothing styles that meet the fit constraints, including: The fused feature vector is input into the multi-head semantic matching module. Multiple parallel semantic matching heads are used to calculate the similarity between the fused feature vector and each clothing style feature vector in the clothing style feature library in different semantic subspaces, thereby obtaining multiple semantic subspace similarity components. Based on the weighted fusion of the multiple semantic subspace similarity components, and according to the comprehensive similarity after weighted fusion, clothing styles with a comprehensive similarity greater than a preset similarity threshold are selected from the clothing style feature library to form a preliminary matching clothing style set. For each clothing style in the preliminary matched clothing style set, the user body shape feature representation in the fused feature vector and the clothing size constraint information in the corresponding clothing style feature vector are extracted, and the geometric fit between the two is calculated. The preliminary matching set of clothing styles is filtered based on the geometric fit, and clothing styles whose geometric fit meets the fit constraint are selectively retained, thus generating a candidate set of clothing styles.
5. The method according to claim 1, characterized in that, The 3D virtual try-on module reconstructs the user's body features into a 3D human body model, and performs physical simulation deformation on the clothing style to generate a virtual try-on effect image. The analysis of the clothing-human fit and visual coordination features contained in the virtual try-on effect image includes: A three-dimensional body surface mesh is constructed based on the location of key body points in the user's body posture feature representation, and its shape is optimized according to the body contour and local morphological features to form a three-dimensional human body model; based on the three-dimensional human body model, clothing mesh data of clothing styles are obtained from the candidate clothing style set, and the clothing mesh data is initially spatially registered with the three-dimensional body surface mesh; For the clothing mesh data and the three-dimensional body surface mesh after initial spatial registration, physical constraints including collision detection constraints between the two and elastic constraints of fabric material are applied, and physical simulation deformation is performed to obtain steady-state deformation results. The steady-state deformation result is combined with the three-dimensional human body model for rendering to generate a virtual try-on effect image; based on the virtual try-on effect image, the spatial distance field distribution between the steady-state deformation result and the three-dimensional body surface mesh is calculated, and the gap distance and contact pressure distribution are obtained according to the spatial distance field distribution to generate the fit characteristics between the clothing and the human body. Based on the fit characteristics, the color matching degree between the clothing area and the skin area, as well as the coordination degree of the outline proportion between the clothing and the human body, are analyzed to generate overall visual coordination characteristics.
6. The method according to claim 5, characterized in that, The steady-state deformation results are combined with the three-dimensional human body model for rendering to generate a virtual try-on effect image; Based on the virtual try-on effect image, the spatial distance field distribution between the steady-state deformation result and the three-dimensional body surface mesh is calculated as follows: The vertex positions of the clothing mesh in the steady-state deformation result are aligned with the vertex positions of the three-dimensional body surface mesh in the three-dimensional human body model in spatial coordinates, and a lighting, material and texture mapping is applied to the steady-state deformation result, thereby generating a virtual try-on effect image that includes the superimposed view of the steady-state deformation result and the three-dimensional human body model. Based on the virtual try-on effect image, the clothing surface area corresponding to the steady-state deformation result and the human body surface area corresponding to the three-dimensional body surface mesh are extracted respectively. For each vertex of the clothing mesh in the clothing surface region, calculate the Euclidean distance between the vertex and the nearest three-dimensional body surface mesh vertex in the human body surface region, and form a distance scalar field with the clothing mesh vertex as the sampling point; The distance scalar field is spatially interpolated and extended over the surface area of the garment to ultimately generate a spatial distance field distribution that covers the entire surface area of the garment.
7. The method according to claim 1, characterized in that, Based on the fit and visual harmony features, combined with the personalized preference patterns formed by the user's historical interaction behavior, the candidate clothing style set is comprehensively scored and ranked, and a personalized recommendation list is output, including: Extract the user's preference intensity for fit and visual harmony for different clothing styles from the user's historical interaction behavior data, and construct a personalized preference pattern based on the relative ratio between the preference intensity for fit and the preference intensity for visual harmony. For each garment style in the candidate garment style set, its corresponding fit characteristics and visual coordination characteristics are mapped to a unified quantization range through numerical range transformation and scale reconstruction; Based on the relative ratio between the fit preference intensity and the visual coordination preference intensity in the personalized preference pattern, the feature weight allocation of the reconstructed fit feature and visual coordination feature in the comprehensive score calculation is dynamically determined. Based on the aforementioned feature weight allocation, the fit feature and the visual coordination feature are combined and superimposed in multiple dimensions to calculate the comprehensive score for each clothing style. All clothing styles in the candidate clothing style set are sorted in descending order according to the comprehensive score, and a preset number of clothing styles with the highest comprehensive score are selected from the results of the descending sort to form a personalized recommendation list.
8. A computer vision-based intelligent clothing personalized customization recommendation system, used to implement the method of any one of claims 1-7, characterized in that, include: The transformation module is used to acquire the user's body image data and the semantic description of clothing needs, decompose the detailed dimensional information in the body image data through a deep learning network, and construct the user's body feature representation; the user's body feature representation is combined with the semantic description of clothing needs and transformed into a unified semantic space to construct a fused feature vector; The generation module is used to establish semantic matching associations in the pre-built clothing style feature library based on the fused feature vectors, and generate a set of candidate clothing styles that meet the fit constraints. The analysis module is used to reconstruct the user's body shape features into a three-dimensional human body model for each clothing style in the candidate clothing style set through the three-dimensional virtual try-on module, and to perform physical simulation deformation on the clothing style to generate a virtual try-on effect image, and to analyze the clothing-human fit and visual coordination features contained in the virtual try-on effect image. The sorting module is used to comprehensively score and sort the candidate clothing styles based on the fit feature and the visual coordination feature, combined with the personalized preference pattern formed by the user's historical interaction behavior, and output a personalized recommendation list. The optimization module is used to present the personalized recommendation list to the user interface, generate feedback signals based on the user's clicks, dwell time and purchase conversion behavior on the recommendation results, and optimize the personalized preference mode based on the feedback signals.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.