Virtual fitting personalized clothing recommendation method based on artificial intelligence
By introducing multimodal semantic map modeling, cross-modal comparison learning network and virtual fitting image generation technology into the virtual fitting and personalized clothing recommendation system, combined with behavior feedback-driven optimization algorithms, the technical defects of the existing system in recommendation accuracy, trial-on effect and personalized matching are solved, and efficient, personalized and intelligent clothing recommendations are achieved.
Patent Information
- Application Number
- CN202510630350.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing virtual fitting and personalized clothing recommendation systems have significant technical defects in the integration of recommendation algorithms and user body modeling, semantic map modeling and multimodal feature alignment, model structure optimization, closed-loop feedback mechanism and fitting image generation, and it is difficult to achieve personalized, visualized and intelligent clothing recommendations.
Using the virtual fitting personalized clothing recommendation method based on artificial intelligence, it integrates multimodal semantic map modeling, cross-modal comparison learning network, virtual fitting image generation technology and behavioral feedback-driven optimization algorithms to extract features from user images, text descriptions, body shape parameters and behavioral data, generate personalized recommendations and multi-angle try-on images, and optimize the recommendation model through closed-loop feedback.
It has achieved the ability to have high recommendation accuracy, authentic try-on effect, strong personality matching and system adaptive optimization, breaking through the technical bottlenecks of the personality adaptability, visual interaction and self-optimization capabilities of traditional recommendation systems.
Smart Images

Figure CN120146972A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a personalized clothing recommendation method for virtual fitting based on artificial intelligence. Background Art
[0002] With the rapid development of e-commerce and mobile Internet, online clothing retail platforms have been rapidly popularized, and users are increasingly inclined to purchase clothing products online. The integration of virtual fitting technology and personalized recommendation technology has become one of the key paths to improve the user experience of e-commerce platforms. Traditional e-commerce platforms mainly rely on rule matching or recommendation systems based on collaborative filtering algorithms to recommend clothing related to users' historical browsing and purchase behaviors. However, such recommendation methods show significant limitations in dealing with problems such as user interest migration, body type difference adaptation, and cold start of new users, and it is difficult to effectively support the personalized, visual, and intelligent clothing recommendation requirements.
[0003] On the other hand, the early development of virtual fitting technology mainly relied on methods based on static templates or simple image synthesis, which performed position overlay or image mapping processing on the user-uploaded avatar or photo and the commodity clothing image. This method not only lacks a sense of reality and immersion, but also cannot accurately simulate the fitting effect according to the user's actual body type parameters, resulting in strong visual deception and low practical reference value of the fitting image. At the same time, existing fitting systems are often isolated from the recommendation system and fail to achieve an end-to-end integrated recommendation process based on user preferences and body type characteristics. In other words, the virtual fitting system only exists as a display module and does not participate in the clothing screening and intelligent matching process.
[0004] In recent years, with the maturity of deep learning technologies, especially technologies such as graph neural networks, contrastive learning, and self-supervised learning, more and more research has attempted to introduce multi-modal feature fusion, semantic graph modeling, and personalized embedding representation methods to improve the accuracy of recommendation systems and user satisfaction. For example, some research uses graph neural networks to perform embedding learning on the relationship graph between users and commodities, or jointly embeds and trains images and texts to improve the semantic expression ability of recommendation results. However, most of these methods focus on the accuracy optimization of "recommendation itself" and have not fully considered how the recommendation results are linked and verified with the user's actual dressing experience, lacking a closed-loop integration mechanism from "recommendation" to "fitting".
[0005] In addition, traditional recommendation systems generally lack the ability to dynamically optimize the model structure. Once the structural parameters and training hyperparameters of the model are set, it is difficult to dynamically adjust them according to feedback during the recommendation task. The current system lacks an optimization algorithm framework that is deeply coupled with user behavior feedback. When facing complex task scenarios (such as multi-layer semantic alignment, multi-angle image generation, multi-objective recommendation effect evaluation), it often falls into problems such as local optimality or unstable convergence, which limits the generalization ability and adaptive generalization ability of the model.
[0006] Furthermore, existing recommendation systems often ignore the modeling of the body shape differences of users during the recommendation process. Most systems default that users have a stable and consistent acceptance of clothing sizes and fits, ignoring the direct relationship between body shape and the try-on effect, resulting in the generated recommendation results lacking persuasiveness and credibility during the try-on stage. Even if some systems allow users to select body shape types, their modeling methods are very rough, only staying at the level of text annotation or template selection, and cannot truly reflect the impact of body shape parameters on the quality of try-on image generation.
[0007] More notably, the current interaction feedback mechanism between virtual try-on and recommendation systems is not yet perfect. The behavioral data such as clicks, browsing, ratings, and collections generated by users after viewing try-on images are not effectively collected and modeled, and even less can they act in reverse on the evolution of the semantic graph structure and the reconstruction of training samples, resulting in the system lacking the ability of continuous learning and personalized evolution. Most traditional recommendation systems make static recommendations based on static models and are difficult to form a closed-loop structure of "recommendation - feedback - optimization - re-recommendation".
[0008] In summary, the existing virtual try-on and personalized clothing recommendation systems have significant technical defects and deficiencies in the following aspects: First, there is a lack of deep integration between the recommendation algorithm and user body shape modeling, resulting in the lack of suitability of the recommendation results for real wearing; second, the semantic graph modeling and embedding method fail to be efficiently aligned with user multi-modal features (image, text, behavior); third, the recommendation model lacks a dynamic optimization mechanism at the structural level and cannot adaptively adjust the structure and parameter configuration according to feedback data; fourth, there is a lack of an end-to-end closed-loop feedback update mechanism, and user behavior feedback fails to be effectively used for the dynamic adjustment of the graph structure and sample strategy; fifth, the virtual try-on image generation process does not combine semantic recommendation results for joint modeling and lacks targeted and personalized rendering capabilities.
[0009] Therefore, how to provide an artificial intelligence-based virtual try-on personalized clothing recommendation method is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention
[0010] An object of the present invention is to provide a personalized clothing recommendation method for virtual fitting based on artificial intelligence. The present invention integrates multimodal semantic graph modeling, cross-modal contrast learning network, virtual fitting image generation technology and behavior feedback-driven optimization algorithm, and details how to realize a closed-loop process of personalized clothing recommendation and multi-angle fitting image generation based on the user's image data, text description, body type parameters and interaction behavior, and has the advantages of high recommendation accuracy, real fitting effect, strong personality matching and strong system adaptive optimization ability.
[0011] A personalized clothing recommendation method for virtual fitting based on artificial intelligence according to an embodiment of the present invention includes the following steps: S1. Collect the user's image data, text description data and the user's historical interaction behavior data to generate a multimodal user feature set; S2. Construct a multimodal heterogeneous semantic graph, and perform modeling and embedding representation on the multimodal heterogeneous semantic graph structure; S3. Extract the structured embedding vectors of the user nodes and clothing nodes from the multimodal heterogeneous semantic graph, and input the structured embedding vectors and the multimodal user feature set into the cross-modal contrast learning network model, and perform semantic alignment training in the shared embedding space by constructing positive and negative sample pairs; S4. Optimize the structure parameters and training hyperparameters of the cross-modal contrast learning network model through the raccoon optimization algorithm to generate an optimized cross-modal contrast learning network model; S5. Apply the optimized cross-modal contrast learning network model to the recommendation task, calculate the semantic matching degree score between the user and the clothing, generate a recommendation candidate set, input the recommendation candidate set and the user body type parameters into the virtual fitting image generation unit, and combine the clothing image to generate a realistic image of the user wearing the recommended clothing, and output an interactive multi-angle fitting view; S6. Collect the user's behavior feedback data on the fitting image to generate a user behavior feedback feature vector, which is used to update the edge weights in the multimodal heterogeneous semantic graph and the composition of the training samples of the cross-modal contrast learning network, and periodically execute steps S2 to S5 to form an adaptive iterative optimization closed-loop personalized recommendation process.
[0012] Optionally, the multimodal user feature set is generated by fusing visual feature vectors, semantic feature vectors and behavior feature vectors, wherein the visual feature vectors are extracted by inputting the user's image data into a convolutional neural network, the semantic feature vectors are extracted by inputting the text description data into a language understanding unit, and the behavior feature vectors are generated by encoding the user's historical behavior data.
[0013] Optionally, the S2 specifically includes: S21. Construct a node set, including user nodes, clothing nodes, and clothing attribute nodes. The user nodes are used to represent user objects with individual identifiers, the clothing nodes are used to represent target clothing objects that can be recommended, and the clothing attribute nodes are used to represent label information such as the style, color, season, and brand of the clothing; S22. Based on the user's historical interaction behavior data, establish an edge relationship between the user node and the clothing node. The edge relationship is used to represent the behavioral associations of click, favorite, purchase, and try-on between the user and the clothing; S23. Based on the clothing metadata and label information, establish an edge relationship between the clothing node and the clothing attribute node. The edge relationship is used to represent the attribute associations of the style, color, and brand possessed by the clothing; S24. Set an initial edge weight for the constructed edge relationship. The edge weight is set according to the association strength between the user interaction frequency, behavior type, and similar labels. The edge weight is used to adjust the recommendation path and the graph neural propagation weight; S25. Use the heterogeneous graph modeling method to perform structural modeling on the multi-modal heterogeneous semantic graph, so that the heterogeneous relationship information is retained between various types of nodes, and at the same time, a complete graph structure representation is established; S26. Based on the graph neural network structure, perform an embedding representation on the multi-modal heterogeneous semantic graph, and encode various types of nodes in the multi-modal heterogeneous semantic graph into structured embedding vectors through a multi-layer information aggregation mechanism. The embedding vectors are used as the input of the cross-modal contrast learning network model.
[0014] Optionally, the specific steps of S3 include: S31. Extract the structured embedding vectors of the user node and the clothing node from the multi-modal heterogeneous semantic graph, which are respectively represented as and , and extract the fused representation vector from the multi-modal user feature set; S32. Introduce the graph fusion gating coefficient , construct a non-linear gating fusion mechanism, and perform nested fusion on the graph structure semantics and modal embeddings to generate the user representation vector : ; Among them, is the Sigmoid function, is the element-wise multiplication, is the multi-layer perceptron network, is the graph fusion gating coefficient; S33. Combine the user representation vector with the clothing structure embedding vector to form a cross-modal sample pair , construct positive sample pairs and negative sample pairs based on the historical interaction information between the user and the clothing; S34. Introduce a semantic multi-level contrast mechanism and set the number of semantic contrast levels , construct contrast tasks with multiple semantic granularities, including overall matching contrast, style attribute contrast, and color semantics contrast, and calculate independent losses for each level; S35. In each contrast layer, introduce a multi-factor-driven dynamic temperature control mechanism, and define the temperature parameter for the th training iteration as: ; Among them, is the basic temperature hyperparameter, , , are the time adjustment factor, similarity variance adjustment factor, and loss sensitivity adjustment factor respectively, represents the variance of the similarity of positive and negative sample pairs in the current batch, represents the contrast loss value of the current batch, is the logarithmic function; S36. Use a cross-modal contrast learning network model to train the semantic alignment of the user representation vector and the clothing representation vector in the shared embedding space by minimizing the weighted fusion multi-level semantic contrast loss function.
[0015] Optionally, the specific content of S4 includes: S41. Set the optimizable structural parameters and training hyperparameters in the cross-modal contrast learning network model to form an optimization target set, including the graph fusion gating coefficient , the number of semantic contrast levels , and the time adjustment factor , the similarity variance adjustment factor , and the loss sensitivity adjustment factor ; S42. Divide the optimization target set into a structural fusion subspace and a training control subspace based on the parameter functional attributes, and initialize two raccoon subpopulations , , and each raccoon individual represents a set of parameter combinations to be optimized; S43. In each subpopulation, execute the local memory-driven mechanism, environmental perturbation mechanism, and collaborative guidance update mechanism of the raccoon optimization algorithm to perform multi-round iterative optimization on the raccoon individuals in the population, and store the optimal raccoon individual in each round in the corresponding subgroup memory bank; S44. Introduce a dynamic memory window mechanism, and set the length of the current iteration memory window as: ; Among them, is the initial window length, is the feedback sensitivity adjustment coefficient, is the fitness score of the current round. The window length is used to control the number of optimal solutions retained in the subgroup memory, is the memory window length during the optimization process of the current round, is the fitness score of the previous round, is a very small positive constant; S45. For the memory banks of each sub-population after each round of optimization, according to the memory window length limit, retain the local optimal individuals and replace the outdated historical solutions; S46. Introduce an attention transfer mechanism for graph spectrum semantic perception. During every round migration period, based on the change in the edge density between user nodes and attribute nodes in the multi-modal heterogeneous semantic graph spectrum, calculate the migration attention vector between populations: ; Among them, , represent the influence weights of the sub-population on the embeddings of user nodes and clothing nodes, represents the change value of the edge weight density from users to attributes in the graph spectrum, represents the sub-population to the sub-population migration attention weight coefficient, is the graph spectrum edge structure change weight adjustment factor, is the normalization operation, and finally perform the migration operation between populations: ; Among them, is the parameter representation vector of the th raccoon individual in the sub-population , is the parameter vector of the current round fitness optimal raccoon individual in the sub-population ; S47. Select the current round optimal individual parameters , from the two sub-populations respectively, and merge them into the global optimal parameter combination ; S48. Define a composite fitness function for triple consistency evaluation as follows: ; Among them, represents the average similarity of positive sample pairs in the th layer semantic space, For the reconstruction error of multi - angle virtual try - on images of recommended clothing, For the ranking quality index between user behavior feedback and recommendation ranking. The ranking quality index is calculated by comparing the positions of the user's actual click, favorite, and purchase behaviors with the recommendation list according to the discounted cumulative gain, and is used to measure the matching degree between the recommendation ranking result and the user's true preference. , , For the balance coefficient; S49. Sort all raccoon individuals according to the calculation result of the fitness function to determine the current global optimal parameter combination. ; S410. Apply the optimal parameter combination to the cross - modal contrast learning network model, update the graph fusion mechanism, semantic contrast hierarchy structure, and temperature regulation strategy, and output the finally optimized cross - modal contrast learning network model.
[0016] Optionally, the S5 specifically includes: S51. Deploy the optimized cross - modal contrast learning network model to the recommendation task module, input the user's fusion feature vector and clothing embedding vector, and perform semantic similarity calculation; S52. According to the semantic similarity calculation result, perform matching degree scoring and ranking on the candidate clothing, and select several clothing items with the highest matching degree to form a recommendation candidate set; S53. Collect the user's body type parameter information, including the body dimension features of height, weight, shoulder width, waist circumference, and hip circumference, and perform standardized modeling on the user's body type parameter information; S54. Input the image features of each piece of clothing in the recommendation candidate set, together with the user's body type parameters and fusion features, into the virtual try - on image generation unit to perform image synthesis operations for the clothing try - on effect; S55. During the image synthesis process, set multiple viewing angles respectively to generate corresponding multi - angle realistic try - on images for each candidate clothing; S56. Organize the generated multi - angle image set into an interactive display interface, and the user can perform visual previews on each recommended clothing, including interactive operations such as image switching, rotation, zooming, and body type fitting effect comparison.
[0017] Optionally, the user's body type parameter information specifically includes the body dimension features of height, weight, shoulder width, waist circumference, and hip circumference, which are used to drive the virtual try - on image generation unit to generate realistic try - on images that match the user's body shape characteristics.
[0018] Optionally, the S6 specifically includes: S61. After the user finishes browsing the try-on images of the recommended clothing, collect the user's behavioral feedback data; S62. Preprocess and encode the collected behavioral feedback data to generate a user behavioral feedback feature vector for characterizing the current preference change of the user; S63. Based on the user behavioral feedback feature vector, update the edge weights in the multi-modal heterogeneous semantic graph, including adjusting the connection strength between the user node and the clothing node and between the user node and the attribute node, to reflect the trend of user interest reconstruction; S64. According to the updated multi-modal heterogeneous semantic graph structure, re-extract the structural embedding representations of the user and the clothing node for generating a new semantic alignment training sample composition, including updating the relationship between the positive sample pair and the negative sample pair; S65. Periodically re-execute the multi-modal semantic graph modeling, cross-modal contrastive learning training, raccoon optimization parameter update, recommendation, and try-on image generation processes to form a dynamically self-updating recommendation iteration loop; S66. After each round of closed-loop optimization cycle ends, adjust the semantic graph structure, matching strategy, and visual synthesis method in the recommendation process according to the cumulative change of the user behavioral data to improve the adaptability and feedback response ability of personalized recommendation.
[0019] Optionally, the user's behavioral feedback data specifically includes interaction information such as clicks, browsing duration, image switching, ratings, and collections, which is used to dynamically adjust the semantic graph structure and training sample composition to optimize the personalized recommendation effect.
[0020] The beneficial effects of the present invention are as follows: The personalized clothing recommendation method based on artificial intelligence proposed by the present invention, on the basis of the existing technology, constructs a complete closed-loop system from semantic understanding, interest matching to visual verification and feedback learning by introducing key technical components such as multi-modal heterogeneous semantic graphs, cross-modal contrastive learning network models, raccoon optimization algorithms, and virtual try-on image generation units. Compared with the traditional recommendation system that only relies on image retrieval or collaborative filtering, the present invention realizes a multi-modal end-to-end processing process from user data collection to try-on image output, and truly breaks through the barriers between the three stages of recommendation, try-on, and feedback.
[0021] The multi-modal heterogeneous semantic graph constructed by the present invention fully integrates user images, text descriptions, behavior records, and clothing attribute data, and performs embedded representation through a graph neural structure, effectively capturing the complex relationship between user interest preferences and clothing semantic tags. The structured embedding extracted from this graph, after being fused with the user's multi-modal features, is fed into a cross-modal contrastive learning network, and the hierarchical semantic alignment mechanism is used to improve the matching accuracy of the user-clothing pair in the semantic space. By introducing the raccoon optimization algorithm to search and optimize the model structure and hyperparameters, the generalization ability and recommendation stability of the model are significantly enhanced, especially in scenarios with large personality differences and severe cold starts, still maintaining good results.
[0022] In addition, the present invention innovatively combines the recommended candidate results with the user's body size parameters, introduces a virtual fitting image generation unit to generate interactive multi-angle fitting images, enabling users to not only "know what to recommend" but also "see the effect after wearing", greatly enhancing the user's participation and recommendation trust. The fitting images not only exist as display results but also, conversely, are linked with the user's behaviors such as clicks, ratings, and collections to form quantifiable behavioral feedback features. The system dynamically updates the edge weights of the semantic graph and the structure of the training samples through this feedback information and retrains the recommendation model periodically, thereby achieving the ability to adaptively capture user interests and provide real-time personalized recommendations.
[0023] In summary, the present invention breaks through the technical bottlenecks of traditional recommendation systems in terms of personality adaptability, visual interactivity, and self-optimization ability, achieving a comprehensive optimization effect of more accurate recommendations, more realistic fitting, more intelligent feedback, and more autonomous systems, and having high practicality and promotion value. Brief Description of the Drawings
[0024] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of a virtual fitting personalized clothing recommendation method based on artificial intelligence proposed by the present invention; Figure 2 is a flowchart of the raccoon optimization algorithm for jointly optimizing the model structure parameters and training hyperparameters of a virtual fitting personalized clothing recommendation method based on artificial intelligence proposed by the present invention. Detailed Description of the Embodiments
[0025] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0026] Refer to Figure 1 andFigure 2 , a personalized clothing recommendation method for virtual fitting based on artificial intelligence, comprising the following steps: S1. Collect the image data, text description data and historical interaction behavior data of the user to generate a multi-modal user feature set; S2. Construct a multi-modal heterogeneous semantic graph, and perform modeling and embedding representation on the structure of the multi-modal heterogeneous semantic graph; S3. Extract the structured embedding vectors of the user nodes and clothing nodes from the multi-modal heterogeneous semantic graph, and input the structured embedding vectors and the multi-modal user feature set into a cross-modal contrast learning network model, and perform semantic alignment training in the shared embedding space by constructing positive and negative sample pairs; S4. Optimize the structure parameters and training hyperparameters of the cross-modal contrast learning network model through the raccoon optimization algorithm to generate an optimized cross-modal contrast learning network model; S5. Apply the optimized cross-modal contrast learning network model to the recommendation task, calculate the semantic matching degree score between the user and the clothing, generate a recommendation candidate set, input the recommendation candidate set and the user body type parameters into the virtual fitting image generation unit together, and combine the clothing image to generate a realistic image of the user wearing the recommended clothing, and output an interactive multi-angle fitting view; S6. Collect the behavioral feedback data of the user on the fitting image to generate a user behavioral feedback feature vector, which is used to update the edge weights in the multi-modal heterogeneous semantic graph and the composition of the training samples of the cross-modal contrast learning network. Periodically execute steps S2 to S5 to form an adaptive iterative optimization closed-loop personalized recommendation process.
[0027] The virtual fitting personalized clothing recommendation method proposed by the present invention constructs a closed-loop recommendation system integrating data collection, multi-modal modeling, semantic alignment, intelligent optimization, visual generation and feedback learning through S1 to S6, and has significant beneficial effects. This method not only integrates multi-modal user features such as images, texts and behaviors, but also constructs a heterogeneous semantic graph to comprehensively express the multi-dimensional relationship between user interests and clothing semantics. Through the training of the cross-modal contrast learning network model and the parameter tuning of the raccoon optimization algorithm, the semantic matching accuracy and the model generalization ability are effectively improved. The system links the recommendation results with the user body type parameters to generate multi-angle fitting images with high realism, realizing the direct implementation from "recommendation" to "visual dressing". More importantly, the behavioral feedback of the user on the fitting image is collected in real time and used to update the semantic graph structure and the composition of the training samples, forming an adaptive optimization closed-loop of "recommendation - fitting - feedback - re-recommendation". This method significantly improves the personalization level, recommendation credibility and user experience of clothing recommendation, and has good practicability and popularization value.
[0028] In this embodiment, the multi-modal user feature set is generated by fusing visual feature vectors, semantic feature vectors, and behavioral feature vectors, where the visual feature vectors are extracted by inputting the user's image data into a convolutional neural network, the semantic feature vectors are extracted by inputting the text description data into a language understanding unit, and the behavioral feature vectors are generated by encoding the user's historical behavior data.
[0029] In this embodiment, S2 specifically includes: S21. Construct a node set, including user nodes, clothing nodes, and clothing attribute nodes. The user nodes are used to represent user objects with individual identifiers, the clothing nodes are used to represent target clothing objects that can be recommended, and the clothing attribute nodes are used to represent label information such as the style, color, season, and brand of the clothing; S22. Based on the user's historical interaction behavior data, establish an edge relationship between the user nodes and the clothing nodes. The edge relationship is used to represent the behavioral associations of click, collection, purchase, and try-on between the user and the clothing; S23. Based on the clothing metadata and label information, establish an edge relationship between the clothing nodes and the clothing attribute nodes. The edge relationship is used to represent the attribute associations of the style, color, and brand possessed by the clothing; S24. Set an initial edge weight for the constructed edge relationship. The edge weight is set according to the association strength between the user interaction frequency, behavior type, and similar labels. The edge weight is used to adjust the recommendation path and the graph neural propagation weight; S25. Use the heterogeneous graph modeling method to perform structural modeling on the multi-modal heterogeneous semantic graph, so that the heterogeneous relationship information is retained between various types of nodes, and at the same time, a complete graph structure representation is established; S26. Based on the graph neural network structure, perform embedding representation on the multi-modal heterogeneous semantic graph, and encode various types of nodes in the multi-modal heterogeneous semantic graph into structured embedding vectors through a multi-layer information aggregation mechanism. The embedding vectors are used as the input of the cross-modal contrast learning network model.
[0030] The present invention systematically constructs a multi-modal heterogeneous semantic graph, which has the ability to accurately express the multi-dimensional relationship between users and clothing, and has significant beneficial effects. By introducing user nodes, clothing nodes and clothing attribute nodes, and combining the user's historical click, favorite, purchase and try-on behaviors, a richly expressive interaction graph structure is constructed, enabling the recommendation system to not only understand user behaviors, but also identify the semantic motivations behind user preferences. At the same time, by using clothing metadata and label information to construct a clothing attribute relationship network, the interpretability and classification ability of clothing semantic information are enhanced. By setting edge weights, the system can dynamically adjust the connection weights between nodes according to the interaction frequency and behavior intensity, providing data support for the propagation path and attention mechanism in the graph neural network. The introduction of the heterogeneous graph structure ensures the independence and scalability of different types of node relationships, while the embedded representation of the graph neural network enables the node semantic features to be uniformly encoded into trainable vector representations, facilitating subsequent cross-modal contrast learning. Overall, this graph modeling method greatly improves the expression ability of the recommendation system in user interest modeling and semantic association learning, providing a high-quality structural basis for subsequent improvement of recommendation effects and personalized semantic alignment.
[0031] In this embodiment, the S3 specifically includes: S31. Extract the structured embedding vectors of the user node and the clothing node from the multi-modal heterogeneous semantic graph, which are respectively represented as and , and extract the fused representation vector from the multi-modal user feature set; S32. Introduce the graph fusion gating coefficient , construct a non-linear gating fusion mechanism, and perform nested fusion on the graph structure semantics and modal embeddings to generate the user representation vector : ; Among them, is the Sigmoid function, is the element-wise multiplication, is the multi-layer perceptron network, is the graph fusion gating coefficient; S33. Combine the user representation vector with the clothing structure embedding vector to form a cross-modal sample pair , and construct positive sample pairs and negative sample pairs according to the historical interaction information between the user and the clothing; S34. Introduce a semantic multi-layer contrast mechanism, set the number of semantic contrast levels , construct contrast tasks at multiple semantic granularities, including overall matching contrast, style attribute contrast and color semantic contrast, and calculate the independent loss for each layer; S35. In each comparison layer, introduce a multi-factor-driven dynamic temperature control mechanism, and define the temperature parameter for the th training iteration as: ; where is the base temperature hyperparameter, , , are the time adjustment factor, the similarity variance adjustment factor, and the loss sensitivity adjustment factor respectively, represents the variance of the similarity between positive and negative sample pairs in the current batch, represents the contrast loss value of the current batch, is the logarithmic function; S36. Use the cross-modal contrast learning network model to perform semantic alignment training on the user representation vector and the clothing representation vector in the shared embedding space by minimizing the weighted fusion multi-layer semantic contrast loss function.
[0032] The present invention proposes a cross-modal contrast learning method that fuses graph structure semantics and multi-modal feature information, and has the remarkable beneficial effects of improving the semantic matching accuracy and model robustness of personalized recommendations. By extracting the graph structure embeddings of users and clothing and combining multi-modal user features, the system uses a gating mechanism to flexibly control the enhancement of graph semantics, so that the final user representation has both structural information and modal perception ability. By constructing positive and negative sample pairs and introducing a semantic multi-layer contrast mechanism, the system can learn the semantic similarity between users and clothing at multiple granularity levels, significantly improving the ability to capture the implicit semantic levels in user preferences. At the same time, a multi-factor-driven dynamic temperature control mechanism is introduced, and the temperature parameter is dynamically adjusted in combination with the time progress, similarity distribution, and training loss, making the gradient of the contrast loss more stable and the model training process more adaptable and generalization ability. Finally, the network model optimizes the representation consistency of users and clothing in the shared embedding space by weighted fusion of multi-layer semantic contrast losses, realizes a recommendation system modeling ability with richer semantic levels, more accurate recommendations, and faster convergence, and provides a high-quality matching score basis for subsequent recommendation candidate generation.
[0033] In this embodiment, the specific content of S4 includes: S41. Set the optimizable structural parameters and training hyperparameters in the cross-modal contrast learning network model to form an optimization target set, including the graph fusion gating coefficient , the number of semantic contrast levels , and the time adjustment factor , the similarity variance adjustment factor , and the loss sensitivity adjustment factor ; S42. Divide the optimization target set into a structural fusion subspace based on the parameter functional attributes and the training control subspace , and initialize two raccoon subpopulations 、 . Each raccoon individual represents a set of parameter combinations to be optimized; S43. In each subpopulation, execute the local memory-driven mechanism, environmental perturbation mechanism, and collaborative guidance update mechanism of the raccoon optimization algorithm to perform multi-round iterative optimization on the raccoon individuals in the population, and store the optimal raccoon individual in each round in the corresponding subgroup memory bank; S44. Introduce a dynamic memory window mechanism. Let the current iteration memory window length be: ; where, is the initial window length, is the feedback sensitivity adjustment coefficient, is the fitness score of the current round. The window length is used to control the number of optimal solutions retained in the subgroup memory, is the memory window length in the current round of the optimization process, is the fitness score of the previous round, is a very small positive constant; S45. For the memory bank of each subpopulation after each round of optimization, according to the memory window length limit, retain the local optimal individuals and replace the outdated historical solutions; S46. Introduce an attention transfer mechanism based on graph spectrum semantic perception. In every round migration period, based on the change in the edge density between user nodes and attribute nodes in the multi-modal heterogeneous semantic graph, calculate the migration attention vector between populations: ; where, 、 represent the influence weights of the subpopulation on the user node and clothing node embeddings, represents the change value of the edge weight density from the user to the attribute in the graph, represents the subpopulation to the subpopulation the attention weight coefficient of migration, is the graph edge structure change weight adjustment factor, is the normalization operation, and finally perform the migration operation between populations: ; where, is the parameter representation vector of the th raccoon individual in the subpopulation , is the subpopulation The parameter vector of the optimal raccoon individual in terms of fitness in the current round; S47. Select the optimal individual parameters in the current round from the two sub-populations , , and combine them into the globally optimal parameter combination ; S48. Define the composite fitness function for triple consistency evaluation as follows: ; Among them, represents the average similarity of positive sample pairs in the -th layer semantic space, is the reconstruction error of the multi-angle virtual try-on image of the recommended clothing, is the ranking quality index between the user behavior feedback and the recommended ranking. The ranking quality index is calculated by comparing the positions of the user's actual click, favorite, and purchase behaviors with the recommended list according to the discounted cumulative gain, and is used to measure the matching degree between the recommended ranking result and the user's true preference, , , is the balance coefficient; S49. Sort all raccoon individuals according to the calculation result of the fitness function to determine the current globally optimal parameter combination ; S410. Apply the optimal parameter combination to the cross-modal contrast learning network model, update the graph fusion mechanism, the semantic contrast hierarchical structure, and the temperature regulation strategy, and output the finally optimized cross-modal contrast learning network model.
[0034] The present invention proposes a method for optimizing the parameters of a cross-modal contrastive learning network based on an improved raccoon optimization algorithm, which focuses on integrating a dynamic memory window mechanism and a sub-population attention transfer mechanism for graph semantic perception, achieving the joint and efficient optimization of structural parameters and training hyperparameters. First, the present invention divides the model parameters into two sub-spaces of structural integration and training control, and initializes the raccoon sub-populations respectively, enabling each type of parameter to adaptively evolve in an independent space. By introducing the dynamic memory window mechanism, the system can dynamically adjust the depth of the memory bank according to the fitness fluctuations during the optimization process, thus avoiding the risk of local convergence to the early optimal solution. Further, the sub-population attention transfer mechanism for graph semantic perception guides the weighted transfer of parameter knowledge between different sub-groups based on the real-time changes in the edge weight density between users and attribute nodes, effectively improving the semantic relevance of the optimization direction and the cross-task generalization ability. Finally, the system jointly evaluates the semantic matching accuracy, image generation quality, and user feedback ranking metrics through a triple consistency fitness function to ensure that the selected optimal parameter combination has comprehensive performance in real recommendation scenarios. Overall, this method effectively improves the model optimization efficiency, the adaptation ability of the recommendation system, and the stability of the training process, and has high practicality and innovation.
[0035] In this embodiment, step S5 specifically includes: S51. Deploy the optimized cross-modal contrastive learning network model to the recommendation task module, input the fused feature vector of the user and the clothing embedding vector, and perform semantic similarity calculation; S52. According to the semantic similarity calculation result, perform matching degree scoring and ranking on the candidate clothing, and select several pieces of clothing with the highest matching degree to form a recommendation candidate set; S53. Collect the body parameter information of the user, including the body dimension features of height, weight, shoulder width, waist circumference, and hip circumference, and perform standardized modeling on the body parameter information of the user; S54. Input the image features of each piece of clothing in the recommendation candidate set, together with the user's body parameters and fused features, into the virtual fitting image generation unit to perform the image synthesis operation of the clothing fitting effect; S55. During the image synthesis process, set multiple viewing angles respectively to generate corresponding multi-angle realistic fitting images for each candidate clothing; S56. Organize the generated multi-angle image set into an interactive display interface, and the user can perform visual previews on each recommended clothing, including interactive operations such as image switching, rotation, zooming, and comparison of body fitting effects.
[0036] The present invention constructs a personalized recommendation execution process that integrates semantic recommendation and visual fitting, which has significant practicality and user experience improvement effects. By deploying the optimized cross-modal contrast learning network model in the recommendation task module, the system can accurately calculate the semantic similarity between the user's fusion features and the clothing embedding, and complete the screening and sorting of candidate clothing based on the matching degree, ensuring that the recommendation results are highly consistent with the user's potential interests. On this basis, the user's body shape parameter information is introduced and standardized modeling is performed, so that the recommendation not only stays at the semantic level, but also realizes the accurate characterization of the user's body shape characteristics. Subsequently, the system jointly inputs the recommendation results and the user's body shape information into the virtual fitting image generation unit to generate multi-angle wearing images with realistic effects, significantly improving the intuitive visualization experience of the recommended content. Through multi-view output and interactive display interface, users can make realistic judgments and personalized decisions on the recommended clothing at the visual level, thereby enhancing user trust and participation. Overall, this method not only improves the accuracy of recommendations, but also opens up the recommendation closed loop from "interest identification" to "wear verification", with significant user friendliness and commercial application value.
[0037] In this embodiment, the user's body parameter information specifically includes body dimension characteristics of height, weight, shoulder width, waist circumference and hip circumference, which are used to drive the virtual fitting image generation unit to generate a simulated fitting image that conforms to the user's body shape characteristics.
[0038] In this implementation manner, S6 specifically includes: S61, after the user finishes browsing the recommended clothing trial images, collecting user behavior feedback data; S62, preprocessing and encoding the collected behavior feedback data to generate a user behavior feedback feature vector for describing the current preference change of the user; S63, based on the user behavior feedback feature vector, updating the edge weights in the multimodal heterogeneous semantic graph, including adjusting the connection strength between the user node and the clothing node, and between the user node and the attribute node, to reflect the user interest reconstruction trend; S64, re-extracting the structural embedding representation of the user and clothing nodes according to the updated multimodal heterogeneous semantic graph structure, for generating a new semantic alignment training sample composition, including updating the relationship between the positive sample pairs and the negative sample pairs; S65, periodically re-execute the multimodal semantic graph modeling, cross-modal contrastive learning training, raccoon optimization parameter update and recommendation, and trial image generation process to form a dynamic self-updating recommendation iterative closed loop; S66. After each closed-loop optimization cycle, the semantic graph structure, matching strategy and visual synthesis method in the recommendation process are adjusted according to the changes in the accumulated user behavior data to improve the adaptability and feedback response capabilities of personalized recommendations.
[0039] The present invention constructs a dynamic adaptive optimization mechanism driven by user behavior feedback, significantly enhancing the learning ability and long-term performance of the personalized recommendation system. After the user finishes browsing the try-on images of the recommended clothing, the method timely collects behavior data such as clicks, dwell time, ratings, and collections, and encodes them into behavior feedback feature vectors to characterize the changes in user preferences. Based on this feedback feature, the system adjusts the edge weights in the multimodal heterogeneous semantic graph to achieve a structural reconstruction of user interests, enabling the connection strength between user nodes and clothing or attribute nodes to evolve dynamically. At the same time, based on the updated graph structure, the system automatically updates the semantic alignment training samples to ensure that the training process continuously conforms to the user's latest preferences. By periodically re-executing the graph modeling, model training, parameter optimization, and recommendation processes, the system constructs a closed-loop architecture of "recommendation - feedback - optimization - re-recommendation" and has the ability of continuous learning. Finally, the system can also adjust the semantic propagation path and image synthesis strategy according to historical behavior changes to further improve the personalized adaptability and response speed. Generally speaking, this method realizes true feedback perception and system self-evolution ability, and has the beneficial effects of long-term personality fitting and continuous improvement of recommendation accuracy.
[0040] In this embodiment, the user's behavior feedback data specifically includes interaction information such as clicks, browsing duration, image switching, ratings, and collections, which is used to dynamically adjust the semantic graph structure and the composition of training samples to optimize the personalized recommendation effect.
[0041] Example 1: To verify the feasibility of the present invention in implementation, the present invention is applied to a large e-commerce platform. The platform selects 200 new users who have not had any purchase behavior in the past 30 days. Among them, 100 people use the existing traditional clothing recommendation system as the control group, and 100 people use the intelligent recommendation system deployed with the method of the present invention as the experimental group. The system completes the complete process from clothing recommendation to realistic image generation and then to feedback learning in an automated and personalized manner.
[0042] When the user first visits the platform, the experimental group users need to upload a clear front half-body photo and fill in their dressing preferences and a brief body type description. The platform automatically identifies the image features, combines the user's browsing history, behavior trajectory, and text information to generate a multimodal user feature vector. The system then models the semantic relationship between the user and the clothing based on the constructed multimodal heterogeneous semantic graph and inputs it into a cross-modal contrastive learning network model for semantic alignment training. The model optimized by the raccoon optimization algorithm is used to generate a personalized clothing recommendation candidate set. After combining the user's body type parameters such as height, weight, shoulder width, waist circumference, and hip circumference, the system generates multi-angle realistic try-on images for each recommended clothing for the user to view and try on interactively online.
[0043] Compared with traditional systems, the system of the present invention has significantly improved the recommendation quality and user experience. From the experimental monitoring data, the accuracy rate of the top-3 recommendations in the experimental group reached 85.1%, which is about 16 percentage points higher than that of the traditional system; the click-through rate of virtual fitting images reached 75.2%, much higher than 42.6% of the control group. The average browsing time of users on the fitting image page increased from 36 seconds to 59 seconds, indicating that users pay great attention to the content of the fitting images; the purchase conversion rate increased from 32.8% to 42.1%, and the body shape satisfaction score also increased from 3.9 to 4.7 points. Users generally reported that the fitting effect is real and the matching degree is high. In addition, the average number of behavioral feedbacks generated by each user in the experimental group was 22 times, which is 1.7 times that of the control group, providing sufficient data for the subsequent adaptive optimization of the system.
[0044] Table 1 Comparison of Key Indicators between the System of the Present Invention and Traditional Recommendation Systems
[0045] According to the data in Table 1, it can be clearly seen that the personalized virtual fitting recommendation system proposed by the present invention is comprehensively superior to the traditional recommendation system in multiple core indicators, demonstrating significant performance advantages and improved user experience. First of all, in terms of the accuracy rate of the top-3 recommendations, the system of the present invention reached 85.1%, while the traditional recommendation system was only 69.1%, with an increase of up to 16 percentage points. This shows that the recommendation model trained through graph modeling, cross-modal contrast learning, and raccoon optimization algorithm can more accurately identify user interests and match the most suitable clothing products, significantly improving the accuracy of recommendations.
[0046] In terms of the ratio of users clicking on virtual fitting images, the click-through rate of users in the experimental group was 75.2%, while that of the control group was only 42.6%. This difference reflects that the virtual fitting images generated by the present invention have higher attractiveness and interactive value, and can effectively guide users to explore the recommendation results in depth, which is a direct manifestation of the authenticity and personality adaptability of the image generation unit. The average browsing duration of users on the recommended fitting page also increased from 36 seconds in the traditional system to 59 seconds, indicating that the system of the present invention can attract users' attention to the recommended content for a longer time, improving user engagement and system stickiness. This "immersive recommendation" experience is particularly crucial for enhancing conversions.
[0047] The system of the present invention also has obvious advantages in terms of the purchase conversion rate, reaching 42.1%, with an increase of nearly 10 percentage points compared to 32.8% of the traditional system. This means that more accurate recommendations and visual fitting effects help users make more confident purchase decisions, greatly enhancing the potential sales volume of products from a commercial perspective.
[0048] In terms of subjective experience, the average score of users' body shape satisfaction with the try-on images reached 4.7 points (out of 5), while that of the traditional system was 3.9 points. This shows that the dressing images generated by the method of the present invention after combining the user's body shape parameters are more in line with the user's real body shape, improving the user's recognition and satisfaction.
[0049] Finally, in terms of the number of behavioral feedbacks, the users of the system of the present invention generated an average of 22 interaction behaviors per person, while the traditional system only had 13 times. This proves that the system not only enhances the user interaction activity, but also accumulates more high-quality training data for the subsequent recommendation optimization of the system, strengthening the continuous learning ability of the model.
[0050] Based on the above analysis, the system of the present invention, by introducing multi-modal feature fusion, semantic graph modeling, intelligent optimization algorithms and visual try-on experience, not only improves the accuracy and personalization of the recommendation results, but also enhances the user interaction experience and commercial conversion value, fully verifying the technical feasibility and market promotion potential of this method in real application scenarios.
[0051] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. A virtual fitting personalized clothing recommendation method based on artificial intelligence, characterized in that: The steps include: S1, collect user image data, text description data and user historical interaction behavior data to generate a multimodal user feature set; S2. Construct a multimodal heterogeneous semantic graph, and model and embed the structure of the multimodal heterogeneous semantic graph; S3, extracting structured embedding vectors of user nodes and clothing nodes from the multimodal heterogeneous semantic graph, and inputting the structured embedding vectors and multimodal user feature sets into the cross-modal contrastive learning network model, and performing semantic alignment training in the shared embedding space by constructing positive and negative sample pairs; S4. Optimizing the structural parameters and training hyperparameters of the cross-modal contrastive learning network model through the raccoon optimization algorithm to generate an optimized cross-modal contrastive learning network model; S5. Apply the optimized cross-modal contrastive learning network model to the recommendation task, calculate the semantic matching score between the user and the clothing, generate a recommendation candidate set, input the recommendation candidate set and the user's body shape parameters into the virtual fitting image generation unit, combine the clothing image to generate a realistic image of the user wearing the recommended clothing, and output an interactive multi-angle fitting view; S6. Collect user behavioral feedback data on the try-on images and generate user behavioral feedback feature vectors, which are used to update the edge weights in the multimodal heterogeneous semantic graph and the training sample composition of the cross-modal contrastive learning network. Periodically execute steps S2 to S5 to form a closed-loop personalized recommendation process with adaptive iterative optimization.
2. The method for recommending personalized clothing based on virtual fitting by artificial intelligence according to claim 1, characterized in that: The multimodal user feature set is generated by fusing a visual feature vector, a semantic feature vector and a behavioral feature vector, wherein the visual feature vector is extracted by inputting the user's image data into a convolutional neural network, the semantic feature vector is extracted by inputting text description data into a language understanding unit, and the behavioral feature vector is generated by encoding the user's historical behavior data.
3. The method for recommending personalized clothing based on virtual fitting by artificial intelligence according to claim 1, characterized in that: The S2 specifically includes: S21, constructing a node set, including a user node, a clothing node, and a clothing attribute node, wherein the user node is used to represent a user object with an individual identifier, the clothing node is used to represent a target clothing object that can be recommended, and the clothing attribute node is used to represent label information of the style, color, season, and brand of the clothing; S22, based on the user's historical interaction behavior data, establishing an edge relationship between the user node and the clothing node, wherein the edge relationship is used to indicate that there is a behavior association between the user and the clothing, such as click, favorite, purchase, and try-on; S23, based on the clothing metadata and label information, establishing an edge relationship between the clothing node and the clothing attribute node, wherein the edge relationship is used to represent the attribute association of the style, color and brand of the clothing; S24, setting an initial edge weight for the constructed edge relationship, the edge weight is set according to the user interaction frequency, behavior type and the strength of association between similar tags, and the edge weight is used to adjust the recommended path and graph neural propagation weight; S25, using a heterogeneous graph modeling method to perform structural modeling on the multimodal heterogeneous semantic graph, so that heterogeneous relationship information is retained between various types of nodes, and a complete graph structure representation is established at the same time; S26. Based on the graph neural network structure, the multimodal heterogeneous semantic graph is embedded and represented, and various nodes in the multimodal heterogeneous semantic graph are encoded into structured embedding vectors through a multi-layer information aggregation mechanism. The embedding vectors are used as input for the cross-modal contrastive learning network model.
4. The method for recommending personalized clothing based on virtual fitting by artificial intelligence according to claim 1, characterized in that: The S3 specifically includes: S31. Extract the structured embedding vectors of user nodes and clothing nodes from the multimodal heterogeneous semantic graph, which are represented as and , and extract the fusion representation vector from the multimodal user feature set ; S32, introduce the gating coefficient of the atlas fusion , construct a nonlinear gated fusion mechanism, nest and fuse the graph structure semantics with the modal embedding, and generate a user representation vector : ; in, is the Sigmoid function, is element-wise multiplication, is a multi-layer perceptron network, is the atlas fusion gating coefficient; S33, user representation vector Embedding vector with clothing structure Constructing cross-modal sample pairs , construct positive sample pairs and negative sample pairs based on the historical interaction information between users and clothing; S34. Introduce a semantic multi-layer comparison mechanism and set the number of semantic comparison levels , construct comparison tasks of multiple semantic granularities, including overall matching comparison, style attribute comparison, and color semantic comparison, and perform independent loss calculation for each layer; S35. In each contrast layer, a multi-factor driven dynamic temperature control mechanism is introduced to define the The temperature parameter for the training iteration is: ; in, is the base temperature hyperparameter, , , They are time adjustment factor, similarity variance adjustment factor and loss sensitivity adjustment factor respectively. Represents the variance of the similarity between the positive and negative sample pairs in the current batch, Represents the contrast loss value of the current batch, is a logarithmic function; S36. Use a cross-modal contrastive learning network model to minimize the weighted fusion multi-layer semantic contrast loss function to train the semantic alignment of user representation vectors and clothing representation vectors in a shared embedding space.
5. The method for recommending personalized clothing based on virtual fitting by artificial intelligence according to claim 1, characterized in that: The S4 specifically includes: S41. Set the optimizable structural parameters and training hyperparameters in the cross-modal contrastive learning network model to form an optimization target set, including the graph fusion gating coefficient , semantic contrast level , and the time adjustment factor , Similarity variance adjustment factor and loss sensitivity adjustment factor ; S42, based on parameter function attributes, divide the optimization target set into structural fusion subspaces and training control subspace , and initialize two raccoon sub-populations , , each raccoon individual represents a set of parameter combinations to be optimized; S43, in each sub-population, executing the local memory driving mechanism, environmental disturbance mechanism and collaborative guidance updating mechanism of the raccoon optimization algorithm, performing multiple rounds of iterative optimization on the raccoon individuals in the population, and storing the best raccoon individuals in each round into the corresponding sub-population memory bank; S44, introduce a dynamic memory window mechanism, and set the current iteration memory window length to: ; in, is the initial window length, is the feedback sensitivity adjustment coefficient, is the fitness score of the current round, and the window length is used to control the number of optimal solutions retained in the subgroup memory. is the memory window length in the current round of optimization, is the fitness score of the previous round, is a very small positive constant; S45, the memory library of each sub-population after each round of optimization is Limit the length of the memory window to retain the local optimal individual and replace the outdated historical solution; S46, introduce the attention transfer mechanism of graph semantic perception, During the round migration cycle, based on the edge density changes between user nodes and attribute nodes in the multimodal heterogeneous semantic graph, the inter-population migration attention vector is calculated: ; in, , represents the influence weight of the subpopulation on the embedding of user nodes and clothing nodes, Indicates the change value of the edge weight density from user to attribute in the graph, Represents subpopulation Subpopulation The attention weight coefficient of the transfer, is the weight adjustment factor for the graph edge structure change, For normalization operation, finally perform inter-population migration operation: ; in, Subpopulation Middle The parameter representation vector of each raccoon individual, Subpopulation The parameter vector of the raccoon individual with the best fitness in the current round; S47, select the optimal individual parameters of the current round from the two sub-populations , , merged into the global optimal parameter combination ; S48. Define the composite fitness function for triple consistency evaluation as follows: ; in, Indicates that the positive sample pair is The average similarity in the layer semantic space, is the reconstruction error of the multi-angle virtual try-on images of the recommended clothing. It is a ranking quality indicator between user behavior feedback and recommendation ranking. , , is the balance coefficient; S49. Sort all raccoon individuals according to the fitness function calculation results to determine the current global optimal parameter combination ; S410, combining the optimal parameters Applied to the cross-modal contrastive learning network model, it updates the graph fusion mechanism, semantic contrast hierarchy structure and temperature control strategy, and outputs the final optimized cross-modal contrastive learning network model.
6. The method for recommending personalized clothing based on virtual fitting by artificial intelligence according to claim 1, characterized in that: The S5 specifically includes: S51, deploying the optimized cross-modal contrastive learning network model to the recommendation task module, inputting the user's fused feature vector and clothing embedding vector, and performing semantic similarity calculation; S52, scoring and sorting the candidate clothing according to the semantic similarity calculation result, and selecting a number of clothing with the highest matching degree to form a recommended candidate set; S53, collecting the user's body parameter information, including body dimension characteristics of height, weight, shoulder width, waist circumference and hip circumference, and performing standardized modeling on the user's body parameter information; S54, inputting the image features of each piece of clothing in the recommended candidate set, the user's body parameters and the fusion features into a virtual fitting image generation unit, and performing an image synthesis operation of the clothing fitting effect; S55, during the image synthesis process, multiple observation perspectives are set respectively, and corresponding multi-angle simulated try-on images are generated for each candidate garment; S56. Organizing the generated multi-angle image collection into an interactive display interface, the user can perform a visual preview of each recommended clothing, including interactive operations of image switching, rotation, scaling and body fitting effect comparison.
7. The method for recommending personalized clothing based on virtual fitting based on artificial intelligence according to claim 6, characterized in that: The user's body parameter information specifically includes body dimension features of height, weight, shoulder width, waist circumference and hip circumference, which are used to drive the virtual fitting image generation unit to generate a simulated fitting image that conforms to the user's body shape features.
8. The method for recommending personalized clothing based on virtual fitting by artificial intelligence according to claim 1, characterized in that: The S6 specifically includes: S61, after the user finishes browsing the recommended clothing trial images, collecting user behavior feedback data; S62, preprocessing and encoding the collected behavior feedback data to generate a user behavior feedback feature vector for describing the current preference change of the user; S63, based on the user behavior feedback feature vector, updating the edge weights in the multimodal heterogeneous semantic graph, including adjusting the connection strength between the user node and the clothing node, and between the user node and the attribute node, to reflect the user interest reconstruction trend; S64, re-extracting the structural embedding representation of the user and clothing nodes according to the updated multimodal heterogeneous semantic graph structure, for generating a new semantic alignment training sample composition, including updating the relationship between the positive sample pairs and the negative sample pairs; S65, periodically re-execute the multimodal semantic graph modeling, cross-modal contrastive learning training, raccoon optimization parameter update and recommendation, and trial image generation process to form a dynamic self-updating recommendation iterative closed loop; S66. After each closed-loop optimization cycle, the semantic graph structure, matching strategy and visual synthesis method in the recommendation process are adjusted according to the changes in the accumulated user behavior data to improve the adaptability and feedback response capabilities of personalized recommendations.
9. The method for recommending personalized clothing based on virtual fitting by artificial intelligence according to claim 8, characterized in that: The user's behavioral feedback data specifically includes clicks, browsing time, image switching, ratings, and collection interaction information, which is used to dynamically adjust the semantic graph structure and training sample composition to optimize the personalized recommendation effect.
Citation Information
Patent Citations
Network intrusion detection method based on distributed improved raccoon algorithm
CN119341839A
Personalized virtual fitting recommendation method and system based on user body type data
CN119599762A
Cited By
Content recommendation method and system for logistics intelligent customer service interaction
CN120561257A
A content recommendation method and system for logistics intelligent customer service interaction
CN120561257B
Precise matching and distributing method and system for garment styles
CN120634691A
A method and system for accurately matching and distributing clothing styles
CN120634691B
Multi-modal clothing recommendation method
CN120744187A