Clothing matching intelligent recommendation system and method based on multi-modal learning

By adopting multimodal learning, graph neural network and cross-modal learning technologies in the clothing matching recommendation system, the problem that existing systems are difficult to utilize multimodal information and deep modeling and matching relationships is solved, and high-quality, personalized and interpretable clothing matching recommendations are achieved.

CN120087201APending Publication Date: 2025-06-03叶娉娉
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510156091.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing clothing matching recommendation system is difficult to make full use of multimodal information, cannot deeply model clothing matching relationships, lacks innovation ability and interpretability, resulting in a lack of coordination, aesthetics and personalization of recommendation results.

Method used

Using an intelligent recommendation system based on multimodal learning, the visual features of clothing are extracted through convolutional neural networks, and text features are extracted through natural language processing. The graph neural network and cross-modal learning methods are used to establish clothing matching diagrams and cross-modal diagrams for deep learning and feature fusion. At the same time, personalized recommendation modules, knowledge graph enhancement modules and style transfer modules are introduced to improve the personalization, interpretability and innovation capabilities of the system.

Benefits of technology

It has achieved comprehensive capture and deep integration of multimodal characteristics of clothing, improved the coordination, aesthetics and personalization of recommendation results, enhanced the innovation ability and interpretability of the system, and significantly improved the accuracy of user experience and recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087201A_ABST
    Figure CN120087201A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent recommendation, in particular to a clothing matching intelligent recommendation system and method based on multi-modal learning, and the system comprises a clothing feature extraction module, a matching relation modeling module, a cross-modal learning module, a cross-modal learning module, a clothing matching intelligent recommendation module and a clothing matching intelligent recommendation module, the receiving module is used for receiving node representation and edge representation sent by the matching relation modeling module; constructing a cross-modal graph, and taking the clothing visual features and the attribute text features as nodes of different modals; using a cross-modal graph neural network to learn a relationship between different modal nodes; a comparative learning mechanism is adopted to enhance the consistency of different modal features; the personalized recommendation module is in communication connection with the cross-modal learning module and is used for receiving scene constraint conditions and personal preferences input by a user; generating a personalized clothing matching recommendation list; dynamically adjusting the recommendation list according to user feedback; the complementarity of different modal information is utilized, and the method can adapt to different types of clothing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent recommendation, in particular to an intelligent clothing matching recommendation system and method based on multimodal learning. Background Art

[0002] With the rapid development of Internet technology and e-commerce, online clothing shopping has become one of the main shopping methods for consumers today. However, faced with a vast amount of clothing products, consumers often have difficulty finding suitable matching solutions. Therefore, an intelligent clothing matching recommendation system has emerged, aiming to provide users with personalized and high-quality clothing matching suggestions.

[0003] In recent years, methods based on collaborative filtering and content recommendation have been widely applied in the field of clothing recommendation. These methods mainly rely on users' historical behavior data and basic attribute information of clothing for recommendation. However, these traditional methods have some obvious limitations. First of all, they often ignore the visual features of clothing, which are crucial for clothing matching. Secondly, these methods are difficult to effectively model the complex matching relationships between clothing, resulting in the lack of coordination and beauty in the recommendation results. In addition, traditional methods usually can only process single-modal information and cannot make full use of the multimodal features of clothing, such as pictures, text descriptions, etc.

[0004] With the development of deep learning technology, some researchers have begun to try to apply convolutional neural networks to the extraction of clothing visual features and use recurrent neural networks to process the text descriptions of clothing. However, these methods still have some problems. First of all, they usually simply splice or fuse visual features and text features, and fail to fully explore the complementarity and correlation between different modalities. Secondly, most of the existing methods adopt shallow model structures and are difficult to capture high-order semantic information and implicit rules in clothing matching. Moreover, these methods often lack explicit modeling of clothing matching knowledge, resulting in the lack of interpretability and credibility of the recommendation results.

[0005] In addition, existing clothing matching recommendation systems usually only focus on recommending existing clothing combinations, lacking innovation and personalized customization capabilities. In practical applications, users may hope to fine-tune or innovate the recommended matching according to their own preferences, but existing systems are difficult to meet this demand. At the same time, most systems cannot provide users with clear and easy-to-understand recommendation reasons, which to a certain extent affects users' acceptance and trust of the recommendation results.

[0006] In view of the above problems, there is an urgent need for an intelligent clothing matching recommendation system that can make full use of multimodal information, deeply model clothing matching relationships, and have innovation capabilities and interpretability. The present invention precisely addresses these challenges and proposes an intelligent clothing matching recommendation system and method based on multimodal learning. Summary of the Invention

[0007] The intelligent clothing matching recommendation system and method based on multimodal learning provided by the present invention effectively solve many problems existing in the existing clothing matching recommendation system through innovatively integrating advanced technologies such as multimodal information processing, graph neural network, knowledge graph, and generative adversarial network, and have significant technical advantages and practical application values.

[0008] The present invention proposes an intelligent clothing matching recommendation system based on multimodal learning, including:

[0009] A clothing feature extraction module, used for:

[0010] Extracting visual features from clothing images using a convolutional neural network;

[0011] Extracting semantic features from clothing text descriptions using natural language processing techniques;

[0012] Fusing the visual features and the semantic features to generate a multimodal clothing feature representation;

[0013] A matching relationship modeling module, communicatively connected to the clothing feature extraction module, used for:

[0014] Receiving the multimodal clothing feature representation sent by the clothing feature extraction module;

[0015] Based on the multimodal clothing feature representation, constructing a clothing matching graph, where each piece of clothing is represented as a node and the matching relationship between clothes is represented as an edge;

[0016] Using a graph neural network to learn the clothing matching graph to obtain node representations and edge representations;

[0017] A cross-modal learning module, communicatively connected to the matching relationship modeling module, used for:

[0018] Receiving the node representations and edge representations sent by the matching relationship modeling module;

[0019] Constructing a cross-modal graph, using clothing visual features and attribute text features as nodes of different modalities;

[0020] Using a cross-modal graph neural network to learn the relationships between nodes of different modalities;

[0021] Adopting a contrastive learning mechanism to enhance the consistency of different modal features;

[0022] A personalized recommendation module, communicatively connected to the cross-modal learning module, used for:

[0023] Receiving the scenario constraint conditions and personal preferences input by the user;

[0024] Generate a personalized clothing matching recommendation list based on the learning results of the cross-modal learning module, combined with the scene constraint conditions and personal preferences;

[0025] Dynamically adjust the recommendation list according to user feedback;

[0026] Among them, the clothing feature extraction module, the matching relationship modeling module, the cross-modal learning module, and the personalized recommendation module constitute an end-to-end training framework, and the overall performance is improved through joint optimization.

[0027] Preferably, the clothing feature extraction module includes:

[0028] A visual feature extraction unit, used for:

[0029] Preprocess the input clothing image, including cropping, scaling, and data augmentation;

[0030] Use a pre-trained convolutional neural network to extract visual features such as the color, style, style, and texture of the clothing;

[0031] Adopt global average pooling to generate a visual feature vector with a fixed dimension;

[0032] A text feature extraction unit, used for:

[0033] Segment and clean the input clothing text description;

[0034] Use a pre-trained word embedding model to convert the text into a vector representation;

[0035] Adopt a recurrent neural network or an attention mechanism to extract the semantic features of the text;

[0036] A feature fusion unit, used for:

[0037] Receive the visual feature vector and the semantic feature vector;

[0038] Design a multi-layer perceptron network to map visual features and semantic features to a common feature space;

[0039] Adopt an attention mechanism to perform weighted fusion on different modality features to generate a final multi-modal clothing feature representation.

[0040] Preferably, the matching relationship modeling module includes:

[0041] A graph construction unit, used for:

[0042] Construct an initial clothing matching graph based on the matching frequency between clothing or expert annotations;

[0043] Use the multi-modal features of each piece of clothing as the initial features of the nodes;

[0044] The weight calculation method of the design edge reflects the coordination degree of clothing matching;

[0045] Graph Convolutional Network Unit, used for:

[0046] Design a multi-layer graph convolutional network structure to learn the representation of nodes and edges;

[0047] In each layer of graph convolution, neighbor node information is aggregated and the central node representation is updated;

[0048] Use skip connections and residual connections to improve the network's expressiveness;

[0049] Attention mechanism unit, used to:

[0050] Design a graph attention layer to learn the importance weights between nodes;

[0051] Implement a multi-head attention mechanism to capture the relationship between nodes from different angles;

[0052] Combining self-attention and neighbor attention can enhance the expressiveness of the model.

[0053] Preferably, the cross-modal learning module comprises:

[0054] Cross-modal graph building blocks for:

[0055] Integrate the visual feature nodes and attribute text feature nodes of clothing into the same graph structure; design heterogeneous edges to connect nodes of different modes;

[0056] Define the intra-modality and inter-modality adjacency matrices;

[0057] Cross-modal graph neural network units for:

[0058] Design a graph neural network structure that can handle heterogeneous nodes and edges;

[0059] Implement modality-specific information transfer and aggregation mechanisms;

[0060] Design a cross-modal information fusion layer to integrate features from different modalities;

[0061] Comparative Learning Unit for:

[0062] Construct positive sample pairs and negative sample pairs, including intra-modal and inter-modal sample pairs;

[0063] Design a contrast loss function to bring positive sample pairs closer together and push negative sample pairs further apart;

[0064] Implement momentum contrastive learning to improve the expressiveness and generalization of the model.

[0065] Preferably, the personalized recommendation module includes:

[0066] A user profile unit for:

[0067] Collecting and analyzing the historical behavior data of users, including browsing, clicking, and purchase records;

[0068] Extracting the static features of users, such as age and gender, and dynamic features, such as fashion preferences;

[0069] Constructing a multi-dimensional user feature vector;

[0070] A scene constraint processing unit for:

[0071] Parsing the scene constraint conditions input by users, such as season, occasion, and style;

[0072] Converting the scene constraints into a quantifiable feature representation;

[0073] Designing the soft and hard rules of the scene constraints;

[0074] A matching calculation unit for:

[0075] Designing a multi-objective optimization function to balance the coordination, personalization, and diversity of clothing matching;

[0076] Implementing a graph-based inference algorithm to perform path search on the clothing matching graph;

[0077] Adopting a reinforcement learning method to optimize the long-term recommendation effect;

[0078] A dynamic adjustment unit for:

[0079] Collecting the real-time feedback of users on the recommendation results;

[0080] Designing an online learning algorithm to update the model parameters in a timely manner;

[0081] Implementing a balance strategy between exploration and exploitation to improve the novelty of the recommendation.

[0082] Preferably, it further includes:

[0083] A multi-task learning module, communicatively connected to the clothing feature extraction module, the matching relationship modeling module, and the cross-modal learning module, for:

[0084] Designing a shared feature extraction network as the basis for multiple tasks;

[0085] Defining multiple related tasks, including clothing attribute prediction, matching relationship classification, and personalized recommendation;

[0086] Designing a task-specific output layer to adapt to the requirements of different tasks;

[0087] Implement a soft parameter sharing mechanism to balance knowledge transfer between tasks;

[0088] Design a multi-task loss function to comprehensively optimize the performance of all tasks.

[0089] Preferably, it further includes:

[0090] A knowledge graph enhancement module, communicatively connected to the collocation relationship modeling module and the cross-modal learning module, for:

[0091] Construct a knowledge graph in the clothing field, including clothing entities, attributes, and relationships;

[0092] Design a knowledge graph embedding method to map entities and relationships into a low-dimensional vector space;

[0093] Implement a reasoning mechanism based on the knowledge graph to supplement prior knowledge of clothing collocation;

[0094] Design a knowledge distillation method to transfer the information of the knowledge graph into a neural network model.

[0095] Preferably, it further includes:

[0096] A style transfer module, communicatively connected to the clothing feature extraction module and the personalized recommendation module, for:

[0097] Implement clothing style transfer based on a generative adversarial network;

[0098] Design a style encoder to extract the style features of clothing;

[0099] Implement an adaptive style mixing method to generate new clothing styles;

[0100] Design a style consistency loss to ensure the quality and diversity of the generated results.

[0101] Preferably, it further includes:

[0102] An interpretability enhancement module, communicatively connected to the collocation relationship modeling module and the personalized recommendation module, for:

[0103] Design an attention visualization method to display the clothing regions and features that the model focuses on;

[0104] Implement a graph-based interpretation method to trace the path of the recommendation decision;

[0105] Generate natural language explanations to illustrate the recommendation reasons and collocation principles;

[0106] Design an interactive interpretation interface to allow users to adjust the recommendation results.

[0107] Intelligent clothing matching recommendation method based on multimodal learning, including the following steps:

[0108] S1. Multimodal clothing feature extraction:

[0109] Extract visual features from clothing images using a convolutional neural network;

[0110] Extract semantic features from clothing text descriptions using natural language processing techniques;

[0111] Fuse the visual features and the semantic features to generate a multimodal clothing feature representation;

[0112] S2. Clothing matching relationship modeling:

[0113] Construct a clothing matching graph based on the multimodal clothing feature representation;

[0114] Use a graph neural network to learn the clothing matching graph to obtain node representations and edge representations;

[0115] S3. Cross-modal learning:

[0116] Construct a cross-modal graph, taking clothing visual features and attribute text features as nodes of different modalities;

[0117] Use a cross-modal graph neural network to learn the relationships between nodes of different modalities;

[0118] Adopt a contrastive learning mechanism to enhance the consistency of different modal features;

[0119] S4. Personalized recommendation generation:

[0120] Receive the scenario constraint conditions and personal preferences input by the user;

[0121] Based on the cross-modal learning results, combine the scenario constraint conditions and personal preferences to generate a personalized clothing matching recommendation list;

[0122] Dynamically adjust the recommendation list according to user feedback;

[0123] S5. Multi-task joint optimization:

[0124] Design a multi-task learning framework to simultaneously optimize clothing attribute prediction, matching relationship classification, and personalized recommendation tasks;

[0125] Adopt a soft parameter sharing mechanism to balance the knowledge transfer between tasks;

[0126] S6. Knowledge graph enhancement:

[0127] Construct a knowledge graph in the clothing field to supplement the prior knowledge of clothing matching;

[0128] Design a knowledge distillation method to transfer the information of the knowledge graph into a neural network model;

[0129] S7. Style transfer and innovation:

[0130] Implement clothing style transfer based on the generative adversarial network;

[0131] Design an adaptive style mixing method to generate new clothing styles;

[0132] S8. Interpretability enhancement:

[0133] Design an attention visualization method to display the decision-making basis of the model;

[0134] Generate natural language explanations to illustrate the recommendation reasons and matching principles;

[0135] Among them, steps S1 to S8 constitute an end-to-end training and inference process, and the overall performance is improved through iterative optimization.

[0136] The beneficial effects of the present invention are mainly reflected in the following aspects:

[0137] First of all, the system of the present invention can comprehensively capture the multi-modal features of clothing. By using the deep learning model to extract the visual features and text features of clothing respectively, and adopting the attention mechanism for effective fusion, the system can represent the attributes and styles of clothing more comprehensively and accurately. This multi-modal feature extraction method not only makes full use of the complementarity of different modal information, but also can adapt to different types of clothing data, significantly improving the richness and robustness of feature representation.

[0138] Secondly, the present invention innovatively introduces a clothing matching relationship modeling method based on the graph neural network. By transforming the clothing matching problem into a graph structure learning problem, the system can effectively capture the complex matching relationships and high-order semantic information between clothing. This method can not only learn explicit matching rules, but also discover potential matching patterns, greatly improving the coordination and aesthetic feeling of the recommendation results.

[0139] Furthermore, the cross-modal learning module of the present invention realizes the deep fusion and complementarity of different modal information. By constructing a cross-modal graph and designing a special message passing mechanism, the system can fully explore the correlation between visual features and text features, so as to generate a more comprehensive and accurate clothing representation. This cross-modal learning method significantly improves the generalization ability of the system and can better handle novel clothing styles and descriptions.

[0140] In addition, the personalized recommendation module of the present invention adopts multi-objective optimization and reinforcement learning methods, which can simultaneously consider the coordination, personalization, and diversity of collocations. By dynamically adjusting the recommendation strategy, the system can continuously adapt to the changes in user preferences and provide more accurate and personalized recommendation services. This method effectively solves the cold start and long-tail problems in traditional recommendation systems and greatly improves the user experience.

[0141] The present invention also innovatively introduces a knowledge graph enhancement and style transfer module, which further improves the performance and functions of the system. The introduction of the knowledge graph injects rich prior knowledge into the recommendation system, improving the reliability and diversity of recommendation results. The style transfer module endows the system with the ability to innovate clothing styles and can generate novel collocation schemes according to user needs, meeting the personalized customization requirements of users.

[0142] Finally, the interpretability enhancement module of the present invention greatly improves the transparency and credibility of recommendation results. Through attention visualization, graph-based interpretation methods, and natural language generation technologies, the system can provide users with intuitive and easy-to-understand reasons for recommendations, helping users understand the recommendation decision-making process. This not only enhances users' trust in the system but also provides valuable insights for fashion designers and retailers.

[0143] Generally speaking, the intelligent clothing collocation recommendation system and method based on multi-modal learning provided by the present invention realize the intelligence, personalization, and interpretability of clothing collocation recommendations through the organic combination of multiple innovative technologies. This system can not only significantly improve the accuracy of recommendations and user satisfaction but also stimulate users' creativity, providing strong technical support for the digital transformation of the fashion industry. In multiple fields such as e-commerce, personal styling consultants, and virtual fitting rooms, the present invention has broad application prospects and great commercial value. Brief Description of the Drawings

[0144] Figure 1 It is the logical block diagram of the overall system of the present invention.

[0145] Figure 2 It is the logical block diagram of the clothing feature extraction module of the present invention.

[0146] Figure 3 It is the logical block diagram of the collocation relationship modeling module of the present invention.

[0147] Figure 4 It is the logical block diagram of the cross-modal learning module of the present invention.

[0148] Figure 5 It is the logical block diagram of the personalized recommendation module of the present invention.

[0149] Figure 6This is the logic block diagram of the knowledge graph enhancement module of the present invention. Detailed implementation manners

[0150] Refer to Figure 1-6 , the present invention provides an intelligent clothing matching recommendation system and method based on multimodal learning. This system can effectively fuse the visual features and text descriptions of clothing, establish the matching relationships between clothing, and generate personalized matching recommendations for users. The following will introduce the detailed implementation manners of the present invention.

[0151] The system of the present invention includes a clothing feature extraction module 1, a matching relationship modeling module 2, a cross-modal learning module 3, and a personalized recommendation module 4. These modules work together to form an end-to-end training framework, and the overall performance is improved through joint optimization.

[0152] Preferably, the clothing feature extraction module 1 is used to extract the feature representations of clothing from multiple modalities. Specifically, this module first extracts visual features from clothing images using a convolutional neural network. In an embodiment of the present invention, a pre-trained ResNet50 network can be used as the feature extractor to extract a 4096-dimensional visual feature vector. Next, this module uses natural language processing techniques to extract semantic features from clothing text descriptions. For example, the BERT model can be used to encode the text to obtain a 768-dimensional semantic feature vector. Finally, this module fuses the visual features and semantic features to generate a multimodal clothing feature representation. The fusion can be achieved through simple feature concatenation or more complex attention mechanisms for weighted fusion.

[0153] The matching relationship modeling module 2 is communicatively connected to the clothing feature extraction module 1 and is used to establish the matching relationships between clothing. This module first receives the multimodal clothing feature representations sent by the clothing feature extraction module 1. Based on these feature representations, the module constructs a clothing matching graph, where each piece of clothing is represented as a node, and the matching relationships between clothing are represented as edges. In a preferred implementation manner of the present invention, the weights of the edges can be determined according to the frequency of clothing matching or expert scores. For example, the edge weight between clothing pairs with a matching frequency exceeding 100 times can be set to 1, between 50 - 100 times to 0.5, and below 50 times to 0.1. After constructing the graph, this module uses a graph neural network to learn the clothing matching graph to obtain node representations and edge representations. A graph convolutional network (GCN) or a graph attention network (GAT) can be used as the basic network structure.

[0154] The cross-modal learning module 3 is communicatively connected to the collocation relationship modeling module 2 and is used to further enhance the fusion of different-modal information. This module first receives the node representation and edge representation sent by the collocation relationship modeling module 2. Then, a cross-modal graph is constructed, with the clothing visual features and attribute text features as nodes of different modalities. In an embodiment of the present invention, two nodes can be created for each piece of clothing, corresponding to the visual features and text features respectively, and an edge is added between these two nodes to indicate that they belong to the same piece of clothing. Next, this module uses the cross-modal graph neural network to learn the relationships between different-modal nodes. A special message passing mechanism can be designed to allow information of different modalities to flow in the graph. Finally, this module adopts a contrastive learning mechanism to enhance the consistency of different-modal features. Specifically, the different-modal features of the same piece of clothing can be used as positive sample pairs, and the features of different pieces of clothing can be used as negative sample pairs, and the model is optimized by minimizing the distance between positive sample pairs and maximizing the distance between negative sample pairs.

[0155] The personalized recommendation module 4 is communicatively connected to the cross-modal learning module 3 and is used to generate personalized clothing collocation recommendations for users. This module first receives the scenario constraint conditions and personal preferences input by the user. The scenario constraints may include factors such as season, occasion, style, etc., while the personal preferences may involve aspects such as color, brand, price, etc. Based on the learning results of the cross-modal learning module 3, combined with these scenario constraint conditions and personal preferences, this module generates a list of personalized clothing collocation recommendations. During the recommendation process, a graph-based reasoning algorithm can be adopted to perform path search on the clothing collocation graph to find the collocation combination that best meets the user's needs. In addition, this module can also dynamically adjust the recommendation list according to user feedback. For example, if the user shows a high click-through rate for a certain style of collocation, the system will increase the weight of this style of collocation in subsequent recommendations.

[0156] The clothing feature extraction module 1 further includes a visual feature extraction unit 11, a text feature extraction unit 12, and a feature fusion unit 13.

[0157] The visual feature extraction unit 11 is responsible for processing the image information of the clothing. First, this unit preprocesses the input clothing image, including cropping, scaling, and data augmentation. For example, the image can be uniformly scaled to a size of 224x224, and operations such as random horizontal flipping, brightness, and contrast adjustment are performed to increase data diversity. Next, this unit uses a pre-trained convolutional neural network to extract visual features such as the color, style, style, and texture of the clothing. In a preferred embodiment of the present invention, the ResNet50 network pre-trained on the ImageNet dataset can be used as the feature extractor. Finally, this unit adopts global average pooling to generate a visual feature vector with a fixed dimension. Generally, the dimension of this vector can be set to 2048 or 4096 to balance the feature expression ability and computational efficiency.

[0158] The text feature extraction unit 12 processes the text description information of the clothing. First, this unit tokenizes and cleans the input clothing text description, removing stop words and irrelevant characters. Then, it uses a pre-trained word embedding model to convert the text into a vector representation. In the present invention, commonly used word embedding models such as Word2Vec or GloVe can be adopted. Finally, this unit uses a recurrent neural network or an attention mechanism to extract the semantic features of the text. For example, a bidirectional LSTM network can be used to process the sequence of word vectors, and the hidden state at the last time step is taken as the semantic representation of the text. Alternatively, a self-attention mechanism can be used to aggregate all word vectors to obtain a global text representation.

[0159] The feature fusion unit 13 is responsible for fusing the visual features and text features to generate a multi-modal clothing feature representation. This unit first receives the feature vectors output by the visual feature extraction unit 11 and the text feature extraction unit 12. Then, it designs a multi-layer perceptron network to map the visual features and semantic features into a common feature space. In an embodiment of the present invention, two independent fully connected layers can be used to process the visual features and text features respectively, and then they are concatenated together. Finally, this unit uses an attention mechanism to perform weighted fusion on different modal features to generate the final multi-modal clothing feature representation. Specifically, the importance weights of the visual features and text features can be calculated, and then weighted summation is performed. The calculation of the attention weights can be achieved through the following formula:

[0160]

[0161] where, f i represents the feature vector of the i-th modality, w is a learnable parameter vector, and α i is the corresponding attention weight. The final multi-modal feature representation can be expressed as:

[0162] f multi = α 1 f visual + α 2 f text ,

[0163] This fusion method can adaptively adjust the importance of different modal information and improve the quality of the feature representation.

[0164] The matching relationship modeling module 2 further includes a graph construction unit 21, a graph convolutional network unit 22, and an attention mechanism unit 23.

[0165] The graph construction unit 21 is responsible for constructing the initial clothing matching graph. First, this unit constructs the initial clothing matching graph based on the matching frequency between clothes or expert annotations. In an embodiment of the present invention, a large-scale clothing matching dataset can be utilized to count the matching frequencies of different clothes. If the matching frequency of two pieces of clothing exceeds a preset threshold (e.g., 100 times), an edge is added between them. Next, this unit takes the multi-modal features of each piece of clothing as the initial features of the nodes. These features are from the output of the clothing feature extraction module 1. Finally, this unit designs a method for calculating the weights of the edges to reflect the coordination degree of clothing matching. For example, the edge weights can be normalized according to the matching frequency or expert scores so that they fall within the interval [0, 1].

[0166] The graph convolutional network unit 22 is used to learn the representations of nodes and edges. This unit designs a multi-layer graph convolutional network structure for learning the representations of nodes and edges. In the present invention, the following graph convolutional operation can be adopted:

[0167]

[0168] where H (l) represents the node feature matrix of the l-th layer, A is the adjacency matrix, D is the degree matrix, W (l) is the learnable weight matrix, and σ is the non-linear activation function. In each layer of graph convolution, this unit aggregates the information of neighboring nodes and updates the representation of the central node. To enhance the expressive power of the network, skip connections and residual connections can be adopted. For example, the outputs of different layers can be concatenated or added to fuse multi-scale information.

[0169] The attention mechanism unit 23 further enhances the expressive power of the model. This unit designs graph attention layers to learn the importance weights between nodes. In a preferred embodiment of the present invention, a multi-head attention mechanism can be adopted to capture the relationships between nodes from different perspectives. Specifically, for node i and its neighboring node j, the attention coefficient can be calculated in the following way:

[0170]

[0171] where h i and h j are the features of nodes i and j respectively, W is the linear transformation matrix, a is the attention vector, and || represents the concatenation operation. Finally, this unit combines self-attention and neighbor attention to enhance the expressive power of the model. Self-attention can capture the importance of the features of the node itself, while neighbor attention focuses on the relationships between nodes.

[0172] Through the collaborative work of the above modules and units, the system of the present invention can effectively extract multi-modal features of clothing, establish the matching relationships between clothing, and generate personalized recommendation results. The system has the following advantages: First, the multi-modal feature extraction can comprehensively capture the visual and semantic information of clothing, improving the richness and accuracy of feature representation. Second, the matching relationship modeling based on the graph neural network can effectively capture the complex relationships between clothing and learn the implicit matching rules. Third, the cross-modal learning mechanism further enhances the fusion of different modal information and improves the generalization ability of the model. Finally, the personalized recommendation module can dynamically adjust the recommendation results according to the specific needs and feedback of users, providing more accurate and personalized services.

[0173] The system of the present invention has broad prospects in practical applications. For example, on e-commerce platforms, the system can recommend suitable clothing combinations for users, improving the shopping experience of users and the sales conversion rate of the platform. In fashion consulting applications, the system can provide personalized dressing suggestions for users to meet the needs of different scenarios and styles. In addition, the system can also be applied to fields such as virtual fitting and fashion design assistance, providing technical support for the digital transformation of the fashion industry.

[0174] The cross-modal learning module 3 of the present invention further includes a cross-modal graph construction unit 31, a cross-modal graph neural network unit 32, and a contrast learning unit 33. These units work together to achieve the deep fusion and learning of different modal information.

[0175] The cross-modal graph construction unit 31 is responsible for constructing the cross-modal graph structure. This unit first integrates the visual feature nodes and attribute text feature nodes of clothing into the same graph structure. In an embodiment of the present invention, each piece of clothing is represented by two nodes, corresponding to its visual feature and text feature respectively. Next, this unit designs heterogeneous edges to connect nodes of different modalities. For example, an edge can be added between the visual feature node and the text feature node of the same piece of clothing, indicating that they describe the same entity. In addition, edges can also be added between nodes of different clothing according to the similarity or matching relationship of clothing. Finally, this unit defines the intra-modal and inter-modal adjacency matrices, providing a basis for subsequent graph neural network learning.

[0176] The cross-modal graph neural network unit 32 is used to learn the node representations in the cross-modal graph. This unit designs a graph neural network structure that can handle heterogeneous nodes and edges. In a preferred embodiment of the present invention, a graph neural network based on meta-paths can be adopted. Specifically, multiple meta-paths can be defined, such as "visual - text - visual" or "text - visual - text", to guide the transfer of information between different modalities. This unit implements a modality-specific information transfer and aggregation mechanism. For example, for visual feature nodes, the following update formula can be designed:

[0177]

[0178] in, represents the representation of node v at layer l, is a set of predefined meta paths, is the set of neighbor nodes reachable through meta-path p, α p,vu is the attention weight, is a learnable weight matrix associated with the meta-path p. Finally, the unit designs a cross-modal information fusion layer to integrate features from different modalities. This fusion process can be achieved using a self-attention mechanism or a gating mechanism.

[0179] The contrastive learning unit 33 is used to enhance the consistency of features of different modalities. The unit first constructs positive sample pairs and negative sample pairs, including intra-modal and inter-modal sample pairs. In one embodiment of the present invention, the visual features and text features of the same garment can be used as positive sample pairs, and the features of different garments can be used as negative sample pairs. Next, the unit designs a contrastive loss function to shorten the distance between positive sample pairs and extend the distance between negative sample pairs. The InfoNCE loss function can be used, which is defined as follows:

[0180]

[0181] Among them, h i and is the feature representation of the positive sample pair, h k is the feature representation of negative samples, sim(·,·) is a similarity function (such as cosine similarity), and τ is a temperature parameter. Finally, the unit implements momentum contrastive learning to improve the expressiveness and generalization of the model. Specifically, a momentum encoder and a feature queue can be maintained to increase the diversity and consistency of negative samples.

[0182] The personalized recommendation module 4 of the present invention further includes a user portrait unit 41, a scene constraint processing unit 42, a matching calculation unit 43 and a dynamic adjustment unit 44. These units work together to generate personalized clothing matching recommendations for users.

[0183] The user profile unit 41 is responsible for constructing a multi-dimensional feature representation of the user. This unit first collects and analyzes the user's historical behavior data, including browsing, clicking, purchase records, etc. In an embodiment of the present invention, a collaborative filtering algorithm can be used to analyze the user's behavior pattern. For example, a matrix factorization-based method can be used to factorize the user-item interaction matrix into low-dimensional latent factors, so as to capture the user's implicit preferences. Next, this unit extracts the user's static features (such as age, gender) and dynamic features (such as fashion preferences). Static features can be obtained from the user registration information, while dynamic features can be obtained by analyzing the user's recent behavior. Finally, this unit constructs a multi-dimensional user feature vector. An embedding layer can be used to convert discrete features into dense vectors and concatenate them with continuous features to form the final user representation.

[0184] The scenario constraint processing unit 42 is used to process the scenario constraint conditions input by the user. This unit first parses the scenario constraint conditions input by the user, such as season, occasion, style, etc. In a preferred embodiment of the present invention, a predefined set of scenario labels, such as "spring", "business", "casual", etc., can be designed, and the user input is mapped to these labels. Next, this unit converts the scenario constraints into a quantifiable feature representation. For example, one-hot encoding or an embedding layer can be used to convert the scenario labels into vectors. Finally, this unit designs the soft and hard rules for scenario constraints. Hard rules can directly filter out clothing that does not meet the scenario requirements, while soft rules can affect the recommendation results by adjusting the similarity scores.

[0185] The matching calculation unit 43 is responsible for generating clothing matching recommendations. This unit first designs a multi-objective optimization function to balance the coordination, personalization, and diversity of clothing matches. The following objective function can be defined:

[0186]

[0187] where S is the set of recommended clothing matches, f coord (S) measures the coordination of the match, f pers (S, u) measures the personalization matching degree with the user u, f div (S) measures the diversity of the recommendation results, and λ 1 , λ 2 and λ 3 are trade-off coefficients. Next, this unit implements a graph-based inference algorithm to perform path search on the clothing matching graph. Random walk or attention-based graph traversal methods can be used to explore potential matching combinations starting from the user's existing clothing. Finally, this unit uses a reinforcement learning method to optimize the long-term recommendation effect. The recommendation process can be modeled as a Markov decision process, and a policy gradient algorithm is used to learn the optimal recommendation strategy.

[0188] The dynamic adjustment unit 44 is used to optimize the recommendation results according to user feedback. This unit first collects the real-time feedback of users on the recommendation results, such as behaviors like clicks, favorites, purchases, etc. In an embodiment of the present invention, a multi-level feedback mechanism can be designed to convert user behaviors into positive or negative signals of different intensities. Next, this unit designs an online learning algorithm to update the model parameters in a timely manner. An incremental learning method, such as online gradient descent, can be adopted to quickly adapt to the latest preferences of users. Finally, this unit implements a balance strategy between exploration and exploitation to improve the novelty of recommendations. For example, the ε-greedy strategy can be used to randomly recommend some new combinations with a certain probability to explore the potential interests of users.

[0189] The system of the present invention further includes a multi-task learning module 5. This module is communicatively connected to the clothing feature extraction module 1, the matching relationship modeling module 2, and the cross-modal learning module 3, and is used to improve the overall performance and generalization ability of the model.

[0190] The multi-task learning module 5 first designs a shared feature extraction network as the basis for multiple tasks. In a preferred embodiment of the present invention, the outputs of the previous several modules can be used as shared features, such as the multi-modal representation and graph structure representation of clothing. Next, this module defines multiple related tasks, including clothing attribute prediction, matching relationship classification, and personalized recommendation. For example, the clothing attribute prediction task can predict attributes such as the color, style, and material of clothing; the matching relationship classification task can determine whether two pieces of clothing are suitable for matching; the personalized recommendation task is to generate matching suggestions for specific users.

[0191] This module also designs task-specific output layers to adapt to the requirements of different tasks. For classification tasks, a softmax layer can be used; for regression tasks, a fully connected layer can be used. In addition, this module implements a soft parameter sharing mechanism to balance the knowledge transfer between tasks. Specifically, independent task-specific layers can be designed for each task, but regularization constraints are added between these layers to encourage parameter similarity. Finally, this module designs a multi-task loss function to comprehensively optimize the performance of all tasks. The losses of different tasks can be combined in a weighted sum manner:

[0192]

[0193] where is the loss function of the i-th task, w i is the corresponding weight coefficient, and M is the total number of tasks. The weights can be automatically learned through grid search or gradient-based methods.

[0194] The system of the present invention further includes a knowledge graph enhancement module 6. This module is communicatively connected to the collocation relationship modeling module 2 and the cross-modal learning module 3, and is used to introduce domain knowledge to improve the interpretability and generalization ability of the model.

[0195] The knowledge graph enhancement module 6 first constructs a knowledge graph in the clothing domain, including clothing entities, attributes, and relationships. In an embodiment of the present invention, knowledge can be extracted from fashion magazines, expert knowledge, and large-scale clothing datasets to construct a knowledge graph containing relationships such as "is a certain type", "suitable for a certain occasion", and "harmonious collocation". Next, this module designs a knowledge graph embedding method to map entities and relationships into a low-dimensional vector space. Classical knowledge graph embedding methods such as TransE and RotatE can be used to learn the vector representations of entities and relationships.

[0196] This module also implements a reasoning mechanism based on the knowledge graph to supplement the prior knowledge of clothing collocation. For example, path patterns in the knowledge graph can be used to infer new collocation relationships, or the attribute information of entities can be used to enhance the feature representation of clothing. Finally, this module designs a knowledge distillation method to transfer the information of the knowledge graph into the neural network model. Specifically, the knowledge graph embedding can be used as an auxiliary task and jointly optimized with the main recommendation task, or the results of knowledge graph reasoning can be used as soft labels to guide model learning.

[0197] By introducing multi-task learning and knowledge graph enhancement, the system of the present invention can better utilize limited data, learn more robust and general feature representations, thereby improving the accuracy and interpretability of recommendations. The addition of these modules enables the system to not only handle the main task of clothing collocation recommendation but also solve multiple related sub-tasks, such as attribute prediction and relationship classification, thereby realizing knowledge sharing and transfer. At the same time, the introduction of the knowledge graph provides the model with rich prior knowledge, helps to handle problems such as data sparsity and cold start, and provides an interpretable basis for recommendation results.

[0198] The system of the present invention further includes a style transfer module 7. This module is communicatively connected to the clothing feature extraction module 1 and the personalized recommendation module 4, and is used to achieve innovation and personalized customization of clothing styles.

[0199] The style transfer module 7 first implements clothing style transfer based on a generative adversarial network (GAN). In a preferred embodiment of the present invention, the architecture of a conditional generative adversarial network (cGAN) can be adopted, where the generator G is responsible for converting the input clothing image into the target style, and the discriminator D is responsible for distinguishing whether the generated image conforms to the target style. The optimization objectives of the generator and the discriminator can be expressed as:

[0200]

[0201] Among them, x is the input clothing image, y is the target style label, and z is the random noise. To preserve the structural information of the clothing, skip connections can be introduced in the generator to construct a network structure similar to U-Net. Next, this module designs a style encoder to extract the style features of the clothing. The output of the intermediate layer of a pre-trained convolutional neural network (such as VGG16) can be used as the style representation. Specifically, the Gram matrix of the feature map can be calculated as the style feature:

[0202]

[0203] Among them, F l is the feature map of the l-th layer, and C l , H l and W l are the number of channels, height, and width of this layer respectively.

[0204] The style transfer module 7 also implements an adaptive style mixing method to generate a new clothing style. In an embodiment of the present invention, the AdaIN (Adaptive Instance Normalization) technology can be used to achieve flexible style regulation. Specifically, the mean and variance of the content features can be adjusted through the following formula to match the statistical characteristics of the style features:

[0205]

[0206] Among them, x is the content feature, y is the style feature, and μ and σ represent the mean and standard deviation respectively.

[0207] Finally, this module designs a style consistency loss to ensure the quality and diversity of the generated results. The content loss and style loss can be considered simultaneously. The content loss ensures that the generated image retains the structure of the original clothing, and the style loss ensures that the generated image conforms to the target style. The total loss function can be expressed as:

[0208]

[0209] Among them, is the content loss, is the style loss, is the adversarial loss, and λ 1 , λ 2 and λ 3 are the trade-off coefficients.

[0210] The system of the present invention further includes an interpretability enhancement module 8. This module is communicatively connected to the collocation relationship modeling module 2 and the personalized recommendation module 4, and is used to improve the interpretability of the recommendation results and user trust.

[0211] The interpretability enhancement module 8 first designs an attention visualization method to display the clothing regions and features that the model focuses on. In one embodiment of the present invention, the Grad-CAM (Gradient-weighted Class Activation Mapping) technique can be used to generate a heatmap, highlighting the image regions that the model pays the most attention to when making decisions. Specifically, the gradient of the target class score with respect to the feature map of the last convolutional layer can be calculated, and then the feature map is weighted and summed:

[0212]

[0213] where is the importance weight of the k-th channel, and A k is the feature map of the k-th channel. Next, the module implements a graph-based interpretation method to trace the path of the recommendation decision. The attention mechanism can be used to analyze the important nodes and edges in the clothing matching graph, so as to explain why certain clothing items are recommended together. For example, the path attention score can be calculated:

[0214] α path = softmax(W 2 tanh(W 1 [h s ||h t ||e path )),

[0215] where h s and h t are the representations of the starting node and the target node respectively, e path is the embedding representation of the path, and W 1 and W 2 are learnable parameter matrices.

[0216]

[0217] where x i is the input feature and template, y i is the target explanation text, and θ is the model parameter.

[0218] Finally, the module designs an interactive interpretation interface that allows users to adjust the recommendation results. A visual control panel can be provided to enable users to adjust the importance of different features, such as color coordination, style matching, occasion suitability, etc. The system will update the recommendation results and the corresponding explanations in real time, thereby improving users' understanding and control of the recommendation process.

[0219] The present invention also provides a clothing matching intelligent recommendation method based on multi-modal learning. This method includes multiple steps, and the specific implementation of each step will be described in detail below.

[0220] Step S1 is multi-modal clothing feature extraction. In this step, first, a convolutional neural network is used to extract visual features from clothing images. A pre-trained ResNet50 network can be used as the feature extractor to extract a 4096-dimensional visual feature vector. Then, natural language processing techniques are used to extract semantic features from clothing text descriptions. The BERT model can be used to encode the text to obtain a 768-dimensional semantic feature vector. Finally, the visual features and semantic features are fused to generate a multi-modal clothing feature representation. The fusion can be achieved through feature concatenation or an attention mechanism.

[0221] Step S2 is clothing matching relationship modeling. Based on the multi-modal clothing feature representation obtained in step S1, a clothing matching graph is constructed. In one embodiment of the present invention, each piece of clothing can be represented as a node, and the matching relationship between clothes is represented as an edge. The weight of the edge can be determined according to the matching frequency or expert scoring. Then, a graph neural network is used to learn the clothing matching graph to obtain node representations and edge representations. A graph convolutional network (GCN) or a graph attention network (GAT) can be used as the basic network structure.

[0222] Step S3 is cross-modal learning. First, a cross-modal graph is constructed, with clothing visual features and attribute text features as nodes of different modalities. Two nodes can be created for each piece of clothing, corresponding to the visual feature and the text feature respectively, and an edge is added between these two nodes to indicate that they belong to the same piece of clothing. Then, a cross-modal graph neural network is used to learn the relationship between nodes of different modalities. A special message passing mechanism can be designed to allow information of different modalities to flow in the graph. Finally, a contrastive learning mechanism is adopted to enhance the consistency of different modal features.

[0223] Step S4 is personalized recommendation generation. First, the scene constraint conditions and personal preferences input by the user are received. This information may include season, occasion, style, color preference, etc. Then, based on the cross-modal learning results, combined with the scene constraint conditions and personal preferences, a personalized clothing matching recommendation list is generated. A graph-based reasoning algorithm can be adopted to perform path search on the clothing matching graph to find the matching combination that best meets the user's needs. Finally, the recommendation list is dynamically adjusted according to user feedback. An online learning algorithm can be designed to update the model parameters in a timely manner to adapt to the user's latest preferences.

[0224] Step S5 is multi-task joint optimization. A multi-task learning framework is designed to optimize clothing attribute prediction, matching relationship classification, and personalized recommendation tasks simultaneously. A shared feature extraction network can be used as the basis for multiple tasks, and then a specific output layer is designed for each task. A soft parameter sharing mechanism is adopted to balance the knowledge transfer between tasks. The similarity of parameters of different tasks can be encouraged through regularization constraints.

[0225] Step S6 is knowledge graph enhancement. First, construct a knowledge graph in the clothing domain to supplement the prior knowledge of clothing matching. Knowledge can be extracted from fashion magazines, expert knowledge, and large-scale clothing datasets. Then, design a knowledge distillation method to transfer the information of the knowledge graph into the neural network model. The knowledge graph embedding can be used as an auxiliary task and jointly optimized with the main recommendation task.

[0226] Step S7 is style transfer and innovation. Implement clothing style transfer based on a generative adversarial network. The architecture of a conditional generative adversarial network (cGAN) can be adopted, where the generator is responsible for converting the input clothing image into the target style, and the discriminator is responsible for distinguishing whether the generated image conforms to the target style. Then, design an adaptive style mixing method to generate new clothing styles. The AdaIN (Adaptive Instance Normalization) technology can be used to achieve flexible style regulation.

[0227] Step S8 is interpretability enhancement. First, design an attention visualization method to display the basis for the model's decision-making. The Grad-CAM technology can be used to generate a heatmap, highlighting the image regions that the model pays the most attention to when making a decision. Then, generate natural language explanations to illustrate the reasons for the recommendation and the matching principles. A combination of a template-based method and a neural language generation model can be adopted to generate more natural and fluent explanatory texts.

[0228] It should be noted that steps S1 to S8 constitute an end-to-end training and inference process, and the overall performance is improved through iterative optimization. In practical applications, some steps can be selectively implemented or the execution order of the steps can be adjusted according to specific requirements and limitations of computing resources.

[0229] Through the above detailed description, the technical solutions of the intelligent clothing matching recommendation system and method based on multimodal learning of the present invention have been fully demonstrated. This system and method can effectively integrate the multimodal information of clothing, establish complex matching relationships, and generate personalized and interpretable recommendation results for users. At the same time, by introducing technologies such as style transfer and knowledge graphs, the innovation ability and knowledge expression ability of the system are further enhanced. These features make the present invention have broad application prospects in the fields of fashion e-commerce, personal styling consultants, virtual fitting, etc.

[0230] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. Intelligent clothing matching recommendation system based on multimodal learning, characterized by ,include: Clothing feature extraction module, used for: Extract visual features from clothing images using convolutional neural networks; Use natural language processing technology to extract semantic features from clothing text descriptions; fusing the visual features and the semantic features to generate a multimodal clothing feature representation; The matching relationship modeling module is connected to the clothing feature extraction module for: Receiving the multimodal clothing feature representation sent by the clothing feature extraction module; Based on the multimodal clothing feature representation, a clothing matching graph is constructed, wherein each clothing item is represented as a node, and matching relationships between clothing items are represented as edges; Using a graph neural network to learn the clothing matching graph, and obtaining node representation and edge representation; The cross-modal learning module is communicatively connected with the collocation relationship modeling module and is used to: Receiving the node representation and edge representation sent by the collocation relationship modeling module; Construct a cross-modal graph, taking clothing visual features and attribute text features as nodes of different modalities; Use cross-modal graph neural networks to learn the relationship between nodes of different modalities; Adopt contrastive learning mechanism to enhance the consistency of different modal features; A personalized recommendation module, which is in communication connection with the cross-modal learning module, is used to: Receive scene constraints and personal preferences input by users; Based on the learning results of the cross-modal learning module, combined with the scene constraints and personal preferences, a personalized clothing matching recommendation list is generated; Dynamically adjust the recommendation list based on user feedback; Among them, the clothing feature extraction module, matching relationship modeling module, cross-modal learning module and personalized recommendation module constitute an end-to-end training framework, and the overall performance is improved through joint optimization.

2. The system according to claim 1, characterized in that , the clothing feature extraction module includes: Visual feature extraction unit, used to: Preprocess the input clothing images, including cropping, scaling, and data augmentation; Use pre-trained convolutional neural networks to extract visual features such as color, style, style, and texture of clothing; Use global average pooling to generate a visual feature vector of fixed dimension; Text feature extraction unit, used for: Segment and clean the input clothing text description; Use pre-trained word embedding models to convert text into vector representations; Use recurrent neural networks or attention mechanisms to extract semantic features of text; Feature fusion unit, used for: Receiving the visual feature vector and the semantic feature vector; Design a multi-layer perceptron network to map visual features and semantic features into a common feature space; The attention mechanism is used to perform weighted fusion of different modal features to generate the final multimodal clothing feature representation.

3. The system according to claim 1, characterized in that , the collocation relationship modeling module includes: Graph building unit, used for: Construct an initial clothing matching graph based on the matching frequency between clothing or expert annotations; The multimodal features of each piece of clothing are used as the initial features of the node; The weight calculation method of the design edge reflects the coordination degree of clothing matching; Graph Convolutional Network Unit, used for: Design a multi-layer graph convolutional network structure to learn the representation of nodes and edges; In each layer of graph convolution, neighbor node information is aggregated and the central node representation is updated; Use skip connections and residual connections to improve the network's expressiveness; Attention mechanism unit, used to: Design a graph attention layer to learn the importance weights between nodes; Implement a multi-head attention mechanism to capture the relationship between nodes from different angles; Combining self-attention and neighbor attention can enhance the expressiveness of the model.

4. The system according to claim 1, characterized in that ,The cross-modal learning module includes: a cross-modal graph building unit, used to: Integrate the visual feature nodes and attribute text feature nodes of clothing into the same graph structure; design heterogeneous edges to connect nodes of different modes; Define the intra-modality and inter-modality adjacency matrices; Cross-modal graph neural network units for: Design a graph neural network structure that can handle heterogeneous nodes and edges; Implement modality-specific information transfer and aggregation mechanisms; Design a cross-modal information fusion layer to integrate features from different modalities; Comparative Learning Unit for: Construct positive sample pairs and negative sample pairs, including intra-modal and inter-modal sample pairs; Design a contrast loss function to bring positive sample pairs closer together and push negative sample pairs further apart; Implement momentum contrastive learning to improve the expressiveness and generalization of the model.

5. The system according to claim 1, characterized in that ,The personalized recommendation module includes: a user portrait unit, used to: Collect and analyze users’ historical behavior data, including browsing, clicking, and purchasing records; Extracting static features of users, such as age and gender, and dynamic features, such as fashion preferences; Construct multi-dimensional user feature vectors; Scene constraint processing unit, used to: Parse the scene constraints entered by the user, such as season, occasion, and style; Convert scene constraints into quantifiable feature representations; Design scenario constraints with soft and hard rules; Matching calculation unit, used for: Design a multi-objective optimization function to balance the coordination, personalization, and diversity of clothing matching; implement a graph-based reasoning algorithm to perform path search on the clothing matching graph; Use reinforcement learning methods to optimize long-term recommendation effects; Dynamic adjustment unit for: Collect real-time feedback from users on recommendation results; Design online learning algorithms and update model parameters in a timely manner; Implement a balanced strategy of exploration and utilization to improve the novelty of recommendations.

6. The system according to claim 1, characterized in that , also includes: The multi-task learning module is communicatively connected with the clothing feature extraction module, the matching relationship modeling module and the cross-modal learning module, and is used to: Design a shared feature extraction network as the basis for multiple tasks; Define multiple related tasks, including clothing attribute prediction, matching relationship classification and personalized recommendation; Design task-specific output layers to meet the needs of different tasks; Implement soft parameter sharing mechanism to balance knowledge transfer between tasks; Design a multi-task loss function to comprehensively optimize the performance of all tasks.

7. The system according to claim 1, characterized in that , also includes: The knowledge graph enhancement module is communicatively connected with the collocation relationship modeling module and the cross-modal learning module, and is used to: Construct a knowledge graph in the clothing field, including clothing entities, attributes and relationships; Design knowledge graph embedding methods to map entities and relationships into low-dimensional vector space; Implement the reasoning mechanism based on knowledge graph to supplement the prior knowledge of clothing matching; Design a knowledge distillation method to transfer the information of the knowledge graph into the neural network model.

8. The system according to claim 1, characterized in that , also includes: The style transfer module is connected to the clothing feature extraction module and the personalized recommendation module for: Realize clothing style transfer based on generative adversarial network; Design style encoder to extract style features of clothing; Implement an adaptive style mixing method to generate new clothing styles; Design style consistency loss to ensure the quality and diversity of generated results.

9. The system according to claim 1, characterized in that , also includes: The explainability enhancement module is connected to the collocation relationship modeling module and the personalized recommendation module for: Design an attention visualization method to show the clothing areas and features that the model focuses on; Implement graph-based explanation methods to track the path of recommendation decisions; Generate natural language explanations to explain the reasons for the recommendation and the matching principles; Design an interactive explanation interface to allow users to adjust the recommendation results.

10. A clothing matching intelligent recommendation method based on multimodal learning based on the system according to any one of claims 1 to 9, characterized in that , including the following steps: S1. Multimodal clothing feature extraction: Extract visual features from clothing images using convolutional neural networks; Use natural language processing technology to extract semantic features from clothing text descriptions; fusing the visual features and the semantic features to generate a multimodal clothing feature representation; S2. Clothing matching relationship modeling: Based on the multimodal clothing feature representation, construct a clothing matching graph; Using a graph neural network to learn the clothing matching graph, and obtaining node representation and edge representation; S3. Cross-modal learning: Construct a cross-modal graph, taking clothing visual features and attribute text features as nodes of different modalities; Use cross-modal graph neural networks to learn the relationship between nodes of different modalities; Adopt contrastive learning mechanism to enhance the consistency of different modal features; S4. Personalized recommendation generation: Receive scene constraints and personal preferences input by users; Based on the cross-modal learning results, combined with scene constraints and personal preferences, a personalized clothing matching recommendation list is generated; Dynamically adjust the recommendation list based on user feedback; S5. Multi-task joint optimization: Design a multi-task learning framework to simultaneously optimize clothing attribute prediction, matching relationship classification, and personalized recommendation tasks; Adopt soft parameter sharing mechanism to balance knowledge transfer between tasks; S6. Knowledge graph enhancement: Construct a knowledge graph in the clothing field to supplement the prior knowledge of clothing matching; Design a knowledge distillation method to transfer the information of the knowledge graph to the neural network model; S7. Style transfer and innovation: Realize clothing style transfer based on generative adversarial network; Design an adaptive style mixing method to generate new clothing styles; S8. Enhanced interpretability: Design attention visualization methods to show the basis of model decisions; Generate natural language explanations to explain the reasons for the recommendation and the matching principles; Among them, steps S1 to S8 constitute an end-to-end training and reasoning process, and the overall performance is improved through iterative optimization.

Citation Information

Cited By

  • Two-dimensional garment making layout generation method and device and computer equipment

    CN120354472A

  • Method, device and computer equipment for generating two-dimensional clothing pattern

    CN120354472B

  • Multi-modal clothing recommendation method

    CN120744187A

  • Image cold start recommendation sorting method and system

    CN121415390A

  • Image cold start recommendation ranking method and system

    CN121415390B