Diversified Explainable Compatibility Modeling Method and System Based on Multimodal Dissociation
Through the comparison and dissociation attribute learning and matching patterns of alignment attributes and candidate attributes, the existing clothing compatibility modeling methods are solved, and the accuracy and interpretability of clothing single items are improved.
Patent Information
- Application Number
- CN202411975483.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The existing clothing compatibility modeling methods are not effective in comprehensive evaluation, and lack fine-grained attribute division and interpretability, resulting in the recommended complementary items not being optimized enough.
The comparative dissociation attribute learning method based on deep mutual information is used to dissociate multimodal features, and the diversity between fashionable items can be calculated by aligning attributes and candidate attributes.
It improves the accuracy of clothing item matching and interpretability of recommendation tasks, and provides a more comprehensive perspective to model the matching process between item attributes.
Smart Images

Figure CN119377901B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of clothing matching solutions, and specifically to a diversified interpretable compatibility modeling method and system based on multimodal dissociation. Background Art
[0002] In recent years, the booming development of the fashion industry has indicated an increasing demand for fashionable clothing. Appropriate clothing matching can reflect a person's good mental outlook. However, due to differences in people's aesthetic concepts, not everyone is good at daily clothing matching, which provides an application background for clothing compatibility modeling solutions.
[0003] The purpose of fashion item compatibility modeling is to evaluate the matching degree between items. Previous work has achieved certain results in complementary item recommendation, but in terms of comprehensive evaluation, there are still the following problems, resulting in poor effects of existing diversified interpretable compatibility modeling:
[0004] (1) Some existing studies attempt to solve the clothing compatibility modeling problem by integrating multimodal information of fashion items. However, these methods usually model the learned fashion item representations in an overall manner and do not effectively divide the fine-grained attributes of items, resulting in a lack of interpretability in the recommendation task.
[0005] (2) A single attribute matching model cannot fully describe the compatibility between fashion items, resulting in the recommended complementary items not being the optimal solution. Summary of the Invention
[0006] To solve the above problems, the present disclosure proposes a diversified interpretable compatibility modeling method and system based on multimodal dissociation. By means of contrast dissociation attribute learning based on deep mutual information, the multimodal features are dissociated, and two matching modes of aligned attributes and candidate attributes are used to measure the compatibility between fashion items, improving the accuracy of item matching.
[0007] According to some embodiments, the present disclosure adopts the following technical solutions:
[0008] A diversified interpretable compatibility modeling method based on multimodal dissociation, comprising:
[0009] Obtaining the multimodal information of two items to be matched;
[0010] Performing feature extraction and fusion on the multimodal information to obtain the multimodal features of each item;
[0011] Dissociating the multimodal features through contrast dissociation attribute learning based on deep mutual information;
[0012] Based on the dissociated multi-modal features, using two matching modes of alignment attributes and candidate attributes, calculate the diversified interpretable compatibility scores of two single items.
[0013] According to some embodiments, the present disclosure adopts the following technical solutions:
[0014] A diversified interpretable compatibility modeling system based on multi-modal dissociation includes:
[0015] An information acquisition module, configured to: acquire the multi-modal information of two single items to be matched;
[0016] A feature extraction module, configured to: perform feature extraction and fusion on the multi-modal information to obtain the multi-modal features of each single item;
[0017] A feature dissociation module, configured to: dissociate the multi-modal features through contrast dissociation attribute learning based on deep mutual information;
[0018] An attribute matching module, configured to: based on the dissociated multi-modal features, use two matching modes of alignment attributes and candidate attributes to calculate the diversified interpretable compatibility scores of two single items.
[0019] According to some embodiments, the present disclosure adopts the following technical solutions:
[0020] A computer program product, including a computer program, which when executed by a processor, implements the diversified interpretable compatibility modeling method based on multi-modal dissociation.
[0021] According to some embodiments, the present disclosure adopts the following technical solutions:
[0022] A non-transitory computer-readable storage medium, which is used to store computer instructions, and when the computer instructions are executed by a processor, implements the diversified interpretable compatibility modeling method based on multi-modal dissociation.
[0023] According to some embodiments, the present disclosure adopts the following technical solutions:
[0024] An electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the diversified interpretable compatibility modeling method based on multi-modal dissociation.
[0025] Compared with the prior art, the beneficial effects of the present disclosure are:
[0026] (1)To avoid the interference of redundant information of single items, the present disclosure designs a contrast dissociation property learning method, which fits the mutual information value between feature attributes through a deep neural network and optimizes the mutual information value between attributes by contrast loss to minimize it, increasing the independence between attributes.
[0027] (2)The present disclosure proposes a diversified and interpretable attribute matching method, which consists of two matching modes of aligned attributes and candidate attributes. The two methods complement each other, modeling the matching process between single-item attributes from a more comprehensive perspective and increasing the interpretability in the complementary single-item recommendation task. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the disclosure. The schematic embodiments and descriptions thereof of the disclosure are used to explain the disclosure and do not constitute an improper limitation of the disclosure.
[0029] Figure 1 is the architecture diagram of Embodiment 1.
[0030] Figure 2 is the visual feature extraction diagram of Embodiment 1.
[0031] Figure 3 is the text feature extraction diagram of Embodiment 1.
[0032] Figure 4 is an example diagram of the matching mode of aligned attributes in Embodiment 1, where (a) is the upper garment 1 to be matched, and (b) and (c) are the lower garments 1-1 and 1-2 to be matched respectively.
[0033] Figure 5 is an example diagram of the matching mode of candidate attributes in Embodiment 1, where (a) and (c) are two matching methods of the same upper garment, and (b) is a schematic diagram of the Müller-Lyer illusion. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0035] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.
[0036] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "comprising" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0037] Embodiment 1
[0038] In an embodiment of the present disclosure, a diversified interpretable compatibility modeling method based on multimodal dissociation is provided, including:
[0039] Obtain the multimodal information of two single items to be matched;
[0040] Extract and fuse the features of the multimodal information to obtain the multimodal features of each single item;
[0041] Dissociate the multimodal features through contrastive dissociation attribute learning based on deep mutual information;
[0042] Based on the dissociated multimodal features, use two matching modes of alignment attributes and candidate attributes to calculate the diversified interpretable compatibility scores of the two single items.
[0043] As an embodiment, the diversified interpretable compatibility modeling method based on multimodal dissociation of the present disclosure, as Figure 1 shown, adopts contrastive dissociation attribute learning based on deep mutual information to extract the feature values of each attribute of the single item, and uses two matching modes of alignment attributes and candidate attributes to measure the compatibility between fashion single items. The specific implementation process is described in detail below:
[0044] I. Problem Definition
[0045] Suppose there is a set of upper garments and a set of lower garments, which are respectively and , where and respectively represent the total number of upper garments and lower garments. For each upper garment and lower garment in the set, both contain information of two modalities (i.e., images and text descriptions); use and to represent the visual features and text features of the upper garment (lower garment) respectively, where, and respectively represent their corresponding feature dimensions; the compatibility modeling method for complementary fashion single items aims to study which lower garment single item when a given upper garment single item It is more compatible with it.
[0046] This embodiment proposes a diversified interpretable compatibility modeling method based on multimodal dissociation , this method can dissociate the multimodal information of a single product, thereby improving the interpretability of the recommendation task. In addition, by simultaneously considering the alignment attribute compatibility score and candidate attribute compatibility scores , and obtain the final diversified interpretable compatibility score, which improves the performance of diversified compatibility modeling.
[0047] 2. Multimodal Feature Extraction and Fusion
[0048] 1. Visual feature extraction
[0049] At present, pre-trained models have been widely used and have made remarkable achievements in the field of computer vision. Therefore, this embodiment uses the pre-trained network ViT to (Bottoms ) to extract visual features from the original image ,like Figure 2 As shown in Figure 2, the specific operations of the ViT model are:
[0050] First, put the top (Bottoms ) is divided into multiple image patches, which are then converted into low-dimensional embeddings through linear projection. The image patch embeddings are then combined with the corresponding position embeddings as a sequence and fed into a 12-layer transformer to capture contextual information. Finally, the final visual features Generated by a Multilayer Perceptron (MLP).
[0051] 2. Text feature extraction
[0052] For the top (Bottoms ) ( ), using the BERT model with strong generalization ability; this embodiment selects the BERT basic version consisting of 12 transformers, such as Figure 3 As shown, with top Take the text feature extraction of BERT model as an example. The text description is decomposed into individual words, which are then converted into word embedding vectors, and finally encoded through a set of transformers to obtain text features. .
[0053] 3. Feature Fusion
[0054] Different modalities represent different attribute information of fashion items. To fully explore the potential of multi-modalities in compatibility modeling, visual features and text features ( ) are connected to obtain multi-modal features . Among them, is the number of dimensions of the multi-modal features, and each dimension is used as a feature attribute.
[0055] III. Contrastive dissociation attribute learning based on deep mutual information
[0056] The compatibility between complementary items is largely affected by the attributes of the items. Existing multi-modal compatibility modeling methods mainly calculate the compatibility directly according to the overall features in the latent compatibility space (such as the style space); however, these methods ignore the influence of repeated and redundant information in multi-modal features on compatibility modeling, which reduces the interpretability of the recommendation process; to enhance the independence of multi-modal features and improve the interpretability of compatibility modeling, this embodiment designs a contrastive dissociation attribute learning method based on deep mutual information, aiming to dissociate the multi-modal features of the items, that is, to dissociate the multi-modal features, which is divided into two major steps:
[0057] (1) Fitting the mutual information value: Different dimensions of multi-modal features represent different feature attributes. Since it is difficult to accurately calculate the mutual information value between high-dimensional continuous representations, a neural network is used to approximate the mutual information value between the feature attributes of the items.
[0058] (2) Optimizing the mutual information value: The smaller the mutual information value between feature attributes, the smaller the degree of dependence and the higher the independence of the attributes, and the better the dissociation effect. Therefore, a contrastive loss is designed to use different feature attributes of the same item as the positive pair and randomly select a certain feature attribute of other items as the negative pair to minimize the mutual information value between the feature attributes of the items.
[0059] Specifically, mutual information is used to measure the degree of dependence between variables. Given two variables and , the stronger their independence and the smaller the degree of dependence, the lower the mutual information, and vice versa. The mutual information can be expressed in the following form:
[0060]
[0061] Among them, represents the joint distribution of variable and variable , and and represent the marginal distributions of the variables respectively.
[0062] This embodiment uses mutual information to quantify the degree of dependence between feature variables in multimodal features, and dissociates multimodal features based on mutual information. Multimodal features For example, Different dimensions of multimodal features, that is, different feature attributes in multimodal features, are finally dissociated into However, it is difficult to accurately calculate multimodal features The mutual information value between medium and high dimensional continuous variables is approximated by calculating the mathematical expectation of the neural network output value.
[0063] According to the above formula, the top No. Dimensional representation (i.e. feature attributes) and Dimensional representation (i.e. feature attributes) It can be expressed as:
[0064]
[0065] In order to further improve the independence between the characteristic attributes of fashion items and reduce the degree of dependence between them, the following contrastive loss is constructed to train the neural network for dissociation by minimizing the degree of dependence (i.e. maximizing independence):
[0066] =:
[0067] in, Represents the negative sample feature attributes randomly extracted from other items. A similar method can be used to obtain the lower garment The dissociation characteristics of .
[0068] 4. Attribute matching based on alignment attributes and candidate attributes
[0069] Based on the dissociated multimodal features of the products, diversified and interpretable attribute matching modeling is carried out, and the compatibility scores between complementary products are evaluated according to the matching scores, which is divided into three steps:
[0070] (1) Matching mode of alignment attributes and modeling the compatibility of alignment attributes: The representation of the same dimension of the top and bottom (the same feature attribute) is recorded as the alignment attribute, and the compatibility score between the alignment attributes is calculated based on the weights between the attributes. This is also the most common matching method that people think of in daily matching, such as matching between colors and stripes.
[0071] (2)Candidate attribute matching pattern for candidate attribute compatibility modeling: To address the issue of the single complementary item attribute matching pattern, candidate attribute compatibility modeling is additionally designed to find the attribute pair with the highest weight between the upper and lower garments, excluding the aligned attributes, as the candidate attributes to supplement the attribute matching method.
[0072] (3)Calculate the compatibility score: The scores of the two attribute matching methods are weighted and summed, and the optimal value of the weight is verified in subsequent experiments. This score is the final compatibility score between complementary items.
[0073] Specifically, based on the dissociated feature representation (i.e., the dissociated multi-modal features) and further complex compatibility rules are designed, including: 1) The matching pattern of aligned attributes. 2) The matching pattern of candidate attributes, learning diverse matching rules between the feature attributes of single items from two perspectives to improve the interpretability of the complementary recommendation task.
[0074] 1. The matching pattern of aligned attributes
[0075] In actual clothing combinations, complementary items are usually selected based on aligned attributes. For example, Figure 4 the lower garments shown in (b) and (c) are both plaid pleated skirts. However, Figure 4 the bow of the upper garment shown in (a) matches the color of the lower garment shown in (b), and both have the text description "blue plaid". Therefore, considering the aligned attribute (i.e., color) between complementary items, Figure 4 the lower garment shown in (b) is more suitable for the upper garment shown in (a) than the lower garment shown in (c).
[0076] Therefore, for the matching pattern of aligned attributes, given an upper garment and a target lower garment , first calculate the alignment weight , which is expressed by the formula:
[0077]
[0078] where are the -th feature attributes of the upper garment and the target lower garment respectively, , and are learnable network parameters, is the sigmoid activation function.
[0079] Based on the alignment weight , calculate the aligned attribute compatibility score :
[0080]
[0081] 2. Matching Patterns of Candidate Attributes
[0082] In fact, relying solely on the alignment attributes between complementary items for compatibility evaluation is insufficient. It is necessary to combine candidate attributes to improve the performance of compatibility modeling. For example, Figure 5 In (a) and (c) of , the "black and white striped vest" (upper garment) and two "black long skirts" (lower garments) are highly matched in terms of alignment attributes (i.e., color). However, when further considering the characteristics of the neckline and hem, Figure 5 a more ideal match appears between the "wide-leg long skirt" and the "V-neck vest" shown in (a) of , which may be mainly attributed to Figure 5 the Müller-Lyer illusion shown in (b) of , which makes the match in (a) of visually elongate a person's height. Therefore, design the matching pattern of candidate attributes and calculate the candidate weights Figure 5 as follows: as follows:
[0083]
[0084] Different from the compatibility modeling of alignment attributes, select the characteristic attribute with the largest candidate weight as the candidate attribute; thus, the candidate attribute compatibility score between the upper garment and the lower garment can be obtained, as shown below:
[0085]
[0086] where, in represents the th characteristic attribute of the upper garment , represents the th characteristic attribute of the target lower garment . The maximum weight is obtained according to Max in the previous formula, indicating that the weight between the th characteristic attribute of the upper garment and the th characteristic attribute of the lower garment except for the alignment attribute is the largest, so it is selected as the candidate attribute.
[0087] V. Target Loss:
[0088] Based on the diversified interpretable compatibility scores and , the total score of the recommendation task can be calculated , namely the diversified interpretable compatibility score, the calculation formula is as follows:
[0089]
[0090] Among them, is a balance parameter used to control the weights of the matching patterns of alignment attributes and candidate attributes.
[0091] Form a positive example set with the clothing items matched by fashion experts , and regard some unobserved combinations as negative examples. To accurately model the compatibility between items, using the Bayesian Personalized Ranking (BPR) framework, the following triples are constructed based on the public dataset :
[0092]
[0093] Among them, the triple represents that the lower clothing in the positive example set is more compatible with the given upper clothing than the lower clothing . Thus, the following loss can be constructed:
[0094]
[0095] Jointly optimize the contrast loss , and obtain the following final objective loss :
[0096]
[0097] Among them, and are non-negative balance parameters.
[0098] VI. Experimental Analysis:
[0099] To comprehensively evaluate the experimental effect of this solution (abbreviated as DICM), comparative experiments are carried out with existing baseline models on two datasets (IQON3000 and Polyvore), and the area under the curve (AUC), mean reciprocal rank (MRR), hit rate (HR@K), and normalized discounted cumulative gain (NDCG@K) are used as evaluation indicators to measure the effect of clothing complementary recommendation. Table 1 shows the performance of different methods on the two datasets based on the AUC and MRR evaluation indicators, and Table 2 shows the performance of the methods with an AUC value not lower than 0.8 in Table 1 on the two datasets based on the HR@K and MRR evaluation indicators, where K is the number of lower clothing items in the candidate sequence:
[0100] Table 1 Performance of different methods based on AUC and MRR evaluation metrics
[0101]
[0102] Table 2 Performance of different methods based on HR@K and NDCG@K evaluation metrics
[0103]
[0104] From these results, it can be found that the DICM of this solution is significantly better than the existing benchmark models on both datasets.
[0105] Example 2
[0106] In an embodiment of the present disclosure, a diversified interpretable compatibility modeling system based on multimodal dissociation is provided, including:
[0107] An information acquisition module, configured to: acquire multimodal information of two single items to be matched;
[0108] A feature extraction module, configured to: perform feature extraction and fusion on the multimodal information to obtain multimodal features of each single item;
[0109] A feature dissociation module, configured to: dissociate the multimodal features through contrast dissociation attribute learning based on deep mutual information;
[0110] An attribute matching module, configured to: based on the dissociated multimodal features, use two matching modes of alignment attributes and candidate attributes to calculate the diversified interpretable compatibility scores of the two single items.
[0111] Example 3
[0112] In an embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the diversified interpretable compatibility modeling method based on multimodal dissociation.
[0113] Example 4
[0114] In an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, and the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the diversified interpretable compatibility modeling method based on multimodal dissociation is implemented.
[0115] Example 5
[0116] In an embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the method for realizing the diversified interpretable compatibility modeling based on multimodal dissociation.
[0117] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate computer-implemented processing, and thus the instructions executed on the computer or other programmable devices provide steps for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0119] Although the specific embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.
Claims
1. A diversified interpretable compatibility modeling method based on multimodal dissociation, characterized by: include: Obtain multimodal information of two items to be matched; Perform feature extraction and fusion on multimodal information to obtain multimodal features of each item; Dissociate multimodal features through contrastive dissociative attribute learning based on deep mutual information; Based on the dissociated multimodal features, the diversified interpretable compatibility scores of the two items are calculated using two matching modes of alignment attributes and candidate attributes. The multimodal information includes images and text descriptions; The matching mode of the alignment attribute is to obtain the alignment weight between the same characteristic attributes of two items through learning, and calculate the alignment attribute compatibility score of the two items to be matched based on the alignment weight; The matching mode of the candidate attributes is to obtain the candidate weights between different feature attributes of two items by learning, and calculate the candidate attribute compatibility scores of the two items to be matched based on the feature attribute with the largest candidate weight.
2. The diversified interpretable compatibility modeling method based on multimodal dissociation according to claim 1, characterized in that: The feature extraction and fusion of multimodal information is specifically as follows: The ViT network and BERT model are used to extract visual features and text features from the images and text descriptions of individual products respectively; The visual features and text features are integrated to obtain multimodal features.
3. The diversified interpretable compatibility modeling method based on multimodal dissociation according to claim 1, characterized in that: The multimodal features are dissociated by learning and training a neural network based on contrast dissociation attributes of deep mutual information, using the trained neural network to dissociate the input multimodal features, and outputting the dissociated multimodal features.
4. The diversified interpretable compatibility modeling method based on multimodal dissociation according to claim 3, characterized in that: The contrastive dissociation attribute learning based on deep mutual information utilizes the mutual information values between feature variables in multimodal features to quantify the degree of dependence between feature variables, and minimizes the degree of dependence between feature variables by constructing contrast loss, thereby training the neural network.
5. A diversified interpretable compatibility modeling system based on multimodal dissociation, characterized by: include: The information acquisition module is configured to: acquire multimodal information of two items to be matched; The feature extraction module is configured to: perform feature extraction and fusion on the multimodal information to obtain the multimodal features of each single product; A feature dissociation module is configured to: dissociate the multimodal features through contrastive dissociation attribute learning based on deep mutual information; The attribute matching module is configured to: calculate the diversified interpretable compatibility scores of two items based on the dissociated multimodal features using two matching modes of alignment attributes and candidate attributes; The multimodal information includes images and text descriptions; The matching mode of the alignment attribute is to obtain the alignment weight between the same characteristic attributes of two items through learning, and calculate the alignment attribute compatibility score of the two items to be matched based on the alignment weight; The matching mode of the candidate attributes is to obtain the candidate weights between different feature attributes of two items by learning, and calculate the candidate attribute compatibility scores of the two items to be matched based on the feature attribute with the largest candidate weight.
6. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the diversified interpretable compatibility modeling method based on multimodal dissociation described in any one of claims 1 to 4 is implemented.
7. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the diversified interpretable compatibility modeling method based on multimodal dissociation as described in any one of claims 1-4 is implemented.
8. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device implements the diversified interpretable compatibility modeling method based on multimodal dissociation as described in any one of claims 1-4.
Citation Information
Patent Citations
Multi-modal emotion recognition method and system based on discriminant learning
CN116226635A
Dress collocation method and system based on attention knowledge extraction, and storage medium
WO2019223302A1