A personalized intelligent recommendation method and system based on clothing wearing tags
Patent Information
- Application Number
- CN202610927634.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-25
AI Technical Summary
本发明解决了现有技术存在的未考虑3D体型精细建模、多模态特征对齐不足、场景冲突处理生硬以及全局穿搭寻优能力弱的问题
1)通过2D照片重建3D人体网格并提取体型特征向量,实现了对用户真实体型的高精度参数化建模,避免了传统推荐中因忽略体型导致的穿搭失衡问题;
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent recommendation technology, and in particular to a personalized intelligent recommendation method and system based on clothing matching tags. Background Technology
[0002] With the rapid development of e-commerce, online clothing shopping has become an important part of people's daily lives. However, faced with a massive amount of clothing products, users often find it difficult to quickly find outfits that match their preferences, body type, and current situation. Existing clothing recommendation technologies are usually based on collaborative filtering or simple deep learning models, and mainly suffer from the following technical shortcomings: 1) Ignoring body shape differences: Existing methods rely heavily on users' historical clicks or purchase records, lacking a detailed depiction of users' actual body shapes. This results in recommended clothing not matching the user's body shape in terms of wearing effect, such as visual imbalance and other problems.
[0003] 2) Superficial multimodal feature fusion: Clothing recommendation involves multiple modalities such as images, text and tags. Existing methods usually use simple splicing and fusion, which fails to effectively eliminate the semantic gap between modalities, resulting in inaccurate feature representation.
[0004] 3) Inflexible handling of conflicts between scene constraints and user preferences: In actual dressing, users' style preferences are often limited by scene constraints such as weather and occasion. Existing methods often directly filter out clothing that does not fit the scene or ignore scene constraints, lacking an intelligent compromise mechanism for conflicts between preferences and constraints.
[0005] 4) Insufficient ability to optimize overall outfit selection: Outfit recommendation is a multi-objective optimization problem that needs to consider visual harmony, body shape balance, and scene suitability. Existing methods usually use a greedy strategy to recommend individual items one by one, which makes it difficult to guarantee the optimality of the overall outfit combination.
[0006] Therefore, there is an urgent need for a personalized intelligent recommendation method that can deeply integrate multimodal features, accurately model body shape, intelligently handle scene conflicts, and perform global optimization. Summary of the Invention
[0007] This invention provides a personalized intelligent recommendation method and system based on clothing matching tags. This invention solves the problems of existing technologies, such as failure to consider detailed 3D body modeling, insufficient multimodal feature alignment, simplistic handling of scene conflicts, and weak global matching optimization capabilities.
[0008] In a first aspect, embodiments of the present invention provide a personalized intelligent recommendation method based on clothing matching tags, the method comprising: Define a clothing outfit tag library that includes clothing attribute tags, user feature tags, and scene constraint tags. Reconstruct a 3D human body mesh from a 2D front and side view photo of a user, generate a body shape parameter vector, and encode it as a 3D body shape feature vector. Based on the clothing tag library, the tag embedding features, visual features and semantic features of each clothing item are extracted, and cross-modal contrastive learning alignment is performed through InfoNCE loss to generate a unified cross-modal embedding vector set for clothing. Calculate the user's user preference vector, scene constraint vector and their conflict index, generate a scene compromise vector, and combine it with the corresponding 3D body feature vector to generate a global state vector; The global state vector is input into the personalized intelligent recommendation model, and the initial candidate outfit feature combination is sampled and output from the unified cross-modal embedding vector set of clothing. Based on the initial candidate outfit feature combinations, an improved artificial rabbit optimization algorithm is used for iterative optimization to generate the globally optimal outfit feature vector. The vector is then retrieved from a unified cross-modal embedding vector set for clothing to output a personalized intelligent recommendation list.
[0009] The technical solution provided in this application has at least the following beneficial effects: 1) By reconstructing 3D human body meshes from 2D photos and extracting body shape feature vectors, high-precision parametric modeling of users' real body shapes is achieved, avoiding the problem of unbalanced clothing caused by ignoring body shape in traditional recommendations; 2) Based on InfoNCE loss, a visual-text-label triple contrast learning mechanism is constructed, which effectively narrows the semantic distance between different modalities of the same product. The unified cross-modal embedding vector generated by the gating mechanism has stronger representation ability and robustness. 3) Introducing a conflict index to determine the degree of conflict between user preferences and scenario constraints, and using a self-attention mechanism to generate scenario compromise vectors, achieves intelligent balance between preferences and constraints, avoiding the rigidity of recommendations caused by simple filtering; 4) An initial candidate combination is generated using a personalized intelligent recommendation model based on the Actor-Critic architecture. An improved artificial rabbit optimization algorithm (combining chaotic mapping initialization, gray wolf cooperative thinking, Cauchy mutation and other strategies) is further introduced to perform global optimization. By comprehensively considering visual coordination, body balance and scene fit, the global optimality and reasonableness of the final recommendation combination are ensured, which significantly improves user experience and recommendation conversion rate.
[0010] In one optional implementation, a clothing tag library is defined, including clothing attribute tags, user feature tags, and scene constraint tags. A 3D human body mesh is reconstructed from a 2D frontal / side profile photo of the user, generating a body shape parameter vector and encoding it as a 3D body shape feature vector, including: Define a clothing outfit tag library that includes clothing attribute tags, user feature tags, and scene constraint tags, and construct the tag word embedding matrix of the clothing outfit tag library; Input a 2D photo of the user's front and side profile into a pre-trained parametric human body model, reconstruct a 3D human body mesh, and fit and output the corresponding body shape parameter vector. The body shape parameter vector is input into a pre-trained nonlinear mapping model to generate a normalized 3D body shape feature vector.
[0011] In one optional implementation, based on a clothing tag library, the tag embedding features, visual features, and semantic features of each clothing item are extracted. Cross-modal contrastive learning alignment is then performed using InfoNCE loss to generate a unified cross-modal embedding vector set for clothing, including: Call the clothing product library and extract the preset clothing attribute tag set, clothing image and descriptive text for each clothing product in the clothing product library; Based on the preset clothing attribute tag set of clothing products, the tag word embedding matrix of the clothing matching tag library is searched to obtain the retrieval tag word vector. The retrieval tag word vector is then input into the tag encoding model to generate the tag embedding feature of the clothing product. The images of the same clothing item are input into a pre-trained visual feature extraction model to generate corresponding visual features. The descriptive text of the same clothing product is input into a pre-trained semantic feature extraction model to generate corresponding semantic features; Based on the InfoNCE loss function, cross-modal contrastive learning alignment is performed on label embedding features, visual features, and semantic features to obtain trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features. Based on the gating mechanism, the trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features are fused together and fully connected dimensionality reduction is performed to obtain the unified cross-modal embedding vector of the current clothing product. Traverse all clothing items in the clothing product library to obtain a unified cross-modal embedding vector set for clothing.
[0012] In one optional implementation, based on the InfoNCE loss function, cross-modal contrastive learning alignment is performed on the label embedding features, visual features, and semantic features to obtain the trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features, including: In the shared embedding space, a triple contrastive learning mechanism including visual-text, visual-label, and text-label is constructed, and the InfoNCE loss function is built. The tag embedding features, visual features, and semantic features of the same clothing product are projected onto a shared embedding space of the same dimension through independent linear mapping layers, and then L2 normalization is performed to obtain the corresponding projected tag embedding features, projected visual features, and projected semantic features. The projected label embedding features, projected visual features, and projected semantic features of any clothing item are used as the positive sample group, while the projected label embedding features, projected visual features, and projected semantic features of other clothing items are used as the negative sample group to obtain the contrastive learning alignment training sample set. Based on the InfoNCE loss function, the contrastive learning aligned training sample set is input, and the backpropagation algorithm is used to optimize the parameters of the visual feature extraction model, semantic feature extraction model, label encoding model and linear mapping layer until the InfoNCE loss converges, resulting in the trained aligned visual feature extraction model, trained aligned semantic feature extraction model and trained aligned label encoding model. Using the trained and aligned visual feature extraction model, the trained and aligned semantic feature extraction model, and the trained and aligned label encoding model, the feature extraction steps for clothing products are repeated to obtain trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features.
[0013] In one optional implementation, the user's preference vector, scene constraint vector, and their conflict index are calculated to generate a scene compromise vector, which is then combined with the corresponding 3D body shape feature vector to generate a global state vector, including: Based on the user's historical interaction records, the user's preference weight for each clothing attribute tag in the clothing matching tag library is calculated using the exponential decay function. The preference weight is then weighted and summed with the word embedding vector of the corresponding clothing attribute tag to obtain the user preference vector. Based on the user's real-time scene labels, a pre-trained mapping layer is used to perform mapping to obtain the corresponding scene constraint vector; Calculate the corresponding conflict index based on the user preference vector and the scenario constraint vector; If the conflict index is greater than 0, a clothing matching conflict is determined to have occurred. A self-attention mechanism is introduced to calculate the compromise weight and generate a scenario compromise vector based on the user preference vector and the scenario constraint vector. The scene compromise vector is combined with the 3D body feature vector to obtain the global state vector; If the conflict index is less than or equal to 0, the user preference vector is used as the scene compromise vector and combined with the 3D body feature vector to obtain the global state vector.
[0014] In one optional implementation, the personalized intelligent recommendation model includes an upper garment recommendation agent, a lower garment recommendation agent, and an accessory recommendation agent, all built on the Actor-Critic architecture. The input states of the top clothing recommendation agent, bottom clothing recommendation agent, and accessory recommendation agent are all global state vectors, and their outputs are top clothing recommendation features, bottom clothing recommendation features, and accessory recommendation features, respectively, which are used to construct a combination of clothing features.
[0015] In one alternative implementation, the global state vector is input into the personalized intelligent recommendation model, and initial candidate outfit feature combinations are sampled and output from a unified cross-modal embedding vector set for clothing, including: The global state vector is input into the top clothing recommendation agent, bottom clothing recommendation agent, and accessory recommendation agent of the personalized intelligent recommendation model, respectively. Based on the global state vector, the top recommendation agent samples and outputs the corresponding initial candidate top recommendation features from the unified cross-modal embedding vector set of clothing. Based on the global state vector, the lower garment recommendation agent samples and outputs the corresponding initial candidate lower garment recommendation features from the unified cross-modal embedding vector set of clothing. Based on the global state vector, the accessory recommendation agent samples and outputs the corresponding initial candidate accessory recommendation features from the unified cross-modal embedding vector set of clothing. The initial candidate top recommendation features, initial candidate bottom recommendation features, and initial candidate accessory recommendation features are combined to obtain the corresponding initial candidate outfit feature combinations.
[0016] In one alternative implementation, based on the initial candidate outfit feature combinations, an improved artificial rabbit optimization algorithm is used for iterative optimization to generate the globally optimal outfit feature vector. This vector is then retrieved from a unified cross-modal embedding vector set for clothing, outputting a personalized intelligent recommendation list, including: The outfit feature vector, composed of the recommended features of tops, bottoms, and accessories, is encoded into the position vector of an individual in the improved artificial rabbit optimization algorithm, and a fitness function is defined. Based on the fitness function and the initial candidate outfit feature combination, the improved artificial rabbit optimization algorithm is used to iteratively optimize the outfit feature vector to obtain the globally optimal outfit feature vector. Based on the globally optimal outfit feature vector, the KNN algorithm is used to search the unified cross-modal embedding vector set of clothing to obtain several actual clothing product combinations whose cosine similarity with the globally optimal outfit feature vector meets a preset threshold. These combinations are then output as a personalized intelligent recommendation list.
[0017] In one alternative implementation, based on the fitness function and the initial candidate outfit feature combination, an improved artificial rabbit optimization algorithm is used to iteratively optimize the outfit feature vector to obtain the globally optimal outfit feature vector, including: The chaotic sequence is generated using the Logistic chaotic mapping iterative formula, and then mapped to the solution space of individuals in the improved artificial rabbit optimization algorithm to obtain an initial population including several initial individuals. Based on the fitness function, calculate the fitness value of each individual in the initial population or the updated population of the previous iteration, and determine the top three individuals in fitness ranking, the second best individual, the third best individual, and the worst individual with the worst fitness based on the fitness value. In the initial population, the worst individual is replaced by the individual corresponding to the initial candidate clothing feature combination, resulting in an optimized population. Calculate the attenuation energy factor. If the attenuation energy factor is greater than 1, proceed to the detour foraging process; otherwise, proceed to the random cave hiding process. The process of foraging by detour is initiated, the running distance is calculated, the cooperative concept of gray wolves is introduced, and the guidance center is calculated based on the best, second best, and third best individuals. Based on the guidance center and the running distance, the position of the optimized population or the population updated in the previous iteration is updated to obtain the population updated in the current iteration. Initiate a random cave hiding process, calculate the running length, generate a one-dimensional cave, and based on the running length and the one-dimensional cave, update the position of the optimized population or the updated population of the previous iteration to obtain the updated population of the current iteration. If the optimal individual remains unchanged for several consecutive iterations, then Cauchy mutation is performed on the optimal individual, and the individual with the better fitness value is retained as the optimal individual. Repeatedly update the position of the population. When the current iteration reaches the maximum number of iterations or the fitness value of the best individual meets the requirements, terminate the iterative update of the population and output the final best individual. Decode the position vector of the optimal individual to obtain the global optimal outfit feature vector.
[0018] Secondly, embodiments of the present invention provide a personalized intelligent recommendation system based on clothing matching tags, used to implement a personalized intelligent recommendation method. The system includes: The clothing and outfit tag definition unit is used to define a clothing and outfit tag library that includes clothing attribute tags, user feature tags, and scene constraint tags. It reconstructs a 3D human body mesh from a 2D front and side view photo of a user, generates a body shape parameter vector, and encodes it into a 3D body shape feature vector. The feature extraction unit is used to extract the tag embedding features, visual features and semantic features of each clothing item based on the clothing matching tag library, and to perform cross-modal contrastive learning alignment through InfoNCE loss to generate a unified cross-modal embedding vector set for clothing. The global state generation unit is used to calculate the user's user preference vector, scene constraint vector and their conflict index, generate the scene compromise vector, and combine it with the corresponding 3D body feature vector to generate the global state vector. The intelligent recommendation unit is used to input the global state vector into the personalized intelligent recommendation model and sample and output the initial candidate outfit feature combination from the uniform cross-modal embedding vector set of clothing. The iterative optimization unit is used to iteratively optimize based on the initial candidate outfit feature combination using an improved artificial rabbit optimization algorithm, generate the globally optimal outfit feature vector, and retrieve it from the unified cross-modal embedding vector set of clothing to output a personalized intelligent recommendation list.
[0019] A third aspect of this invention provides an electronic device, which includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, such that the at least one processor can perform the method proposed in the first aspect of the present invention.
[0020] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described in the first aspect of the present invention. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the electronic device structure of the hardware operating environment involved in the embodiments of the present invention; Figure 2 This is a flowchart illustrating the steps of a personalized intelligent recommendation method based on clothing matching tags provided in an embodiment of the present invention. Figure 3 This is a functional unit diagram of a personalized intelligent recommendation system based on clothing matching tags provided in an embodiment of the present invention. Detailed Implementation
[0022] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0023] The present invention will be further described below with reference to the accompanying drawings.
[0024] Reference Figure 1 , Figure 1 This is a schematic diagram of the electronic device structure of the hardware operating environment involved in the embodiments of the present invention.
[0025] like Figure 1 As shown, the electronic device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0026] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0027] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and an electronic program for a personalized intelligent recommendation system based on clothing matching tags.
[0028] exist Figure 1 In the electronic device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the electronic device of the present invention can be set in the electronic device. The electronic device calls the electronic program of the personalized intelligent recommendation system based on clothing matching tags stored in the memory 1005 through the processor 1001, and executes the personalized intelligent recommendation method based on clothing matching tags provided in the embodiment of the present invention.
[0029] Reference Figure 2 The present invention provides a personalized intelligent recommendation method based on clothing matching tags, the method comprising: S201: Define a clothing tag library including clothing attribute tags, user feature tags, and scene constraint tags; reconstruct a 3D human body mesh from a 2D front and side view photo of a user; generate body shape parameter vectors and encode them as 3D body shape feature vectors. S202: Based on the clothing tag library, extract the tag embedding features, visual features and semantic features of each clothing item, and perform cross-modal contrastive learning alignment through InfoNCE loss to generate a unified cross-modal embedding vector set for clothing. S203: Calculate the user's user preference vector, scene constraint vector and their conflict index, generate a scene compromise vector, and combine it with the corresponding 3D body feature vector to generate a global state vector; S204: Input the global state vector into the personalized intelligent recommendation model, and sample and output the initial candidate outfit feature combination from the uniform cross-modal embedding vector set of clothing. S205: Based on the initial candidate outfit feature combination, the improved artificial rabbit optimization algorithm is used for iterative optimization to generate the global optimal outfit feature vector, and then the vector is retrieved from the unified cross-modal embedding vector set of clothing to output a personalized intelligent recommendation list.
[0030] The technical solution provided in this application has at least the following beneficial effects: 1) By reconstructing 3D human body meshes from 2D photos and extracting body shape feature vectors, high-precision parametric modeling of users' real body shapes is achieved, avoiding the problem of unbalanced clothing caused by ignoring body shape in traditional recommendations; 2) Based on InfoNCE loss, a visual-text-label triple contrast learning mechanism is constructed, which effectively narrows the semantic distance between different modalities of the same product. The unified cross-modal embedding vector generated by the gating mechanism has stronger representation ability and robustness. 3) Introducing a conflict index to determine the degree of conflict between user preferences and scenario constraints, and using a self-attention mechanism to generate scenario compromise vectors, achieves intelligent balance between preferences and constraints, avoiding the rigidity of recommendations caused by simple filtering; 4) An initial candidate combination is generated using a personalized intelligent recommendation model based on the Actor-Critic architecture. An improved artificial rabbit optimization algorithm (combining chaotic mapping initialization, gray wolf cooperative thinking, Cauchy mutation and other strategies) is further introduced to perform global optimization. By comprehensively considering visual coordination, body balance and scene fit, the global optimality and reasonableness of the final recommendation combination are ensured, which significantly improves user experience and recommendation conversion rate.
[0031] In one optional implementation, a clothing tag library is defined, including clothing attribute tags, user feature tags, and scene constraint tags. A 3D human body mesh is reconstructed from a 2D frontal / side profile photo of the user, generating a body shape parameter vector and encoding it as a 3D body shape feature vector, including: S2011: Define a clothing outfit tag library including clothing attribute tags, user feature tags, and scene constraint tags, and construct the tag word embedding matrix of the clothing outfit tag library; In this embodiment, the clothing attribute tags include style category, style positioning, color scheme, fabric material, layering, and cut. Styles and categories: Tops, bottoms, coats, dresses, suits, shoes and accessories; Style categories: commuter, street, retro, minimalist, sporty, French, dark, etc. Color schemes: cool colors, warm colors, Morandi colors, basic black, white and gray, highly saturated contrasting colors, etc. Fabric materials: cotton, linen, silk, wool, knitwear, chemical fiber, leather, chiffon, etc.; Layering of clothing: inner layer, middle layer for warmth, outer layer for wind protection, and accessories for decoration; Cut styles: fitted, slim fit, loose, oversized, A-line, H-line, X-line, etc.; The user feature tags include body shape data, skin color features, style preferences, and material and color preferences; Body shape data: apple shape, pear shape, H shape, X shape, inverted triangle shape (corresponding to 3D body shape parametric features); Skin tone characteristics: cool skin tone, warm skin tone, neutral skin tone; Style preferences: Preference style words extracted from user history (e.g., preference for retro, preference for minimalism); Material and color preferences: Implicit preferences such as frequently clicked color tags and frequently purchased material tags; The scenario constraint labels include weather parameters, occasion attributes, and time nodes; Weather parameters: temperature range (e.g., <10℃, 10-20℃, >25℃), weather conditions (sunny, rainy, snowy, strong wind); Occasion attributes: Work commuting, daily leisure, dinner parties, sports and outdoor activities, dating, travel; Time periods: spring, summer, autumn, winter, and daytime and nighttime; S2012: Input the user's front and side 2D photo into the pre-trained parametric human body model, reconstruct a 3D human body mesh, and fit and output the corresponding body shape parameter vector. The parametric human body model is constructed based on the Skinned Multi-Person Linear Model (SMPL) algorithm. S2013: Input the body shape parameter vector into a pre-trained nonlinear mapping model to generate a normalized 3D body shape feature vector. The nonlinear mapping model is constructed based on the Multilayer Perceptron (MLP) algorithm. In this embodiment, the nonlinear mapping model contains three hidden layers (with dimensions of 64, 128, and 256, respectively), the activation function is ReLU, and the final output layer outputs a 256-dimensional 3D body feature vector normalized to the interval [-1, 1].
[0032] In one optional implementation, based on a clothing tag library, the tag embedding features, visual features, and semantic features of each clothing item are extracted. Cross-modal contrastive learning alignment is then performed using InfoNCE loss to generate a unified cross-modal embedding vector set for clothing, including: S2021: Call the clothing product library and extract the preset clothing attribute tag set, clothing image and description text for each clothing product in the clothing product library; S2022: Based on the preset clothing attribute tag set of clothing products, search in the tag word embedding matrix of the clothing matching tag library to obtain the retrieval tag word vector, input the retrieval tag word vector into the tag encoding model to generate the tag embedding feature of the clothing product. The tag encoding model is built based on the Transformer encoder. In this embodiment, the label encoding model includes a 6-layer encoder, and the multi-head attention mechanism is set to 8 heads, outputting 512-dimensional label embedding features; S2023: Input the clothing images of the same clothing product into a pre-trained visual feature extraction model to generate corresponding visual features. The visual feature extraction model is constructed based on the Residual Network (ResNet) algorithm. In this embodiment, the output of the last pooling layer of the visual feature extraction model is extracted and then reduced to 512 dimensions by a fully connected layer to obtain the visual features. S2024: Input the description text of the same clothing product into the pre-trained semantic feature extraction model to generate corresponding semantic features. The semantic feature extraction model is built based on the Bidirectional Encoder Representations from Transformers (BERT)-base model. In this embodiment, the output representation of the [CLS] token of the semantic feature extraction model is reduced to 512 dimensions through a fully connected layer to obtain the semantic features; S2025: Based on the InfoNCE loss function, cross-modal contrastive learning alignment is performed on label embedding features, visual features and semantic features to obtain trained and aligned label embedding features, trained and aligned visual features and trained and aligned semantic features. S2026: Based on the gating mechanism, the trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features are fused together and fully connected dimensionality reduction is performed to obtain the unified cross-modal embedding vector of the current clothing product. S2027: Traverse all clothing items in the clothing product library to obtain a unified cross-modal embedding vector set for clothing.
[0033] In one optional implementation, based on the InfoNCE loss function, cross-modal contrastive learning alignment is performed on the label embedding features, visual features, and semantic features to obtain the trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features, including: S20251: In the shared embedding space, construct a triple contrastive learning mechanism that includes visual-text, visual-label, and text-label, and construct the InfoNCE loss function; The formula for the InfoNCE loss function is as follows: In the formula, Loss due to InfoNCE; For visual-text loss, visual-label loss, and text-label loss; In the formula, For apparel product indexing; The total number of clothing items; For the first Label embedding features, visual features, and semantic features of clothing products; For the first Label embedding features, visual features, and semantic features of clothing products; It is an exponential function; This is the function for calculating cosine similarity. Temperature coefficient; S20252: Project the tag embedding features, visual features, and semantic features of the same clothing product onto a shared embedding space of the same dimension through independent linear mapping layers, and perform L2 normalization to obtain the corresponding projected tag embedding features, projected visual features, and projected semantic features. S20253: The projected label embedding features, projected visual features, and projected semantic features of any clothing item are used as the positive sample group, and the projected label embedding features, projected visual features, and projected semantic features of other clothing items are used as the negative sample group to obtain the contrastive learning alignment training sample set. S20254: Based on the InfoNCE loss function, input the contrastive learning aligned training sample set, use the backpropagation algorithm to optimize the parameters of the visual feature extraction model, semantic feature extraction model, label encoding model and linear mapping layer until the InfoNCE loss converges, and obtain the trained aligned visual feature extraction model, trained aligned semantic feature extraction model and trained aligned label encoding model. S20255: Using the trained and aligned visual feature extraction model, the trained and aligned semantic feature extraction model, and the trained and aligned label encoding model, the feature extraction steps for clothing products are repeated to obtain trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features.
[0034] In one optional implementation, the user's preference vector, scene constraint vector, and their conflict index are calculated to generate a scene compromise vector, which is then combined with the corresponding 3D body shape feature vector to generate a global state vector, including: S2031: Based on the user's historical interaction records (such as clicks, favorites, and add-to-cart), use the exponential decay function to calculate the user's preference weight for each clothing attribute tag in the clothing matching tag library, and then perform a weighted sum with the word embedding vector of the corresponding clothing attribute tag to obtain the user preference vector. S2032: Based on the user's real-time scene tags (e.g., "commuting to work", "10-20℃", "sunny day"), use a pre-trained mapping layer to perform mapping to obtain the corresponding scene constraint vector; S2033: Calculate the corresponding conflict index based on the user preference vector and the scenario constraint vector, using the following formula: In the formula, Conflict index; These are the user preference vector and the scenario constraint vector; This is the function for calculating cosine similarity. The preset conflict determination threshold is used; assuming the current user preference is "street sports style" and the scenario is "commuting to work," the cosine similarity is low, and a conflict is determined. >0, a conflict occurs; S2034: If the conflict index is greater than 0, a clothing conflict is determined to have occurred. A self-attention mechanism is introduced to calculate the compromise weight and generate a scene compromise vector based on the user preference vector and the scene constraint vector. The formula is as follows: In the formula, Compromise vectors for the scene; To compromise on weighting; These are the user preference vector and the scenario constraint vector; The preset feature space dimension; It is the transpose symbol; It is a normalized exponential function; S2035: Combine the scene compromise vector with the 3D body feature vector to obtain the global state vector; S2036: If the conflict index is less than or equal to 0, the user preference vector is used as the scene compromise vector and combined with the 3D body feature vector to obtain the global state vector.
[0035] In one optional implementation, the personalized intelligent recommendation model includes an upper garment recommendation agent, a lower garment recommendation agent, and an accessory recommendation agent, all built on the Actor-Critic architecture. The input states of the top clothing recommendation agent, bottom clothing recommendation agent, and accessory recommendation agent are all global state vectors, and the outputs are top clothing recommendation features, bottom clothing recommendation features, and accessory recommendation features, respectively, which are used to construct a combination of clothing features; The top recommendation agent is responsible for outputting "top recommendation features," specifically making decisions about tops, coats, and other upper-body clothing. It samples and retrieves only from the top subset within the unified cross-modal embedding vector set for clothing. The lower garment recommendation agent is responsible for outputting "lower garment recommendation features". It makes decisions specifically for lower body clothing such as pants and skirts, and samples and retrieves only from the lower garment subset in the unified cross-modal embedding vector set of clothing.
[0036] Accessory Recommendation Agent: Responsible for outputting "accessory recommendation features", making decisions specifically for decorative items such as shoes, bags, and jewelry, and sampling and retrieving only from the accessory subset in the unified cross-modal embedding vector set of clothing.
[0037] In one alternative implementation, the global state vector is input into the personalized intelligent recommendation model, and initial candidate outfit feature combinations are sampled and output from a unified cross-modal embedding vector set for clothing, including: S2041: Input the global state vector into the top clothing recommendation agent, bottom clothing recommendation agent and accessory recommendation agent of the personalized intelligent recommendation model respectively; S2042: Based on the global state vector, the top recommendation agent is used to sample and output the corresponding initial candidate top recommendation features from the unified cross-modal embedding vector set of clothing. S2043: Based on the global state vector, the lower garment recommendation agent samples and outputs the corresponding initial candidate lower garment recommendation features from the unified cross-modal embedding vector set of clothing. S2044: Based on the global state vector, use the accessory recommendation agent to sample and output the corresponding initial candidate accessory recommendation features from the unified cross-modal embedding vector set of clothing. S2045: Combine the initial candidate top recommendation features, initial candidate bottom recommendation features, and initial candidate accessory recommendation features to obtain the corresponding initial candidate outfit feature combination.
[0038] In one alternative implementation, based on the initial candidate outfit feature combinations, an improved artificial rabbit optimization algorithm is used for iterative optimization to generate the globally optimal outfit feature vector. This vector is then retrieved from a unified cross-modal embedding vector set for clothing, outputting a personalized intelligent recommendation list, including: S2051: Encode the outfit feature vector composed of the recommended features of the top, bottom and accessories into the position vector of the individual in the improved artificial rabbit optimization algorithm, and define the fitness function; The formula for the fitness function is: In the formula, To improve the individual rabbit optimization algorithm X The corresponding fitness value; To score visual coordination, a unified cross-modal embedding vector set is used to calculate the mean feature similarity between individual items within the outfit set, in order to evaluate the visual and semantic coordination of the outfit. To score body size visual balance, a pre-trained discriminative network (such as a multilayer perceptron) is used to group individuals... X The corresponding clothing features are fused with 3D body shape features to assess the visual balance of the clothing on the current user's body shape. To calculate the scene fit score, calculate the individual X The cosine similarity between the corresponding outfit aggregation features and the scene compromise vector is used to evaluate the degree to which outfits meet scene constraints and user preferences. For fitness weight coefficients, satisfying ; In the formula, For individuals X The total number of clothing items included in the corresponding initial candidate outfit feature combination; This is an index for individual items, which include tops, bottoms, and accessories. The first one obtained based on the unified cross-modal embedding vector set Item and the first A unified cross-modal embedding vector for each individual item; This is the function for calculating cosine similarity. In the formula, The Sigmoid activation function maps the discriminant network output to the (0,1) interval, representing the normalized score of body balance. It is a linear rectification activation function; The weight matrices and bias terms for the hidden and output layers of the pre-trained discriminant network; It is a 3D body shape feature vector; The aggregated features are obtained by average pooling of individual clothing items; The first one obtained based on the unified cross-modal embedding vector set A unified cross-modal embedding vector for each individual item; For individual item indexes; For individuals X The total number of clothing items included in the corresponding initial candidate outfit feature combination; This is the vector concatenation operator; In the formula, This is the function for calculating cosine similarity. The aggregated features are obtained by average pooling of individual clothing items; Compromise vectors for the scene; S2052: Based on the fitness function and the initial candidate outfit feature combination, the improved artificial rabbit optimization algorithm is used to iteratively optimize the outfit feature vector to obtain the globally optimal outfit feature vector; S2053: Based on the global optimal outfit feature vector, the K-Nearest Neighbors (KNN) algorithm is used to search the unified cross-modal embedding vector set of clothing to obtain several actual clothing product combinations whose cosine similarity with the global optimal outfit feature vector meets the preset threshold. These combinations are then output as a personalized intelligent recommendation list. In this embodiment, based on the globally optimal top recommendation feature, globally optimal bottom recommendation feature, and globally optimal accessory recommendation feature in the globally optimal outfit feature vector, the KNN algorithm is used to search in the corresponding top, bottom, and accessory subsets in the unified cross-modal embedding vector set of clothing. This yields several actual clothing items whose cosine similarity with the globally optimal top recommendation feature, globally optimal bottom recommendation feature, and globally optimal accessory recommendation feature respectively meets a preset threshold. These items are then combined into several combinations of actual clothing items and output as a personalized intelligent recommendation list.
[0039] In one alternative implementation, based on the fitness function and the initial candidate outfit feature combination, an improved artificial rabbit optimization algorithm is used to iteratively optimize the outfit feature vector to obtain the globally optimal outfit feature vector, including: S20521: Use the Logistic chaotic mapping iterative formula to generate a chaotic sequence, and map the chaotic sequence to the solution space of individuals in the improved artificial rabbit optimization algorithm to obtain an initial population including several initial individuals; The formula is: In the formula, For the first n+ 1. n There are several chaotic variables whose values range from [0,1]. This is the stability coefficient, typically 4; n Index for chaotic variables; In the formula, For the initial population, the first i An initial individual; For the first i One chaotic variable; To determine the upper and lower bounds of the solution space; i For individual indexes; t Index for iteration count; S20522: Based on the fitness function, calculate the fitness value of each individual in the initial population or the updated population of the previous iteration, and determine the top three individuals in fitness ranking, the second best individual, the third best individual, and the worst individual with the worst fitness based on the fitness value. S20523: In the initial population, replace the worst individual with the individual corresponding to the initial candidate clothing feature combination to obtain an optimized population; S20524: Calculate the attenuation energy factor. If the attenuation energy factor is greater than 1, proceed to the detour foraging process; otherwise, proceed to the random cave hiding process. The formula for the attenuation energy factor is: In the formula, For the first t The decay energy factor of the next iteration; t Index for iteration count; This represents the maximum number of iterations. A single random number between (0, 1) is used to control random fluctuations in energy. S20525: Initiate the detour foraging process, calculate the running length, introduce the gray wolf cooperation concept, calculate the guidance center based on the best individual, the second best individual, and the third best individual, and update the position of the optimized population or the population updated in the previous iteration based on the guidance center and the running length to obtain the updated population in the current iteration. The formula for the running distance is: In the formula, The running distance; The base of the exponent; t Index for iteration count; This represents the maximum number of iterations. A uniformly distributed random number between (0, 1); The formula for the guidance center is: In the formula, for t+ The guiding center for a single iteration; For the first t+ The first, second, and third potential movement vectors in one iteration; In the formula, For the first t The iteration of the ... i Each updated individual, in the initial iteration, In the optimized population, the first i An initial individual; For the first t+ 1, t The first, second, and third potential movement vectors of the next iteration; For the first t The next iteration Distance vectors between the best, second-best, and third-best individuals; This represents the vector of the first, second, and third control coefficients; These are the first, second, and third oscillation coefficients; A uniformly distributed random number between (0, 1); For control coefficients and oscillation coefficients; As a control factor;t This represents the current iteration number; For the first t The best, second-best, and third-best individuals in the next iteration; The formula for updating the location of the detour foraging route is: In the formula, for t+ The guiding center for a single iteration; For the first t+ 1 ,t The iteration of the ... i A newer individual; For running operators; The running distance; As a dimension mask vector, only a portion of dimensions are randomly selected for updating, realizing sparse dimension search and improving population diversity; A uniformly distributed random number between (0, 1); This is the rounding function; These are random numbers distributed according to a standard normal distribution between (0, 1). S20526: Initiate the random cave hiding process, calculate the running length, generate a one-dimensional cave, and based on the running length and the one-dimensional cave, update the position of the optimized population or the updated population of the previous iteration to obtain the updated population of the current iteration. The formula for the single-dimensional cave is: In the formula, For the first t The iteration of the ... i An updated, randomly selected, single-dimensional cave vector; For the first t The nonlinear energy hiding coefficient of the next iteration; A uniformly distributed random number between (0, 1); It is a one-dimensional mask vector, with only the target cave dimension being 1 and the other dimensions being 0; For the overall dimension; This is the index of the dimension value of the one-dimensional mask vector; This is the floor function operator; For the first t The iteration of the ... i A newer individual; The formula for updating the location of the random cave hiding place is: In the formula, For the first t+ 1 ,t The iteration of the ...i A newer individual; For running operators; The running distance; It is a dimensional mask vector; A uniformly distributed random number between (0, 1); S20527: If the optimal individual remains unchanged for several consecutive iterations, then Cauchy mutation is performed on the optimal individual, and the individual with the better fitness value is retained as the optimal individual. The formula is as follows: In the formula, For the first t+ Individuals after one iteration of Cauchy mutation and the optimal individual; This is the variable asynchronous length coefficient, used to control the disturbance amplitude; commonly used values are 0.01-0.1. These are standard Cauchy random numbers with a mean of 0 and a scale of 1. S20528: Repeatedly update the position of the population. When the current iteration count reaches the maximum iteration count or the fitness value of the best individual meets the requirements, terminate the iterative update of the population and output the final best individual. S20529: Decode the position vector of the optimal individual to obtain the global optimal outfit feature vector.
[0040] This invention also provides a personalized intelligent recommendation system 300 based on clothing matching tags, see below. Figure 3 The system may include the following units: Clothing and Outfit Tag Definition Unit 301 is used to define a clothing and outfit tag library that includes clothing attribute tags, user feature tags, and scene constraint tags. It reconstructs a 3D human body mesh from a 2D front and side view photo of the user, generates a body shape parameter vector, and encodes it into a 3D body shape feature vector. The feature extraction unit 302 is used to extract the tag embedding features, visual features and semantic features of each clothing item based on the clothing matching tag library, and perform cross-modal contrastive learning alignment through InfoNCE loss to generate a unified cross-modal embedding vector set for clothing. The global state generation unit 303 is used to calculate the user's user preference vector, scene constraint vector and their conflict index, generate a scene compromise vector, and generate a global state vector by combining the corresponding 3D body feature vector. The intelligent recommendation unit 304 is used to input the global state vector into the personalized intelligent recommendation model and sample and output the initial candidate outfit feature combination from the uniform cross-modal embedding vector set of clothing. The iterative optimization unit 305 is used to perform iterative optimization based on the initial candidate outfit feature combination using an improved artificial rabbit optimization algorithm, generate the globally optimal outfit feature vector, and retrieve it from the unified cross-modal embedding vector set of clothing to output a personalized intelligent recommendation list.
[0041] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. Memory, used to store computer programs; When a processor executes a program stored in memory, it implements the personalized intelligent recommendation method based on clothing matching tags of the present invention.
[0042] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. The communication interface is used for communication between the aforementioned terminal and other devices. The memory can include Random Access Memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory can also be at least one storage device located remotely from the aforementioned processor.
[0043] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0044] Furthermore, to achieve the above objectives, embodiments of the present invention also propose a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the personalized intelligent recommendation method based on clothing matching tags according to embodiments of the present invention.
[0045] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable hardware devices (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0046] The embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (apparatus), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0047] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0048] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0049] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. "And / or" indicates that either one or both can be chosen. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.
[0050] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A personalized intelligent recommendation method based on clothing matching tags, characterized in that, The method includes: Define a clothing tag library that includes clothing attribute tags, user feature tags, and scene constraint tags. Reconstruct a 3D human body mesh from a 2D front and side view photo of a user, generate a body shape parameter vector, and encode it as a 3D body shape feature vector. Based on the clothing tag library, the tag embedding features, visual features and semantic features of each clothing item are extracted, and cross-modal contrastive learning alignment is performed through InfoNCE loss to generate a unified cross-modal embedding vector set for clothing. Calculate the user's user preference vector, scene constraint vector and their conflict index, generate a scene compromise vector, and combine it with the corresponding 3D body feature vector to generate a global state vector; The global state vector is input into the personalized intelligent recommendation model, and the initial candidate outfit feature combination is sampled and output from the unified cross-modal embedding vector set of clothing. Based on the initial candidate outfit feature combinations, an improved artificial rabbit optimization algorithm is used for iterative optimization to generate the globally optimal outfit feature vector. The vector is then retrieved from a unified cross-modal embedding vector set for clothing to output a personalized intelligent recommendation list.
2. The personalized intelligent recommendation method based on clothing matching tags according to claim 1, characterized in that, Define a clothing outfit tag library including clothing attribute tags, user feature tags, and scene constraint tags. Reconstruct a 3D human body mesh from a 2D front and side view photo of the user, generate body shape parameter vectors, and encode them into 3D body shape feature vectors, including: Define a clothing outfit tag library that includes clothing attribute tags, user feature tags, and scene constraint tags, and construct the tag word embedding matrix of the clothing outfit tag library; Input a 2D photo of the user's front and side profile into a pre-trained parametric human body model, reconstruct a 3D human body mesh, and fit and output the corresponding body shape parameter vector. The body shape parameter vector is input into a pre-trained nonlinear mapping model to generate a normalized 3D body shape feature vector.
3. The personalized intelligent recommendation method based on clothing matching tags according to claim 2, characterized in that, Based on a clothing tag library, the tag embedding features, visual features, and semantic features of each clothing item are extracted. Cross-modal contrastive learning alignment is then performed using InfoNCE loss to generate a unified cross-modal embedding vector set for clothing, including: Call the clothing product library and extract the preset clothing attribute tag set, clothing image and descriptive text for each clothing product in the clothing product library; Based on the preset clothing attribute tag set of clothing products, the tag word embedding matrix of the clothing matching tag library is searched to obtain the retrieval tag word vector. The retrieval tag word vector is then input into the tag encoding model to generate the tag embedding feature of the clothing product. The images of the same clothing item are input into a pre-trained visual feature extraction model to generate corresponding visual features. The descriptive text of the same clothing product is input into a pre-trained semantic feature extraction model to generate corresponding semantic features; Based on the InfoNCE loss function, cross-modal contrastive learning alignment is performed on label embedding features, visual features, and semantic features to obtain trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features. Based on the gating mechanism, the trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features are fused together and fully connected dimensionality reduction is performed to obtain the unified cross-modal embedding vector of the current clothing product. Traverse all clothing items in the clothing product library to obtain a unified cross-modal embedding vector set for clothing.
4. The personalized intelligent recommendation method based on clothing matching tags according to claim 3, characterized in that, Based on the InfoNCE loss function, cross-modal contrastive learning alignment is performed on label embedding features, visual features, and semantic features to obtain trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features, including: In the shared embedding space, a triple contrastive learning mechanism including visual-text, visual-label, and text-label is constructed, and the InfoNCE loss function is built. The tag embedding features, visual features, and semantic features of the same clothing product are projected onto a shared embedding space of the same dimension through independent linear mapping layers, and then L2 normalization is performed to obtain the corresponding projected tag embedding features, projected visual features, and projected semantic features. The projected label embedding features, projected visual features, and projected semantic features of any clothing item are used as the positive sample group, while the projected label embedding features, projected visual features, and projected semantic features of other clothing items are used as the negative sample group to obtain the contrastive learning alignment training sample set. Based on the InfoNCE loss function, the contrastive learning aligned training sample set is input, and the backpropagation algorithm is used to optimize the parameters of the visual feature extraction model, semantic feature extraction model, label encoding model and linear mapping layer until the InfoNCE loss converges, resulting in the trained aligned visual feature extraction model, trained aligned semantic feature extraction model and trained aligned label encoding model. Using the trained and aligned visual feature extraction model, the trained and aligned semantic feature extraction model, and the trained and aligned label encoding model, the feature extraction steps for clothing products are repeated to obtain trained and aligned label embedding features, trained and aligned visual features, and trained and aligned semantic features.
5. The personalized intelligent recommendation method based on clothing matching tags according to claim 4, characterized in that, Calculate the user's preference vector, scene constraint vector, and their conflict index to generate a scene compromise vector. Combine this with the corresponding 3D body shape feature vector to generate a global state vector, including: Based on the user's historical interaction records, the user's preference weight for each clothing attribute tag in the clothing matching tag library is calculated using the exponential decay function. The preference weight is then weighted and summed with the word embedding vector of the corresponding clothing attribute tag to obtain the user preference vector. Based on the user's real-time scene labels, a pre-trained mapping layer is used to perform mapping to obtain the corresponding scene constraint vector; Calculate the corresponding conflict index based on the user preference vector and the scenario constraint vector; If the conflict index is greater than 0, a clothing matching conflict is determined to have occurred. A self-attention mechanism is introduced to calculate the compromise weight and generate a scenario compromise vector based on the user preference vector and the scenario constraint vector. The scene compromise vector is combined with the 3D body feature vector to obtain the global state vector; If the conflict index is less than or equal to 0, the user preference vector is used as the scene compromise vector and combined with the 3D body feature vector to obtain the global state vector.
6. The personalized intelligent recommendation method based on clothing matching tags according to claim 5, characterized in that, The personalized intelligent recommendation model includes an intelligent agent for recommending upper garments, an intelligent agent for recommending lower garments, and an intelligent agent for recommending accessories, all of which are built based on the Actor-Critic architecture. The input states of the top clothing recommendation agent, bottom clothing recommendation agent, and accessory recommendation agent are all global state vectors, and their outputs are top clothing recommendation features, bottom clothing recommendation features, and accessory recommendation features, respectively, which are used to construct a combination of clothing features.
7. The personalized intelligent recommendation method based on clothing matching tags according to claim 6, characterized in that, The global state vector is input into the personalized intelligent recommendation model, and the initial candidate outfit feature combination is output from the unified cross-modal embedding vector set of clothing, including: The global state vector is input into the top clothing recommendation agent, bottom clothing recommendation agent, and accessory recommendation agent of the personalized intelligent recommendation model, respectively. Based on the global state vector, the top recommendation agent samples and outputs the corresponding initial candidate top recommendation features from the unified cross-modal embedding vector set of clothing. Based on the global state vector, the lower garment recommendation agent samples and outputs the corresponding initial candidate lower garment recommendation features from the unified cross-modal embedding vector set of clothing. Based on the global state vector, the accessory recommendation agent samples and outputs the corresponding initial candidate accessory recommendation features from the unified cross-modal embedding vector set of clothing. The initial candidate top recommendation features, initial candidate bottom recommendation features, and initial candidate accessory recommendation features are combined to obtain the corresponding initial candidate outfit feature combinations.
8. The personalized intelligent recommendation method based on clothing matching tags according to claim 7, characterized in that, Based on the initial candidate outfit feature combinations, an improved artificial rabbit optimization algorithm is used for iterative optimization to generate the globally optimal outfit feature vector. This vector is then retrieved from a unified cross-modal embedding vector set for clothing, outputting a personalized intelligent recommendation list, including: The outfit feature vector, composed of the recommended features of tops, bottoms, and accessories, is encoded into the position vector of an individual in the improved artificial rabbit optimization algorithm, and a fitness function is defined. Based on the fitness function and the initial candidate outfit feature combination, the improved artificial rabbit optimization algorithm is used to iteratively optimize the outfit feature vector to obtain the globally optimal outfit feature vector. Based on the globally optimal outfit feature vector, the KNN algorithm is used to search the unified cross-modal embedding vector set of clothing to obtain several actual clothing product combinations whose cosine similarity with the globally optimal outfit feature vector meets a preset threshold. These combinations are then output as a personalized intelligent recommendation list.
9. The personalized intelligent recommendation method based on clothing matching tags according to claim 8, characterized in that, Based on the fitness function and the initial candidate outfit feature combination, an improved artificial rabbit optimization algorithm is used to iteratively optimize the outfit feature vector to obtain the globally optimal outfit feature vector, including: The chaotic sequence is generated using the Logistic chaotic mapping iterative formula, and then mapped to the solution space of individuals in the improved artificial rabbit optimization algorithm to obtain an initial population including several initial individuals. Based on the fitness function, calculate the fitness value of each individual in the initial population or the updated population of the previous iteration, and determine the top three individuals in fitness ranking, the second best individual, the third best individual, and the worst individual with the worst fitness based on the fitness value. In the initial population, the worst individual is replaced by the individual corresponding to the initial candidate clothing feature combination, resulting in an optimized population. Calculate the attenuation energy factor. If the attenuation energy factor is greater than 1, proceed to the detour foraging process; otherwise, proceed to the random cave hiding process. The process of foraging by detour is initiated, the running distance is calculated, the cooperative concept of gray wolves is introduced, and the guidance center is calculated based on the best, second best, and third best individuals. Based on the guidance center and the running distance, the position of the optimized population or the population updated in the previous iteration is updated to obtain the population updated in the current iteration. Initiate a random cave hiding process, calculate the running length, generate a one-dimensional cave, and based on the running length and the one-dimensional cave, update the position of the optimized population or the updated population of the previous iteration to obtain the updated population of the current iteration. If the optimal individual remains unchanged for several consecutive iterations, then Cauchy mutation is performed on the optimal individual, and the individual with the better fitness value is retained as the optimal individual. Repeatedly update the position of the population. When the current iteration reaches the maximum number of iterations or the fitness value of the best individual meets the requirements, terminate the iterative update of the population and output the final best individual. Decode the position vector of the optimal individual to obtain the global optimal outfit feature vector.
10. A personalized intelligent recommendation system based on clothing matching tags, used to implement the personalized intelligent recommendation method as described in any one of claims 1-9, characterized in that, The system includes: The clothing and outfit tag definition unit is used to define a clothing and outfit tag library that includes clothing attribute tags, user feature tags, and scene constraint tags. It reconstructs a 3D human body mesh from a 2D front and side view photo of a user, generates a body shape parameter vector, and encodes it into a 3D body shape feature vector. The feature extraction unit is used to extract the tag embedding features, visual features and semantic features of each clothing item based on the clothing matching tag library, and to perform cross-modal contrastive learning alignment through InfoNCE loss to generate a unified cross-modal embedding vector set for clothing. The global state generation unit is used to calculate the user's user preference vector, scene constraint vector and their conflict index, generate the scene compromise vector, and combine it with the corresponding 3D body feature vector to generate the global state vector. The intelligent recommendation unit is used to input the global state vector into the personalized intelligent recommendation model and sample and output the initial candidate outfit feature combination from the uniform cross-modal embedding vector set of clothing. The iterative optimization unit is used to iteratively optimize based on the initial candidate outfit feature combination using an improved artificial rabbit optimization algorithm, generate the globally optimal outfit feature vector, and retrieve it from the unified cross-modal embedding vector set of clothing to output a personalized intelligent recommendation list.