Multimodal and multitask skin care recommendation method based on VLSNR+MMOE combined fusion structure

Through the multimodal and multi-task skin care recommendation method with the VLSNR+MMOE combined fusion structure, the subjectivity, data silos, lack of personalization and real-time problems of the existing skin care product recommendation system are solved, and accurate user skin feature recognition and personalized recommendations are achieved, which improves the efficiency of the recommendation system and user satisfaction.

CN118839041BActive Publication Date: 2025-09-30SOUTH CHINA UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410903698.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2025-09-30
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

Existing skin care product recommendation systems have problems such as strong subjectivity, data silos, insufficient personalization, high computational complexity, and lack of real-time performance, making it difficult to accurately identify user skin characteristics and provide personalized recommendations.

Method used

A multimodal and multi-task skin care recommendation method based on the VLSNR+MMOE combined fusion structure is adopted. The skin image features are extracted through the CNN model to construct a skin care product knowledge graph. The factorization machine and multimodal attention encoder are combined to generate skin care product vectors and user interest vectors. The multi-head attention mechanism and GRU model are used to process time series features. Finally, the CTR, CVR and user satisfaction prediction results are generated through the MMoE model.

Benefits of technology

It achieves more accurate user demand analysis, improves the recall efficiency and accuracy of personalized skin care product recommendations, and enhances the real-time performance and user experience of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118839041B_ABST
    Figure CN118839041B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal, multitask skin care recommendation method based on a VLSNR+MMOE combined fusion structure, comprising the following steps: extracting skin features from skin images; constructing a skin care product knowledge graph; inputting the skin features of the skin images and skin care product features into a factor decomposition machine for training to generate a Top K candidate skin care product set; encoding using a multimodal attention encoder based on a VLSNR model; calculating the relationship between features at different positions based on a multi-head attention mechanism, and outputting a time series feature vector based on a GRU model; splicing the time series feature vector, a product image and text feature fusion vector, a user image and text feature fusion vector, a personalized feature vector, and a context feature vector into an MMoE model, and outputting CTR, CVR, and user satisfaction prediction results. The present invention can accurately identify user skin features and, in combination with a skin care product knowledge base, make personalized recommendations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data matching recommendation, and in particular to a multimodal multitasking skin care recommendation method based on a VLSNR+MMOE combined fusion structure. Background Art

[0002] Traditional skincare product recommendation systems primarily rely on the experience of offline beauty consultants or user-completed questionnaires. While these methods can provide personalized recommendations to a certain extent, they have significant limitations. To address these issues, data-driven skincare product recommendation systems have emerged in recent years. These systems leverage user data, machine learning, and image processing techniques to provide users with more accurate and personalized skincare product recommendations. Currently, collaborative filtering or content-based recommendation systems can capture users' implicit preferences by analyzing their purchase history and behavioral data, recommending skincare products favored by other users. However, these methods generally suffer from the following problems: (a) Subjectivity and accuracy. Traditional questionnaires and rule-based systems rely on user self-reports and expert knowledge, which are highly subjective. Users' understanding of their own skin problems may not be accurate, while expert knowledge bases have limited coverage and are unable to cope with diverse and dynamically changing user needs. (b) Data silos and cold start problems. Collaborative filtering systems rely on a large amount of user behavior data, but their recommendation performance is poor for new users or when data is sparse. Such systems often lack sufficient data support when facing new users or new products, resulting in reduced accuracy and reliability of recommendation results. (c) Lack of personalization. Early recommendation systems struggled to fully capture and understand users' personalized needs, especially specific skin problems and preferences. Factors such as skin type, skin problems, and lifestyle habits have a significant impact on skincare product selection, but early systems were unable to fully integrate this information for recommendation. (d) Computational complexity and scalability issues. Although content recommendation systems can provide recommendations based on skincare product ingredients and efficacy, they are computationally expensive and difficult to scale when faced with large-scale data. Such systems need to process a large amount of skincare product information and user characteristics, which significantly increases computational complexity and storage requirements, affecting the system's response speed and user experience. (e) Lack of real-time and dynamic adjustment. Early systems generally lacked the ability to monitor and dynamically adjust the user's skin condition in real time. The user's skin condition may change with seasonal changes and changes in lifestyle habits. Early systems found it difficult to capture these changes in a timely manner and adjust the recommendation strategy, resulting in a decrease in the timeliness and applicability of the recommendation results. Summary of the Invention

[0003] In order to overcome the defects and shortcomings of the existing technology, the present invention provides a multimodal and multi-task skin care recommendation method based on the VLSNR+MMOE combined fusion structure. The present invention can solve the problems of strong subjectivity, data silos, insufficient personalization, high computational complexity, and lack of real-time performance in the existing skin care product recommendation system. It can accurately identify the user's skin characteristics and make personalized recommendations in combination with the skin care product knowledge base.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] The present invention provides a multi-modal multi-task skin care recommendation method based on a VLSNR+MMOE combined fusion structure, comprising the following steps:

[0006] Extract skin features of skin images based on CNN model;

[0007] Crawl skin care product information, parse skin care product instructions and user reviews, convert them into RDF triples, define the skin care product ontology model, and build a skin care product knowledge graph;

[0008] Extract skin care product features from the skin care product knowledge graph, input the skin features of the skin image and the skin care product features into the factorization machine for training, generate skin care product vectors and user interest vectors based on the trained parameters of the factorization machine, and calculate the cosine similarity between the two to generate the top K candidate skin care product set;

[0009] The multimodal attention encoder based on the VLSNR model is used to encode and generate the user history behavior sequence vector, the image and text feature vectors of the top K candidate skin care products, the personalized feature vector, the context feature vector, the product image and text feature fusion vector, and the user image and text feature fusion vector;

[0010] Based on a multi-layer perceptron, a nonlinear transformation is performed on the feature vector formed by fusing the user's historical behavior sequence vector and the image and text feature vectors of the top K candidate skincare products to extract high-level features. The relationship between features at different positions is calculated using a multi-head attention mechanism, and the time series feature vector is output based on the GRU model.

[0011] The time series feature vector, product image and text feature fusion vector, user image and text feature fusion vector, personalized feature vector, and context feature vector are concatenated and input into the MMoE model to output the prediction results of CTR, CVR, and user satisfaction.

[0012] As a preferred technical solution, skin features of skin images are extracted based on a CNN model, specifically including:

[0013] Acquire skin images and perform image preprocessing;

[0014] The preprocessed skin image is passed through the multi-layer convolutional layer and pooling layer of the CNN model to extract low-level to high-level features of the image. The convolutional layer uses the ReLU activation function and expands the feature map output by the convolutional layer into a one-dimensional feature vector through the fully connected layer or the GAP layer.

[0015] As a preferred technical solution, a knowledge graph of skin care products is constructed, specifically including:

[0016] Crawl skin care product information;

[0017] Analyze skin care product instructions and user reviews based on NLP technology;

[0018] Define the skin care product ontology model and create an OWL ontology;

[0019] The skin care product instructions and user review information are converted into RDF triples, and the RDF triples are stored in the Neo4j or AllegroGraph graph database to construct a skin care product knowledge graph.

[0020] As a preferred technical solution, skin care product vectors and user interest vectors are generated based on the trained parameters of the factorization machine and the cosine similarity between the two is calculated to generate the Top K candidate skin care product set, specifically including:

[0021] After factorization machine training, the embedding vectors of each feature of skin features and skin care product features, as well as the assigned weights of each feature, are obtained;

[0022] The feature embedding vectors of all skin care product subsets are accumulated to form the representation vector of each skin care product, which is expressed as:

[0023]

[0024] Among them, P i represents the representation vector of the i-th skin care product, w productj is the weight of the jth skin care product feature, E productj is the embedding vector of the j-th skin care product feature;

[0025] For skin care product feature vector P i Perform k-NN indexing, generate index tables and store them in the Fasis database;

[0026] The skin features of the skin image input by the user are extracted and weighted summed with the corresponding distribution weights obtained after factor decomposition machine training to obtain the current user's interest vector U, which is expressed as:

[0027]

[0028] Among them, w userjis the weight of the j-th skin feature, E skinj is the embedding vector of the j-th skin feature;

[0029] Calculate the user interest vector U and each skin care product vector P stored in the Fasis database i The similarity scores are sorted from high to low to generate the Top K candidate skin care product set.

[0030] As a preferred technical solution, a multimodal attention encoder based on the VLSNR model is used for encoding, specifically including:

[0031] Obtain historical user interaction data on skincare products, encode it based on the Clip-Crossmodel Attention module of the multimodal attention encoder, and generate a user historical behavior sequence vector;

[0032] Obtain multimodal information of the Top K candidate skincare products and generate image and text feature vectors of the Top K candidate skincare products based on the Clip-Crossmodel Attention module encoding;

[0033] Based on the Clip-Crossmodel Attention module, the user's preferences and product attributes are encoded to generate personalized feature vectors;

[0034] Based on the Clip-Crossmodel Attention module, the context information is encoded to generate a context feature vector;

[0035] Based on the Clip-Crossmodel Attention module, the skin care product image and text features crawled from the website and the skin features extracted from the user input skin image are encoded to generate the product image and text feature fusion vector and the user image and text feature fusion vector.

[0036] As a preferred technical solution, the time series feature vector, product image and text feature fusion vector, user image and text feature fusion vector, personalized feature vector, and context feature vector are spliced ​​and input into the MMoE model to output the prediction results of CTR, CVR, and user satisfaction, specifically including:

[0037] After vector splicing, the vectors are input into multiple expert networks to extract feature relationships. After vector splicing, they are passed to the gating network of each task. Each gating network generates the weight of each expert network through the gating mechanism according to the input features. The gating network performs weighted summation on the output of the expert network based on the calculated weights to obtain the feature representation of each task, which is specifically expressed as:

[0038]

[0039] Among them, T task represents the task feature representation, α i represents the weight of the i-th expert network, E i represents the output of the i-th expert network, and N represents the number of expert networks;

[0040] The task feature representation is passed to the corresponding task layer to calculate and generate the prediction results of CTR, CVR and user satisfaction.

[0041] As a preferred technical solution, the task feature representation is transferred to the corresponding task layer to calculate the multi-task loss function, which is specifically expressed as:

[0042] L=αL CTR +βL CVR +γL Satisfaction

[0043] Among them, L is the total loss function, α is the weight coefficient of the CTR task, and L CTR is the click-through rate CTR loss, β is the weight coefficient of the CVR task, L CVR is the conversion rate CVR loss, γ is the weight coefficient of the user satisfaction task, L Satisfaction It is the loss of user satisfaction;

[0044] L CTR and L CVR The error of click-through rate prediction is calculated using the cross entropy loss function, which is expressed as:

[0045]

[0046] Among them, y CTR is the actual click rate label, with a value of 0 or 1. is the click rate value predicted by the model, with a value range of [0,1], y CVR It is the actual conversion rate label, with a value of 0 or 1. is the conversion rate value predicted by the model, ranging from [0, 1];

[0047] L Satisfaction The error of user satisfaction prediction is calculated using the mean square error (MSE) loss function, which is expressed as:

[0048]

[0049] Among them, y Satisfaction is the actual user satisfaction label, is the user satisfaction value predicted by the model.

[0050] In order to achieve the above second purpose, the present invention adopts the following technical solutions:

[0051] A multimodal multitask skin care recommendation system based on a VLSNR+MMOE combined fusion structure is used to implement the multimodal multitask skin care recommendation method based on the VLSNR+MMOE combined fusion structure. The system includes: a skin feature extraction module, a skin care product knowledge graph construction module, a candidate skin care product set construction module, a multimodal attention encoding module, a multi-head attention calculation module, a time series feature vector construction module, a vector splicing module, and a prediction result generation module;

[0052] The skin feature extraction module is used to extract skin features of the skin image based on the CNN model;

[0053] The skin care product knowledge graph construction module is used to crawl skin care product information, parse skin care product instructions and user comments, and convert them into RDF triples, define the skin care product ontology model, and construct the skin care product knowledge graph;

[0054] The candidate skin care product set construction module is used to extract skin care product features from the skin care product knowledge graph, input the skin features and skin care product features of the skin image into a factor decomposition machine for training, generate skin care product vectors and user interest vectors based on the trained parameters of the factor decomposition machine, and calculate the cosine similarity between the two to generate a Top K candidate skin care product set;

[0055] The multimodal attention encoding module is used to encode based on the multimodal attention encoder of the VLSNR model to generate a user historical behavior sequence vector, a picture and text feature vector of the top K candidate skin care products, a personalized feature vector, a context feature vector, a product picture and text feature fusion vector, and a user picture and text feature fusion vector;

[0056] The multi-head attention calculation module is used to perform nonlinear transformation on the feature vector obtained by fusing the user's historical behavior sequence vector and the image and text feature vectors of the TopK candidate skin care products based on a multi-layer perceptron, extract high-level features, and calculate the relationship between features at different positions based on the multi-head attention mechanism;

[0057] The time series feature vector construction module is used to output the time series feature vector based on the GRU model;

[0058] The vector splicing module is used to splice the time series feature vector, the product image and text feature fusion vector, the user image and text feature fusion vector, the personalized feature vector, and the context feature vector;

[0059] The prediction result generation module is used to input the spliced ​​vectors into the MMoE model and output the prediction results of CTR, CVR and user satisfaction.

[0060] In order to achieve the third purpose above, the present invention adopts the following technical solutions:

[0061] A storage medium stores a program, which, when executed by a processor, implements the multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure as described above.

[0062] In order to achieve the fourth purpose, the present invention adopts the following technical solutions:

[0063] A computing device includes a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, it implements the multimodal and multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure as described above.

[0064] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0065] (1) In response to the problem of limited multimodal data fusion capabilities, the present invention adopts a multimodal data fusion technology solution to solve the technical problems of processing and fusing data from multiple sources, including skin images uploaded by users, skin care product information, user behavior data, etc., achieving a more comprehensive and accurate analysis of user needs and providing personalized skin care product recommendations.

[0066] (2) To address the problem of insufficient recall efficiency and accuracy, the present invention adopts a factorization machine (FM) and cosine similarity calculation, combined with the KNN indexing technology of the Fasis database, to solve the problem of efficiency and accuracy in generating candidate skin care product sets, thereby achieving the technical effect of improving the efficiency and accuracy of the recall stage and shortening the recommendation response time.

[0067] (3) To address the problem of poor deep feature extraction and fusion effects, the present invention utilizes the MultiModelAttention Encoder technical solution to solve the problem of deep encoding and fusion of user historical behavior sequences and skin care product graphic features, achieving the technical effect of capturing complex feature relationships, improving feature representation quality, and enhancing model prediction capabilities.

[0068] (4) In order to solve the problem of insufficient processing of time series features, the present invention uses the technical solution of processing time series features through the GRU model to solve the problem of time dependency between user behavior and skin care product candidate sets, and achieves the technical effect of better understanding changes in user behavior and providing time-sensitive recommendation results.

[0069] (5) In response to the problem of insufficient personalization and contextual awareness, the present invention combines the technical solution of user personalized characteristics and contextual information (such as season, climate, time, etc.), solves the problem of personalized and contextually aware recommendations, and achieves the technical effect of providing skin care products that are more in line with the user's current needs, enhancing user experience, and increasing user satisfaction and trust in the recommendation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 Schematic diagram of the process of the multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure of the present invention;

[0071] Figure 2 Schematic diagram of the overall implementation architecture of the multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure of the present invention;

[0072] Figure 3 A schematic diagram of a process for generating a candidate skin care product set based on a factorization machine and cosine similarity calculation according to the present invention;

[0073] Figure 4 A schematic diagram of the process of performing multimodal fusion encoding and generating time series feature vectors in the present invention;

[0074] Figure 5 This is a flow chart of the present invention for generating prediction results of CTR, CVR, and user satisfaction based on MMoE through three expert networks and a gating mechanism. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0076] Example 1

[0077] like Figure 1 、 Figure 2 As shown, this embodiment provides a multi-modal multi-task skin care recommendation method based on a VLSNR+MMOE combined fusion structure, comprising the following steps:

[0078] S1: Extract skin features of skin images based on CNN model;

[0079] Generally speaking, skin images uploaded by users may contain problems such as noise and uneven lighting, which directly affect the accuracy of feature extraction. To improve the accuracy of feature extraction, the image is first preprocessed. Median filtering or bilateral filtering techniques are used to reduce the noise in the image, and adaptive histogram equalization (CLAHE) is used to balance the uneven lighting in the image.

[0080] These preprocessing steps are performed offline, taking into account computational efficiency and processing quality requirements, thus ensuring high-quality images fed into the CNN model. After preprocessing, a pretrained CNN model (such as VGG16, ResNet50, or InceptionV3) is used to extract the characteristic categories of the skin image. Through transfer learning, the model is fine-tuned to suit skin feature extraction. These models can efficiently extract important features in the image, such as hue, glossiness, and pore distribution. The specific steps are as follows:

[0081] S11: User uploads skin image via mobile device or computer;

[0082] S12: After receiving the image uploaded by the user, preprocessing is first performed: median filtering or bilateral filtering technology is used to reduce noise in the image, and adaptive histogram equalization (CLAHE) is used to balance the uneven illumination in the image;

[0083] S13: Input the pre-processed image generated in step S12 into the CNN model;

[0084] S14: Extract low-level to high-level features of the image through multiple convolutional layers and pooling layers. The convolutional layer uses the ReLU activation function.

[0085] S15: Expand the feature map output by the convolutional layer into a one-dimensional feature vector through a fully connected layer or a GAP (Global Average Pooling) layer.

[0086] The preprocessing method of this embodiment can significantly improve the quality of the input image, thereby improving the accuracy of subsequent feature extraction; at the same time, the use of the pre-trained model can quickly and accurately extract high-quality features, ensuring the basic data quality of the recommendation system.

[0087] This embodiment refers to a variety of professional websites and academic articles to understand the skin features that can be extracted and detected. These reference materials include "Dingxiang Doctor", "Chinese Medical Association Dermatology Branch Guidelines", "Journal of Dermatology and Venereology", "Journal of Dermatological Science", "The British Journal of Dermatology" and other authoritative journals, as well as professional skin care websites such as "WebMD", "Healthline" and "Mayo Clinic". Through these materials, 50 skin features that can be extracted are determined. These features include skin color: hue, saturation, brightness, skin gloss: reflectivity, pore size and distribution, number and distribution of blackheads, number and distribution of acne, wrinkle depth and distribution, number and distribution of spots, number and distribution of red blood vessels, skin texture: roughness, smoothness, oil secretion, moisture content, skin elasticity, skin uniformity, epidermal thickness, epidermal hydration status, skin surface temperature, capillary distribution, skin microcirculation status, inflammatory markers, skin cell density, ultraviolet damage marks, skin pH value, sebum secretion rate, acne scar condition, skin barrier function, dermis thickness, collagen content, elastin content, glycosaminoglycan content, skin metabolism rate, skin sensitivity, acne type, degree of redness and swelling, degree of dryness, pigmentation, blood circulation, lipid layer integrity, cell activity, cell renewal rate, stratum corneum status, UV reflectivity, UV absorption rate, number of aging spots, number and distribution of acne, skin firmness, epidermal water evaporation rate, skin redness, skin softness, skin translucency, skin dullness, etc.

[0088] S2: Constructing a skincare product knowledge graph based on RDF and OWL technologies: Skincare product information is crawled from skincare product e-commerce platforms and official websites. Natural language processing (NLP) technology is used to parse skincare product instructions and user reviews. The extracted skincare product information and product review information are standardized and structured, and converted into RDF triples. Multiple RDF triples constitute the skincare product knowledge graph. Considering the consistency and scalability requirements of the data, this data processing and storage work is completed in an offline environment. The specific steps are as follows:

[0089] S21: Collect detailed information about skin care products using web crawler technology, including ingredients, efficacy, applicable skin types, user reviews, etc. This embodiment crawls skin care product information from skin care product e-commerce platforms and official websites;

[0090] S22: Use NLP technology to analyze skin care product instructions and user reviews to extract valuable information;

[0091] S23: Define the skin care product ontology model, including skin care product categories, ingredients, efficacy, etc., and create an OWL ontology using tools such as Protégé;

[0092] S24: Convert the skin care product information obtained in step S22 into an RDF triple, such as: <skin care product A><containing ingredients><ingredient B>;

[0093] S25: Generate and manage RDF triple data using Jena or RDFLib, store the RDF triples obtained in step S22 in a Neo4j or AllegroGraph graph database, and use SPARQL query language to retrieve and analyze skin care product information;

[0094] Through the construction of a knowledge graph, this embodiment can systematically organize and manage skin care product information, improve data consistency and scalability, and enhance the intelligence level of the recommendation system.

[0095] The detailed information dataset of skin care products used in this embodiment comes from skin care product e-commerce platforms, brand official websites, user reviews, and professional beauty websites, such as Tmall, JD.com, Little Red Book (RED), Beauty Practice, Dingxiang Doctor, Taobao, Suning, Vip.com, official websites of skin care brands (such as Lancôme, Estée Lauder, Shiseido, etc.), user review websites (such as Douban and Zhihu), and professional beauty websites (such as YOKA Fashion Network and Meizhuang Xindi). Through web crawler technology, comprehensive information about skin care products can be obtained, including but not limited to the following 50 features: product name, brand name, product type, applicable skin type, main ingredients, auxiliary ingredients, efficacy, usage method, frequency of use, product packaging specifications, applicable population, applicable age group, production date, shelf life, product price, product origin, manufacturer, product color, product smell, product texture, product pH value, sun protection index, product certification, product series, product label, user evaluation score, number of user reviews, whether the product contains alcohol, preservatives, fragrances, whether it has been allergy tested, whether it is suitable for pregnant women, whether it has been tested on animals, whether it contains hormones, mineral oil, silicone oil, dyes, heavy metals, skin condition feedback after use, product permeability, absorption speed, moisturizing degree, freshness, stability, safety, antioxidant properties, antibacterial properties, anti-inflammatory properties, repair properties and additional functions, etc.

[0096] S3: Generate the corresponding candidate set based on the factorization machine (FM) and cosine similarity calculation. Figure 3 The specific steps are as follows:

[0097] S31: Using the CNN model pre-trained in step S1, online calculation is performed to extract skin features of the user's skin image;

[0098] S32: Extracting skin care product features from the skin care product knowledge graph in step S2;

[0099] S33: Input the skin features and skincare product features obtained in the first two steps into a factorization machine (FM) model for training to obtain the embedding vector of each feature of the skin features and skincare product features, as well as the assigned weight of each feature, that is, to obtain the parameters of the FM model;

[0100] S34: For each skin care product, use the feature vector and weight calculated by the FM model, first calculate offline and then accumulate the feature embedding vectors of all skin care product subsets to form the representation vector of each skin care product Among them, P i represents the representation vector of the i-th skin care product, w productj is the weight of the jth skin care product feature, E productj is the embedding vector of the j-th skin care product feature.

[0101] S35: The skin care product feature vector P generated in step S34 i Perform k-NN indexing and generate index tables through offline calculation;

[0102] S36: storing the generated index table in the Fasis database for quick retrieval;

[0103] S37: Use the CNN model pre-trained in step S1 to extract the skin feature information input by the user, and perform a weighted summation with the parameters of the FM model trained in step S33 to obtain the current user's interest vector U. The specific formula is: Among them, U represents the user's interest vector, w userj is the weight of the j-th skin feature, E skinj is the embedding vector of the jth skin feature. Use the k-nearest neighbor (k-NN) algorithm to index in the Fasis database and quickly retrieve similar skin care products;

[0104] S38: Using cosine similarity, online calculation of the user interest vector U and each skin care product vector P stored in the Fasis database i The similarity is calculated as follows:

[0105]

[0106] S39: Sort the similarity scores from high to low to generate the Top K candidate skin care product set as the recall result.

[0107] In this embodiment, the generated user interest vector U is stored in MongoDB as a record of the user's historical skin characteristics for subsequent recommendation optimization and personalized adjustment;

[0108] In the existing recall stage, traditional methods often take a long time to search in a massive database of skin care products, resulting in low efficiency of the recall process; due to the limitations of feature extraction and matching algorithms, the accuracy of the recall results is often not high; traditional recall methods do not take into account the personalized needs of users and cannot be based on the specific skin characteristics of users. The present invention builds a recall model based on factor decomposition machine (FM), which can efficiently calculate the feature vector of skin care products and improve the calculation efficiency of recall. The user interest vector U and the skin care product vector P are calculated by cosine similarity. i The similarity between them can accurately reflect the actual needs of users and improve the accuracy of recall results. It can also generate personalized user interest vectors U based on the skin images uploaded by users to achieve personalized recommendations and meet the specific needs of users.

[0109] S4: Based on the Variable Length Sequential Neural Recommender (VLSNR) model, this model is a recommendation system model that processes user behavior sequences and candidate sets. It aims to utilize user historical behavior information and the characteristics of candidate items to provide efficient personalized recommendations. By introducing a multimodal attention encoder (MultiModelAttentionEncoder) to encode user historical behavior sequences and candidate sets, it can better capture the complex feature relationships of user behavior and improve the accuracy of recommendations.

[0110] like Figure 4 The specific steps are as follows:

[0111] S41: Collect historical user interaction data on specific skin care products, such as clicks, favorites, and adding to shopping carts;

[0112] S42: Use the Clip-Crossmodel Attention module in the MultiModel Attention Encoder to encode the user historical behavior offline calculation in step S41 to generate a user historical behavior sequence vector;

[0113] S43: The top K candidate skincare products to be recommended are obtained from the recall phase in step S3, including multimodal information of the skincare products, such as images and text descriptions of the skincare products. These information is input into the Clip-Crossmodel Attention module for encoding, and multimodal feature vectors are generated by offline calculation, i.e., the image-text feature vectors of the top K candidate skincare products.

[0114] S44: Combine user preferences and product attributes to generate personalized features, such as the user's sensitivity to specific components, perform embedding vectorization, and generate personalized feature vectors through offline calculations.

[0115] S45: embedding and vectorizing context information such as the current season, climate conditions, and time (morning or evening), and generating a context feature vector through offline calculation;

[0116] S46: Input the skin care product image and text features crawled from the website and the user's skin features extracted by CNN into the MultiModel Attention Encoder for fusion embedding to obtain the skin care product image and text feature fusion vector (i.e., the product image and text feature fusion vector) and the user image and text feature fusion vector;

[0117] S47: Offline calculation is performed based on the Clip-Crossmodel Attention to generate a multimodal fusion code for the user's historical behavior sequence vector generated after encoding in step S42 and the skin care product image and text features corresponding to the top K candidate skin care products in the recall phase;

[0118] S48: Through the multi-layer perceptron (MLP), the fused feature vector is nonlinearly transformed to extract high-level features.

[0119] In the existing multimodal fusion encoding process, traditional methods often fail to fully utilize all feature information when fusing features of different modalities (such as image features, text features, etc.), resulting in information loss and poor recommendation effects. In addition, the complex relationships and interactions between multimodal features are difficult to capture effectively, and traditional methods cannot fully utilize this information, affecting the accuracy of the recommendation system. Furthermore, the user's personalized preferences and contextual information (such as season, climate, time, etc.) are not fully utilized, resulting in the recommendation system being unable to dynamically adjust according to the user's real-time status. In response to these problems, the present invention proposes a multimodal fusion encoding based on the VLSNR model through the MultiModel Attention Encoder, which can fully fuse the user's historical behavior sequence and the candidate set of skin care product image and text features, and use the multi-head attention mechanism to capture the complex relationship between multimodal features, thereby improving the accuracy of the recommendation system. In addition, by embedding vectorization of user preferences and contextual information, personalized features can be generated to achieve dynamic personalized recommendations, thereby enhancing the real-time and accuracy of recommendations.

[0120] S5: Based on Multi-Head Attention and GRU to capture the feature relationship of different positions in the input sequence, combined Figure 4 The specific steps are as follows:

[0121] S51: Obtain the user's historical behavior sequence vector from the MultiModel Attention Encoder in step S4;

[0122] S52: Obtaining the image and text feature vectors of the top K candidate skin care products;

[0123] S53: Input the user's historical behavior sequence vector and the image and text feature vectors of the top K candidate skin care products into the Multi-Head Attention mechanism, and use the multi-head attention mechanism to online calculate the relationship between features at different positions;

[0124] S54: Perform weighted summation on the calculation results and generate a fused feature vector through online calculation;

[0125] S55: The feature vector fused in step S54 is used as input and input into the GRU model time step by time;

[0126] S56: The GRU model updates the hidden state sequentially according to the input feature vector to capture the dependency relationship in the time series;

[0127] S57: The GRU model processes all input feature vectors at time steps;

[0128] S58: Output the processed feature vector as a time series feature vector.

[0129] In this embodiment, the multi-head attention mechanism uses multiple independent attention heads working in parallel to learn the relationships between features from different subspaces. For text features and image features, the distribution forms Q, K, and V. For each attention head, the dot product of Q and K is calculated, and then the attention weight is calculated through the Softmax function and applied to the V vector:

[0130]

[0131] Then fuse the attention results of the image and text:

[0132]

[0133] Finally, the outputs of all attention heads are concatenated and linearly transformed to obtain the final multimodal feature representation:

[0134] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O

[0135] Among them, Wo is the parameter matrix for linear transformation.

[0136] In the existing time series feature processing process, common problems mainly include the difficulty in capturing the feature relationship at different positions in the input sequence and insufficient processing of time dependencies. When processing user behavior data, traditional methods often ignore the interactive relationship between different features and the dependencies in the time series, resulting in poor performance of the recommendation system. To address these problems, the present invention is based on Multi-Head Attention and GRU to capture the feature relationship at different positions in the input sequence, and uses the multi-head attention mechanism to effectively capture the complex relationship between features at different positions in the input sequence. It also processes time dependencies through the GRU model to generate high-quality time series feature vectors, thereby improving the accuracy and dynamic response capabilities of the recommendation system.

[0137] S6: Based on Multi-gate Mixture-of-Experts (MMoE), prediction results of CTR, CVR and user satisfaction are generated through multiple expert networks and gating mechanisms.

[0138] In this embodiment, the Multi-gate Mixture of Experts (MMoE) network is a model for multi-task learning. It uses multiple expert networks and a gating mechanism to handle different tasks, effectively capturing the feature relationships between tasks and improving prediction accuracy.

[0139] like Figure 5 The specific steps are as follows:

[0140] S61: Concatenate the time series feature vector generated by GRU, the skin care product image and text feature fusion vector, the user image and text feature fusion vector, the personalized feature vector, and the context feature vector into a long vector as the input of the MMoE model;

[0141] S62: Input the long vectors concatenated in step S61 into multiple expert networks;

[0142] S63: Each expert network independently processes input features, extracts feature relationships, and generates independent feature outputs through online calculations through multi-layer neural networks;

[0143] S64: Pass the concatenated long vector to the gating network of each task;

[0144] S65: Design an independent gating network for each task (CTR prediction, CVR prediction, user satisfaction prediction). Each gating network generates the weight of each expert network according to the input features through a gating mechanism (such as the Softmax function) according to the requirements of different tasks. The specific steps are as follows:

[0145] Input feature processing: The gating network receives the concatenated input feature vector;

[0146] Weight calculation: Calculate the weight of each expert network based on the input features and task requirements to ensure that the sum of the weights is 1;

[0147] Feature weighting: Multiply the output of the expert network by the calculated weight to generate a weighted feature representation;

[0148] S66: To obtain the requirements for CTR, CVR, and user satisfaction, the gating network performs weighted summation on the output of the expert network based on the calculated weights to obtain the feature representation of each task. The specific formula is: T task represents the task feature representation, α i represents the weight of the i-th expert network, E i represents the output of the i-th expert network.

[0149] S67: After all tasks have their feature representations, the task feature representations are passed to the corresponding task-specific layers;

[0150] S68: The task-specific layer further processes the feature representation and generates online calculations to generate prediction results for CTR, CVR, and user satisfaction.

[0151] S69: Calculate the multi-task loss function to optimize the prediction performance of each task. The multi-task loss function formula is as follows:

[0152] L=αL CTR +βL CVR +γL Satisfaction

[0153] Among them, L is the total loss function, α is the weight coefficient of the CTR task, and L CTR is the click-through rate CTR loss, β is the weight coefficient of the CVR task, L CVR is the conversion rate CVR loss, γ is the weight coefficient of the user satisfaction task, L Satisfaction It is the loss of user satisfaction.

[0154] L CTR and L CVR The cross entropy loss function is used to calculate the error of click-through rate prediction. The loss function formula is as follows:

[0155]

[0156] Among them, y CTR is the actual click rate label, with a value of 0 or 1. is the click rate value predicted by the model, with a value range of [0,1], y CVR It is the actual conversion rate label, with a value of 0 or 1. It is the conversion rate value predicted by the model, and its value range is [0, 1].

[0157] L Satisfaction The mean square error (MSE) loss function is used to calculate the error of user satisfaction prediction. The loss function formula is as follows:

[0158]

[0159] Among them, y Satisfaction is the actual user satisfaction label, usually a continuous value. is the user satisfaction value predicted by the model.

[0160] In this embodiment, the task feature representation is combined with other task-related features, and the combined features are nonlinearly transformed through a fully connected layer or other neural network layer. An activation function (such as Sigmoid or Softmax) is used to generate the final prediction result. By integrating these prediction results, skin care products can be recommended more accurately and personalized, thereby improving user satisfaction and the overall performance of the system.

[0161] In this embodiment, the CTR prediction results are used to evaluate the probability of users clicking on recommended skin care products, thereby optimizing the display order of the recommendation list; the CVR prediction results are used to evaluate the probability of users purchasing recommended skin care products, ensuring that the recommended skin care products meet user needs and improving conversion rates; and the user satisfaction prediction results are used to evaluate user satisfaction with the recommendation results, thereby improving user experience.

[0162] In the existing multi-task prediction process, common problems mainly include conflicts and coordination problems between tasks and insufficient processing of multi-task feature relationships. When performing multi-task prediction, traditional methods are often unable to effectively resolve conflicts between tasks, resulting in the prediction accuracy of certain tasks being affected. In addition, traditional methods do not adequately process the feature relationships of different tasks and cannot fully utilize the potential correlations between tasks. To address these problems, the present invention generates CTR, CVR and user satisfaction prediction results based on MMoE through three expert networks and a gating mechanism. Through multiple expert networks and gating mechanisms, it can effectively process multi-task feature relationships, solve task conflict problems, and improve prediction accuracy; at the same time, through the multi-task loss function, it can balance the weights of each task and further improve the performance of the overall system.

[0163] Example 2

[0164] This embodiment provides a multimodal multi-task skin care recommendation system based on a VLSNR+MMOE combined fusion structure, which is used to implement the multimodal multi-task skin care recommendation method based on the VLSNR+MMOE combined fusion structure of the above-mentioned embodiment 1. The system includes: a skin feature extraction module, a skin care product knowledge graph construction module, a candidate skin care product set construction module, a multimodal attention encoding module, a multi-head attention calculation module, a time series feature vector construction module, a vector splicing module, and a prediction result generation module;

[0165] In this embodiment, the skin feature extraction module is used to extract skin features of the skin image based on the CNN model;

[0166] In this embodiment, the skin care product knowledge graph construction module is used to crawl skin care product information, parse skin care product instructions and user reviews, and convert them into RDF triples, define the skin care product ontology model, and construct the skin care product knowledge graph;

[0167] In this embodiment, the candidate skin care product set construction module is used to extract skin care product features from the skin care product knowledge graph, input the skin features and skin care product features of the skin image into the factorization machine for training, generate skin care product vectors and user interest vectors based on the trained parameters of the factorization machine, and calculate the cosine similarity between the two to generate the top K candidate skin care product set;

[0168] In this embodiment, the multimodal attention encoding module is used to encode the multimodal attention encoder based on the VLSNR model to generate the user history behavior sequence vector, the image and text feature vectors of the top K candidate skin care products, the personalized feature vector, the context feature vector, the product image and text feature fusion vector, and the user image and text feature fusion vector;

[0169] In this embodiment, the multi-head attention calculation module is used to perform nonlinear transformation on the feature vector formed by fusing the user's historical behavior sequence vector and the image and text feature vectors of the top K candidate skin care products based on a multi-layer perceptron, extract high-level features, and calculate the relationship between features at different positions based on the multi-head attention mechanism;

[0170] In this embodiment, the time series feature vector construction module is used to output the time series feature vector based on the GRU model;

[0171] In this embodiment, the vector splicing module is used to splice the time series feature vector, the product image and text feature fusion vector, the user image and text feature fusion vector, the personalized feature vector, and the context feature vector;

[0172] In this embodiment, the prediction result generation module is used to input the spliced ​​vectors into the MMoE model and output the prediction results of CTR, CVR and user satisfaction.

[0173] Example 3

[0174] This embodiment provides a storage medium, which may be a ROM, RAM, disk, CD or other storage medium. The storage medium stores one or more programs. When the program is executed by the processor, the multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure of embodiment 1 is implemented.

[0175] Example 4

[0176] This embodiment provides a computing device, which includes a processor and a memory, wherein the memory stores one or more programs. When the processor executes the program stored in the memory, the multimodal multi-task skin care recommendation based on the VLSNR+MMOE combined fusion structure of the above-mentioned embodiment 1 is implemented.

[0177] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A multimodal and multitasking skin care recommendation method based on a VLSNR+MMOE combined fusion structure, characterized in that: The steps include: Extract skin features of skin images based on CNN model; Crawl skin care product information, parse skin care product instructions and user reviews, convert them into RDF triples, define the skin care product ontology model, and build a skin care product knowledge graph; Extract skin care product features from the skin care product knowledge graph, input the skin features of the skin image and the skin care product features into the factorization machine for training, generate skin care product vectors and user interest vectors based on the trained parameters of the factorization machine, and calculate the cosine similarity between the two to generate the top K candidate skin care product set; The multimodal attention encoder based on the VLSNR model is used to encode and generate the user history behavior sequence vector, the image and text feature vectors of the top K candidate skin care products, the personalized feature vector, the context feature vector, the product image and text feature fusion vector, and the user image and text feature fusion vector; Based on a multi-layer perceptron, a nonlinear transformation is performed on the feature vector formed by fusing the user's historical behavior sequence vector and the image and text feature vectors of the top K candidate skincare products to extract high-level features. The relationship between features at different positions is calculated using a multi-head attention mechanism, and the time series feature vector is output based on the GRU model. The time series feature vector, product image and text feature fusion vector, user image and text feature fusion vector, personalized feature vector, and context feature vector are concatenated and input into the MMOE model to output the prediction results of CTR, CVR, and user satisfaction.

2. The multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure according to claim 1, characterized in that: The skin features of skin images are extracted based on the CNN model, including: Acquire skin images and perform image preprocessing; The preprocessed skin image is passed through the multi-layer convolutional layer and pooling layer of the CNN model to extract low-level to high-level features of the image. The convolutional layer uses the ReLU activation function and expands the feature map output by the convolutional layer into a one-dimensional feature vector through the fully connected layer or the GAP layer.

3. The multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure according to claim 1, characterized in that: Construct a knowledge graph for skin care products, specifically including: Crawl skin care product information; Analyze skin care product instructions and user reviews based on NLP technology; Define the skin care product ontology model and create an OWL ontology; The skin care product instructions and user review information are converted into RDF triples, and the RDF triples are stored in the Neo4j or AllegroGraph graph database to construct a skin care product knowledge graph.

4. The multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure according to claim 1, characterized in that: Based on the trained parameters of the factorization machine, we generate skin care product vectors and user interest vectors and calculate their cosine similarity to generate the top K candidate skin care product set. Specifically, we include: After factorization machine training, the embedding vectors of each feature of skin features and skin care product features, as well as the assigned weights of each feature, are obtained; The feature embedding vectors of all skin care product subsets are accumulated to form the representation vector of each skin care product, which is expressed as: Among them, P i represents the representation vector of the i-th skin care product, w productj is the weight of the jth skin care product feature, E productj is the embedding vector of the j-th skin care product feature; For skin care product feature vector P i Perform k-NN indexing, generate index tables and store them in the Fasis database; The skin features of the skin image input by the user are extracted and weighted summed with the corresponding distribution weights obtained after factor decomposition machine training to obtain the current user's interest vector U, which is expressed as: Among them, w userj is the weight of the jth skin feature, E skinj is the embedding vector of the j-th skin feature; Calculate the user interest vector U and each skin care product vector P stored in the Fasis database i The similarity scores are sorted from high to low to generate the Top K candidate skin care product set.

5. The multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure according to claim 1, characterized in that: The multimodal attention encoder based on the VLSNR model is used for encoding, specifically including: Obtain historical user interaction data on skincare products, encode it based on the Clip-CrossmodelAttention module of the multimodal attention encoder, and generate a user historical behavior sequence vector; Obtain multimodal information of the Top K candidate skincare products and generate image and text feature vectors of the Top K candidate skincare products based on the Clip-Crossmodel Attention module encoding; Based on the Clip-Crossmodel Attention module, the user's preferences and product attributes are encoded to generate personalized feature vectors; Based on the Clip-Crossmodel Attention module, the context information is encoded to generate a context feature vector; Based on the Clip-Crossmodel Attention module, the skin care product image and text features crawled from the website and the skin features extracted from the user input skin image are encoded to generate the product image and text feature fusion vector and the user image and text feature fusion vector.

6. The multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure according to claim 1, characterized in that: The time series feature vector, product image and text feature fusion vector, user image and text feature fusion vector, personalized feature vector, and context feature vector are concatenated and input into the MMOE model to output the prediction results of CTR, CVR, and user satisfaction, including: After vector splicing, the vectors are input into multiple expert networks to extract feature relationships. After vector splicing, they are passed to the gating network of each task. Each gating network generates the weight of each expert network through the gating mechanism according to the input features. The gating network performs weighted summation on the output of the expert network based on the calculated weights to obtain the feature representation of each task, which is specifically expressed as: Among them, T task represents the task feature representation, α i represents the weight of the i-th expert network, E i represents the output of the i-th expert network, and N represents the number of expert networks; The task feature representation is passed to the corresponding task layer to calculate and generate the prediction results of CTR, CVR and user satisfaction.

7. The multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure according to claim 6, characterized in that: The task feature representation is passed to the corresponding task layer and the multi-task loss function is calculated, which is specifically expressed as: L=αL CTR +βL CVR +γL Satisfaction Among them, L is the total loss function, α is the weight coefficient of the CTR task, and L CTR is the click-through rate CTR loss, β is the weight coefficient of the CVR task, L CVR is the conversion rate CVR loss, γ is the weight coefficient of the user satisfaction task, L Satisfaction It is the loss of user satisfaction; L CTR and L CVR The error of click-through rate prediction is calculated using the cross entropy loss function, which is expressed as: Among them, y CTR is the actual click rate label, with a value of 0 or 1. is the click rate value predicted by the model, with a value range of [0,1], y CVR It is the actual conversion rate label, with a value of 0 or 1. is the conversion rate value predicted by the model, ranging from [0, 1]; L Satisfaction The error of user satisfaction prediction is calculated using the mean square error loss function, which is expressed as: Among them, y Satisfaction is the actual user satisfaction label, is the user satisfaction value predicted by the model.

8. A multimodal and multitask skin care recommendation system based on a VLSNR+MMOE combined fusion structure, characterized by: A multimodal multitask skin care recommendation method based on a VLSNR+MMOE combined fusion structure for implementing any one of claims 1-7, the system comprising: a skin feature extraction module, a skin care product knowledge graph construction module, a candidate skin care product set construction module, a multimodal attention encoding module, a multi-head attention calculation module, a time series feature vector construction module, a vector splicing module, and a prediction result generation module; The skin feature extraction module is used to extract skin features of the skin image based on the CNN model; The skin care product knowledge graph construction module is used to crawl skin care product information, parse skin care product instructions and user comments, and convert them into RDF triples, define the skin care product ontology model, and construct the skin care product knowledge graph; The candidate skin care product set construction module is used to extract skin care product features from the skin care product knowledge graph, input the skin features and skin care product features of the skin image into a factor decomposition machine for training, generate skin care product vectors and user interest vectors based on the trained parameters of the factor decomposition machine, and calculate the cosine similarity between the two to generate a Top K candidate skin care product set; The multimodal attention encoding module is used to encode based on the multimodal attention encoder of the VLSNR model to generate a user historical behavior sequence vector, a picture and text feature vector of the top K candidate skin care products, a personalized feature vector, a context feature vector, a product picture and text feature fusion vector, and a user picture and text feature fusion vector; The multi-head attention calculation module is used to perform nonlinear transformation on the feature vector obtained by fusing the user's historical behavior sequence vector and the image and text feature vectors of the top K candidate skin care products based on a multi-layer perceptron, extract high-level features, and calculate the relationship between features at different positions based on the multi-head attention mechanism; The time series feature vector construction module is used to output the time series feature vector based on the GRU model; The vector splicing module is used to splice the time series feature vector, the product image and text feature fusion vector, the user image and text feature fusion vector, the personalized feature vector, and the context feature vector; The prediction result generation module is used to input the spliced ​​vectors into the MMOE model and output the prediction results of CTR, CVR and user satisfaction.

9. A storage medium storing a program, characterized in that: When the program is executed by a processor, the multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure as described in any one of claims 1 to 7 is implemented.

10. A computing device comprising a processor and a memory for storing a program executable by the processor, characterized in that When the processor executes the program stored in the memory, it implements the multimodal multitasking skin care recommendation method based on the VLSNR+MMOE combined fusion structure as described in any one of claims 1-7.