AI make-up personalized generation method and system combining user portraits and fashion trends
By combining user profiles with popular trends to create personalized AI-powered makeup, and utilizing a hybrid model of generative adversarial networks and physical rendering engines, the problem of opaque decision-making and privacy protection in existing technologies is solved. This achieves a closed-loop optimization of high-fidelity makeup effects and user trust, thereby improving personalization and recommendation accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUJIAN LUYE FAIRY MIRROR SYSTEM INTEGRATION CO LTD
- Filing Date
- 2025-11-26
- Publication Date
- 2026-04-24
AI Technical Summary
Existing AI-powered personalized makeup generation methods suffer from opaque recommendation decision-making processes, physical distortion of virtual makeup effects, and contradictions between user privacy protection and continuous model optimization. These issues lead to fragmented user experiences and a lack of trust, making it difficult to achieve large-scale commercial applications.
By combining user profiles and popular trends, dynamic context vectors are generated. A hybrid model of generative adversarial networks and physical rendering engines is used to generate high-fidelity makeup solutions. Interpretability analysis and federated learning are introduced for closed-loop optimization to ensure user privacy and security.
It achieves high-fidelity makeup effect simulation, decision-making transparency, and user trust, improves personalization and recommendation accuracy, and continuously optimizes the model while protecting user privacy.
Smart Images

Figure CN121921387A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an AI-powered personalized makeup generation method and system that combines user profiles with fashion trends. Background Technology
[0002] In recent years, with the deep integration of artificial intelligence technology into computer vision and recommendation systems, personalized makeup recommendation technology has made significant progress. Existing solutions mostly match users' static features (such as skin tone and face shape) with macro-level trend data, generating virtual try-on effects through deep learning models. This technology has initially achieved the function of filtering suitable options from massive amounts of products and, relying on augmented reality (AR) technology, provided users with a preliminary interactive experience, laying the technological foundation for the development of this field.
[0003] Current technologies still have many limitations. Recommendation systems often employ "black box" models, with opaque decision-making processes that make it difficult for users to understand the rationale behind recommendations, leading to a lack of trust. Virtual makeup try-on effects generally lack physical realism, especially under complex lighting conditions, where the makeup effect deviates significantly from the actual application, resulting in a disjointed user experience. Existing methods also present a conflict between data utilization and user privacy protection, making it difficult to achieve continuous model optimization and personalized evolution while protecting sensitive biometric information. These shortcomings collectively hinder the large-scale commercial application and deep user adoption of these technologies. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides an AI-powered personalized makeup generation method that combines user profiles and fashion trends to address the problems of opaque recommendation decision-making processes and physical distortion of virtual makeup effects in existing AI-powered personalized makeup generation methods, the contradiction between user privacy protection and continuous model optimization, and how to construct a closed-loop personalized makeup generation system that combines high-fidelity experience, decision interpretability, and privacy security.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides an AI-powered personalized makeup generation method that combines user profiles with fashion trends, characterized by comprising the following steps:
[0008] Acquire and fuse the user's multimodal context data to generate a structured dynamic context vector;
[0009] Based on dynamic context vectors, global trends are individually weighted to generate a personalized weighted trend vector.
[0010] Personalized weighted trend vectors are input into the hybrid generative model to generate high-fidelity makeup schemes. The hybrid generative model simultaneously executes texture generation by the generative adversarial network and optical simulation by the physical rendering engine.
[0011] Perform interpretability analysis on the generated high-fidelity makeup solutions to generate a visual decision report that includes the weights of the reasons for recommendation;
[0012] User feedback is received based on visualized decision reports, and personalized weighted trend vectors are dynamically adjusted accordingly to iteratively optimize high-fidelity makeup solutions.
[0013] Finally, the optimized makeup scheme is generated and stored, and the user profile model is updated based on this interaction to complete the closed-loop learning.
[0014] As a preferred embodiment of the AI-powered personalized makeup generation method combining user profiles and trend analysis described in this invention, the specific steps for acquiring and fusing the user's multimodal context data to generate a structured dynamic context vector are as follows:
[0015] The ambient light sensor of the mobile device collects the color temperature and intensity data of the ambient light, and at the same time calls the front camera to capture the real-time video stream of the user's face to extract the coordinates of key facial points and micro-expression action units.
[0016] The collected raw ambient light data and facial video stream data are respectively input into a lightweight convolutional neural network for feature extraction to obtain quantized ambient light feature vectors and facial dynamic feature vectors.
[0017] On the user terminal device, the ambient lighting feature vector, facial dynamic feature vector and locally stored user historical preference vector are aggregated by the federated averaging algorithm to form a preliminary local fusion vector.
[0018] The local fusion vector is input into a recurrent neural network for time series modeling, and finally outputs a structured dynamic context vector representing the user's current integrated state.
[0019] As a preferred embodiment of the AI-powered personalized makeup generation method combining user profiles and trend analysis described in this invention, the step of generating a personalized weighted trend vector by applying a personalized weighted trend vector to the global trend based on a dynamic context vector includes the following specific steps:
[0020] The global trend map, generated by analyzing global fashion data sources using graph neural networks, is obtained from the cloud. This map represents the popularity of different makeup attributes in vector form.
[0021] The generated dynamic context vector is multiplied by the global trend vector to calculate the initial correlation score between the user profile and each trend element.
[0022] Using an attention mechanism network, the initial relevance score is modulated a second time based on the intensity of the emotional state in the dynamic context vector, thereby enhancing the weight of trend elements that have a high degree of resonance with the user's current emotions.
[0023] The modulated weights of each trend element are weighted and fused with the original global popular trend vector to output a unique, personalized weighted trend vector.
[0024] As a preferred embodiment of the AI-powered personalized makeup generation method combining user profiles and trend analysis described in this invention, the method involves inputting a personalized weighted trend vector into a hybrid generation model to generate high-fidelity makeup schemes. The hybrid generation model simultaneously executes texture generation via a generative adversarial network and optical simulation via a physically based rendering engine. The specific steps are as follows:
[0025] The personalized weighted trend vector is used as a conditional input and fed into a pre-trained conditional generative adversarial network generator to generate a preliminary two-dimensional makeup texture map.
[0026] The personalized weighted trend vector is mapped to the material parameter space of the physical rendering engine, and a material model matching the trend is called from the predefined physical material library, including gloss, metallicity and subsurface scattering parameters.
[0027] The generated 2D makeup texture map is UV-mapped and aligned with the user's real-time facial model, and the physical rendering engine is driven to perform ray tracing-based lighting calculations based on the current ambient lighting data and material model to simulate the real optical performance of makeup on the curved surface of the face.
[0028] The texture mapping and physically rendered results are fused and post-processed at the pixel level to output the final high-fidelity makeup solution image.
[0029] As a preferred embodiment of the AI-powered personalized makeup generation method combining user profiles and trend analysis described in this invention, the specific steps of performing interpretability analysis on the generated high-fidelity makeup scheme to generate a visual decision report containing recommendation reason weights are as follows:
[0030] Attribution analysis was used to calculate the contribution of each feature in the dynamic context vector and personalized weighted trend vector during the generation process to each pixel in the final makeup scheme.
[0031] Features with high contribution are clustered and summarized into core decision factors, including color compatibility, trend fit and sentiment matching.
[0032] Assign a quantifiable weight percentage to each core decision factor and generate a corresponding natural language description fragment;
[0033] The weight percentages are integrated with natural language description fragments and overlaid on the sidebar of the high-fidelity makeup solution image to form a visual decision report.
[0034] As a preferred embodiment of the AI-powered personalized makeup generation method combining user profiles and trend analysis described in this invention, the specific steps of receiving user feedback based on a visualized decision report and dynamically adjusting the personalized weighted trend vector accordingly to iteratively optimize the high-fidelity makeup scheme are as follows:
[0035] Analyze user feedback commands provided through natural language or interface sliders to the visual decision report, and translate the commands into intentions to adjust the weights of specific decision factors;
[0036] Based on the stated adjustment intention, the weight values of the corresponding trend elements in the personalized weighted trend vector are corrected in reverse.
[0037] The revised personalized weighted trend vector is then input back into the hybrid generation model to trigger a new round of high-fidelity makeup scheme generation.
[0038] The newly generated makeup scheme and the revised decision factors are then combined to form a visual decision report, which is presented to the user for confirmation or further optimization until the user is satisfied.
[0039] As a preferred embodiment of the AI-powered personalized makeup generation method combining user profiles and trend analysis described in this invention, the finalized and optimized makeup scheme is stored, and the user profile model is updated based on this interaction to complete closed-loop learning. The specific steps are as follows:
[0040] The user-confirmed final makeup plan, the corresponding personalized weighted trend vector, and the user interaction log are packaged together to form a complete session record.
[0041] Differential privacy technology is used to add noise to the session record for desensitization, ensuring that it cannot be traced back to the individual's identity;
[0042] The anonymized session records are encrypted and transmitted to the federated learning server as training data to update the global user profile model.
[0043] After the federated learning aggregation update is completed on the server side, the updated model parameters are distributed to user terminals, thereby completing closed-loop learning without centrally collecting the original data.
[0044] Secondly, this invention provides an AI-powered personalized makeup generation system that combines user profiles with popular trends, including:
[0045] The dynamic perception module acquires and fuses the user's multimodal context data to generate a structured dynamic context vector;
[0046] The trend modulation module, based on dynamic context vectors, performs personalized weighting on global popular trends to generate a personalized weighted trend vector;
[0047] The hybrid rendering module inputs personalized weighted trend vectors into the hybrid generative model to generate high-fidelity makeup schemes. The hybrid generative model simultaneously executes texture generation by the generative adversarial network and optical simulation by the physical rendering engine.
[0048] The decision analysis module performs interpretability analysis on the generated high-fidelity makeup solutions to generate a visual decision report that includes the weights of the recommendation reasons;
[0049] The interaction optimization module receives user feedback based on a visual decision report and dynamically adjusts the personalized weighted trend vector accordingly to iteratively optimize the high-fidelity makeup solution.
[0050] The closed-loop update module finalizes and stores the optimized makeup scheme, and updates the user profile model based on this interaction to complete the closed-loop learning.
[0051] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the AI-based personalized makeup generation method combining user profiles and trend information as described in the first aspect of the present invention.
[0052] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the AI-based personalized makeup generation method combining user profiles and trend information as described in the first aspect of the present invention.
[0053] The beneficial effects of this invention are as follows: By constructing a collaborative mechanism of dynamic context vectors and personalized weighted trend vectors, the real-time status of users and macro-popular trends are effectively integrated, significantly improving the accuracy and personalization of recommendations; by adopting a hybrid model of generative adversarial networks and physical rendering engines, the technical bottlenecks of light and shadow distortion and unrealistic material representation in virtual makeup try-on are solved, achieving high-fidelity simulation of makeup effects; the introduction of an interpretability analysis module to generate visual decision reports breaks the "black box" limitation of recommendation systems, enhancing user trust and participation in decision-making; based on a closed-loop update architecture of federated learning and differential privacy, a balance between continuous model optimization and privacy security is achieved while ensuring that sensitive data such as user biometrics does not leave the device, thus constructing a complete technical solution that combines dynamic adaptability, realistic experience, decision transparency, and data security. Attached Figure Description
[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1 This is a flowchart of the AI-powered personalized makeup generation method that combines user profiles and fashion trends in Example 1. Detailed Implementation
[0056] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0057] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0058] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0059] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides an AI-powered personalized makeup generation method that combines user profiles with fashion trends. The method includes the following steps:
[0060] Acquire and fuse the user's multimodal context data to generate a structured dynamic context vector;
[0061] Based on dynamic context vectors, global trends are individually weighted to generate a personalized weighted trend vector.
[0062] Personalized weighted trend vectors are input into the hybrid generative model to generate high-fidelity makeup schemes. The hybrid generative model simultaneously executes texture generation by the generative adversarial network and optical simulation by the physical rendering engine.
[0063] Perform interpretability analysis on the generated high-fidelity makeup solutions to generate a visual decision report that includes the weights of the reasons for recommendation;
[0064] User feedback is received based on visualized decision reports, and personalized weighted trend vectors are dynamically adjusted accordingly to iteratively optimize high-fidelity makeup solutions.
[0065] Finally, the optimized makeup scheme is generated and stored, and the user profile model is updated based on this interaction to complete the closed-loop learning.
[0066] It should be noted that,
[0067] This step involves real-time acquisition of ambient light parameters, including color temperature and light intensity data, from the user's environment using the mobile terminal device's ambient light sensor. Simultaneously, a high-resolution front-facing camera captures a sequence of dynamic facial video at 30 frames per second. The acquired raw video stream is processed by a facial keypoint detection algorithm to extract 468 three-dimensional coordinate points, and a facial motion coding system analyzes the intensity of muscle motor units corresponding to micro-expressions. Ambient light data and facial dynamic data are input into two parallel lightweight convolutional neural networks for feature distillation. The former outputs an ambient light vector containing spectral features, while the latter generates a facial dynamic vector representing the expression state. In the user's local storage, the system retrieves the user's historical preference vector, processed with differential privacy, and uses a federated averaging algorithm to weight and aggregate the three feature vectors to form a preliminary local fusion vector. This fusion vector is then fed into a gated recurrent unit network for time-series modeling, analyzing the feature evolution patterns over twenty consecutive time steps, ultimately outputting a structured dynamic context vector with temporal correlation.
[0068] The system retrieves a global trend map generated by a graph neural network architecture from a cloud-based distributed database. This map is constructed by semantically parsing millions of fashion news items, social media content, and runway images, encoding the popularity index of different makeup attributes in high-dimensional vector form. The local terminal performs tensor multiplication on the dynamic context vector generated in step one and the global trend vector from the cloud, calculating cosine similarity to obtain an initial association matrix between the user profile and each trend element. Subsequently, an attention mechanism network is activated. This network performs a non-linear transformation on the initial association matrix based on the intensity features of the emotional state in the dynamic context vector, significantly enhancing the weights of trend elements with high resonance with the user's current emotional state while suppressing interference from irrelevant elements. Finally, the modulated weight matrix is weighted and fused with the original trend vector to generate a personalized weighted trend vector with user-specific characteristics.
[0069] This step employs a hybrid architecture that integrates a conditional generative adversarial network (GAN) with a physically based rendering engine. First, a personalized weighted trend vector is input as a conditional parameter into a pre-trained GAN generator. Deconvolution operations generate a high-resolution 2D makeup texture map that meets the trend requirements. Simultaneously, the system maps the trend vector to the material parameter space of the physically based rendering engine, precisely matching corresponding material models from a pre-built physical material library, including physical properties such as surface gloss, metallic reflectivity, and skin subsurface scattering coefficient. Then, UV mapping technology precisely fits the generated texture map onto the user's 3D facial mesh model, driving the ray-tracing-based physically based rendering engine. This, combined with real-time ambient lighting data, simulates the light reflection, refraction, and diffusion phenomena of makeup products on the facial surface. Finally, multi-channel blending technology is used to pixel-level fuse the texture map and the physically based rendering result, and after color correction, the final high-fidelity makeup solution image is output.
[0070] The system employs an inter-layer correlation propagation algorithm to perform reverse attribution analysis on key decision nodes in the generation process. It constructs a complete decision influence map by calculating the contribution value of each feature dimension in the dynamic context vector and personalized weighted trend vector to each pixel of the final makeup scheme. Feature dimensions with contribution values exceeding a preset threshold are clustered to identify three to five core decision factors, such as color compatibility index, trend fit coefficient, and sentiment matching index. Each core factor is then assigned a quantifiable weight percentage, and the numerical results are transformed into easily understandable text descriptions using a natural language generation template. Finally, these text descriptions are integrated with the corresponding weight data, and a visual decision report containing color-coded bar charts and text descriptions is generated through a graphical interface engine and overlaid on the interactive sidebar of the makeup scheme image.
[0071] The system uses a natural language processing module to analyze user feedback on the decision report via voice or text, while simultaneously monitoring user slider operations on the interactive interface. These two input signals are then converted into instructions to adjust the weights of specific decision factors. Based on the identified adjustment intent, the system traces back to the corresponding trend element weight parameters in the personalized weighted trend vector and uses a gradient descent algorithm to directionally correct the weight values. The updated personalized weighted trend vector is then re-input into the hybrid generation model, triggering a new round of makeup scheme generation. The newly generated scheme and the corrected decision factors are then processed again by the interpretability analysis module to generate an updated visual report, forming a complete optimization loop. This iterative process continues until the user confirms satisfaction or the preset maximum number of iterations is reached.
[0072] The system packages the user-confirmed final makeup scheme image, the corresponding personalized weighted trend vector parameter set, and complete user interaction logs into a single data package. It then uses serialization encoding technology to generate a standard-format session log file. A differential privacy protection mechanism is applied to this session log, adding random noise conforming to a Laplace distribution to de-identify sensitive personal information, ensuring that the data cannot be used to infer an individual's identity. The processed encrypted data is uploaded to a federated learning server cluster via a secure transmission protocol, serving as training samples for updating the global user profile model. The server aggregates uploaded data from millions of terminals, executes a federated averaging algorithm to update the central model parameters, and then differentially distributes the optimized model parameters to each user terminal, thereby achieving continuous model evolution while absolutely protecting user privacy.
[0073] Specifically, the steps for acquiring and fusing the user's multimodal context data to generate a structured dynamic context vector are as follows:
[0074] The ambient light sensor of the mobile device collects the color temperature and intensity data of the ambient light, and at the same time calls the front camera to capture the real-time video stream of the user's face to extract the coordinates of key facial points and micro-expression action units.
[0075] The collected raw ambient light data and facial video stream data are respectively input into a lightweight convolutional neural network for feature extraction to obtain quantized ambient light feature vectors and facial dynamic feature vectors.
[0076] On the user terminal device, the ambient lighting feature vector, facial dynamic feature vector and locally stored user historical preference vector are aggregated by the federated averaging algorithm to form a preliminary local fusion vector.
[0077] The local fusion vector is input into a recurrent neural network for time series modeling, and finally outputs a structured dynamic context vector representing the user's current integrated state.
[0078] It should be noted that,
[0079] This step activates the mobile terminal's ambient light sensor to continuously collect ambient light color temperature and intensity data at a sampling frequency of 60 times per second, while simultaneously using the front-facing 12-megapixel camera to capture a high-definition video stream of the user's face at 30 frames per second. The acquired video stream data is used to locate 468 3D facial coordinate points in real time using a facial keypoint detection algorithm, including key areas such as eyelid contours, lip boundaries, and the bridge of the nose. Simultaneously, a facial motion coding system is used to analyze the video sequence frame by frame, identifying and quantifying the activation intensity of 44 basic facial motion units, particularly micro-expression features such as eyebrow raising and mouth corner stretching. The ambient light sensor data undergoes noise reduction and standardization processing by a digital signal processor, outputting a standard-format ambient light parameter set. All acquired data is transmitted to the local processing unit via a secure encrypted channel, providing a complete raw data source for subsequent feature extraction.
[0080] By using multi-sensor synchronous acquisition technology, a precise correspondence between ambient lighting and facial dynamic data was achieved, providing a complete data foundation for subsequent analysis. A high-precision facial motion analysis algorithm was adopted to capture micro-expression changes that are difficult to identify using traditional methods, providing a reliable basis for emotional state analysis. Finally, a complete preprocessing process from raw data to structured features was established.
[0081] This step inputs the preprocessed ambient light data into a specially designed lightweight convolutional neural network (CNN), which contains five convolutional layers and three fully connected layers. It extracts the spatial distribution patterns of illumination features through a sliding window of the convolutional kernels, ultimately outputting a 128-dimensional ambient light feature vector. Facial video stream data is first used to generate a facial mesh sequence using keypoint coordinates, then input into another parallel lightweight CNN. This network uses 3D convolutional kernels to process temporal data, extracting muscle movement patterns and facial expression changes from consecutive video frames, outputting a 256-dimensional facial dynamic feature vector. Both neural networks employ depthwise separable convolutional structures to reduce computation and enable real-time inference on mobile devices. Batch normalization is used during feature extraction to ensure numerical stability, all intermediate feature maps are processed using the ReLU activation function, and the final output feature vector is L2 normalized and stored in a temporary buffer.
[0082] Through a specially optimized lightweight neural network architecture, efficient feature extraction capabilities are achieved on mobile devices. By using 3D convolution to process temporal facial data, dynamic features of facial expression changes are effectively captured. Compared with traditional static feature extraction methods, it can better reflect the user's true state and provide high-quality feature input for subsequent multimodal fusion.
[0083] This step retrieves the user's historical preference vector, processed with differential privacy, from the encrypted storage area of the terminal device. This vector contains the user's past makeup selection records and rating data. The system concatenates the ambient lighting feature vector, facial dynamic feature vector, and historical preference vector according to preset weights to form an initial concatenated vector. Then, a localized variant of the federated averaging algorithm is used to fuse the three feature vectors into a preliminary 512-dimensional fused vector through a weighted averaging operation. The weights are dynamically adjusted based on feature confidence, with ambient lighting features having a weight of 0.3, facial dynamic features having a weight of 0.4, and historical preference vector having a weight of 0.3. A moving average technique is used during the fusion process to smooth weight changes and ensure the stability of the output vector. The final generated local fused vector is standardized and temporarily stored in the device's memory for subsequent processing.
[0084] By localizing the federated averaging algorithm, privacy-preserving fusion of multi-source features is achieved. A dynamic weight allocation mechanism is adopted to automatically adjust the contribution ratio according to the confidence level of different features, ensuring that the fusion result reflects both real-time status and historical preferences. Compared with fixed-weight fusion methods, it is more adaptable to changes in user status.
[0085] This step inputs the locally fused vector sequence into a recurrent neural network composed of gated recurrent units. This network contains two hidden layers, each with 128 neurons, and learns the dependencies of features over time through a gating mechanism. The network processes the fused vector sequence over 20 consecutive time steps using a sliding window approach, with each time step spaced 100 milliseconds apart. During time series modeling, updating the gating controls the retention ratio of historical information, and resetting the gating determines the degree of fusion for the current input. Through the coordinated work of these two gating mechanisms, the network can capture the gradual change patterns of the user's state. Finally, the hidden state of the last time step is mapped to a 256-dimensional dynamic context vector through a fully connected layer. This vector contains comprehensive information about the user's environmental adaptability, emotional state, and preference tendencies, and carries temporal evolution characteristics.
[0086] By using recurrent neural networks for time series modeling, the features of isolated moments are elevated into a dynamic context that includes the laws of temporal evolution. The gating mechanism effectively captures the continuous changes in user state, enabling the system to understand the process of state transition rather than just the instantaneous state, thus providing contextual information with a time dimension for personalized recommendations.
[0087] Specifically, the step of generating a personalized weighted trend vector by individually weighting the global trend based on the dynamic context vector involves the following steps:
[0088] The global trend map, generated by analyzing global fashion data sources using graph neural networks, is obtained from the cloud. This map represents the popularity of different makeup attributes in vector form.
[0089] The generated dynamic context vector is multiplied by the global trend vector to calculate the initial correlation score between the user profile and each trend element.
[0090] Using an attention mechanism network, the initial relevance score is modulated a second time based on the intensity of the emotional state in the dynamic context vector, thereby enhancing the weight of trend elements that have a high degree of resonance with the user's current emotions.
[0091] The modulated weights of each trend element are weighted and fused with the original global popular trend vector to output a unique, personalized weighted trend vector.
[0092] It should be noted that,
[0093] The system accesses a cloud-based distributed database via a secure application programming interface (API) to obtain a global trend graph generated by a graph neural network (GNN) architecture. This GNN analyzes diverse data sources, including social media platforms, digital content from fashion magazines, fashion week runway images, and e-commerce sales data, to construct a fashion knowledge graph containing millions of nodes. Each node in the graph represents a specific makeup attribute, including hue, saturation, brightness, texture type, and gloss level. Edge weights between nodes represent the strength of the association between attributes. The GNN extracts node features through multi-layer graph convolution operations, then aggregates global information through a graph attention mechanism, ultimately encoding each makeup attribute into a 128-dimensional feature vector. These feature vectors are organized hierarchically according to attribute categories, forming a complete trend graph, and the data is updated periodically to maintain its timeliness.
[0094] By deeply mining diverse fashion data sources through graph neural networks, the system transforms massive unstructured data into structured trend maps. By adopting a hierarchical vector coding method, the system can accurately capture the popularity changes of different makeup attributes, providing a comprehensive and accurate trend benchmark for subsequent personalized weighting, overcoming the shortcomings of traditional methods in that the trend information is one-sided and the updates are lagging.
[0095] This step performs a tensor multiplication operation between a locally generated 256-dimensional dynamic context vector and a 128-dimensional global trend vector obtained from the cloud. First, the two vectors are dimensionally aligned, and the dynamic context vector is projected onto a 128-dimensional feature space through a fully connected layer. Then, an outer product operation is performed to generate a 128×128 two-dimensional correlation matrix. Each element of this matrix represents the initial correlation strength between a specific feature dimension in the dynamic context vector and the corresponding dimension in the trend vector. Next, the matrix is summed row by row to compress it into a 128-dimensional initial correlation score vector, where each score quantifies the degree of matching between the user's current state and a single trend element. The entire calculation process is completed on the neural network processor of the mobile device, using fixed-point arithmetic to optimize computational efficiency and ensure real-time requirements.
[0096] By using high-dimensional space mapping through tensor multiplication, deep feature interaction between user profiles and popular trends is achieved. The outer product operation is used to capture the second-order association between features, which can reveal more complex matching relationships than the traditional dot product operation, providing richer semantic information for subsequent personalized modulation and effectively improving the accuracy of trend recommendation.
[0097] This step constructs a modulation network based on a scaled dot product attention mechanism. The query vector is derived from the sentiment state feature segment of the dynamic context vector, while the key and value vectors are derived from the initial relevance score vector. First, a 32-dimensional sentiment state feature sub-vector is extracted from the 256-dimensional dynamic context vector and used as the query input for the attention mechanism. Then, the 128-dimensional initial relevance score vector is used as the key and value inputs, respectively, and projected onto a 64-dimensional feature space through a linear transformation. The attention mechanism calculates the dot product similarity between the query vector and all key vectors and applies the Softmax function to generate an attention weight distribution. This weight distribution is then weighted and summed with the value vector to produce a sentiment-modulated relevance score vector. Specifically, when a user is detected to be in a strong emotional state, the attention mechanism automatically amplifies the weights of trend elements that align with that emotional tone.
[0098] By modulating emotion perception through an attention mechanism, dynamic adaptation of trend recommendations to users' emotional states is achieved. By adopting a query mechanism based on emotion features, the system can intelligently adjust the weight allocation of trend elements according to the user's current emotional state, thereby enhancing the emotional fit of the recommendation results and improving the emotional resonance of the user experience.
[0099] This step transforms the sentiment-modulated 128-dimensional relevance score vector into a weighted coefficient vector using a sigmoid activation function, ensuring that each weight value falls between 0 and 1. These weighted coefficients are then element-wise multiplied with the original global trend vector to achieve personalized weighting of trend elements. The weighted trend vector undergoes L2 normalization to eliminate the influence of dimensions and maintain consistent vector magnitudes. Next, the normalized weighted trend vector is concatenated with the original dynamic context vector to form a 384-dimensional extended feature vector. This extended vector is then processed through a multilayer perceptron with 128 neurons for feature compression and refinement, ultimately outputting a 256-dimensional personalized weighted trend vector. The entire fusion process employs residual connection technology to preserve important information from the original trend features.
[0100] By modulating the weight coefficients and trend vectors element by element, a precise transformation from global trends to personalized trends is achieved. The use of a multilayer perceptron for feature refinement ensures that the output vector retains both the commonalities of trends and reflects individual characteristics, ultimately generating trend guidance that truly matches the user's personality and current state, providing precise directional guidance for subsequent makeup generation.
[0101] Specifically, the personalized weighted trend vector is input into the hybrid generative model to generate high-fidelity makeup schemes. The hybrid generative model simultaneously executes texture generation by the generative adversarial network and optical simulation by the physically based rendering engine. The specific steps are as follows:
[0102] The personalized weighted trend vector is used as a conditional input and fed into a pre-trained conditional generative adversarial network generator to generate a preliminary two-dimensional makeup texture map.
[0103] The personalized weighted trend vector is mapped to the material parameter space of the physical rendering engine, and a material model matching the trend is called from the predefined physical material library, including gloss, metallicity and subsurface scattering parameters.
[0104] The generated 2D makeup texture map is UV-mapped and aligned with the user's real-time facial model, and the physical rendering engine is driven to perform ray tracing-based lighting calculations based on the current ambient lighting data and material model to simulate the real optical performance of makeup on the curved surface of the face.
[0105] The texture mapping and physically rendered results are fused and post-processed at the pixel level to output the final high-fidelity makeup solution image.
[0106] It should be noted that,
[0107] This step employs a conditional generative adversarial network based on the U-Net architecture. The generator receives a 256-dimensional personalized weighted trend vector as conditional input. The generator first projects the conditional vector onto a high-dimensional feature space through a fully connected layer, then concatenates it with a random noise vector to form the initial input. This input undergoes a series of upsampling blocks, each containing a transposed convolutional layer, a batch normalization layer, and a ReLU activation function, progressively increasing the feature map resolution from 4×4 to 512×512. During decoding, the generator fuses features extracted in the encoder stage through skip connections, ensuring the detail quality of the output texture. The final output layer uses the Tanh activation function to generate a two-dimensional makeup texture map conforming to the RGB color space. This map contains color distribution information for areas such as eyeshadow, lip gloss, and blush, maintaining semantic consistency with the input trend vector.
[0108] By using the guided generation mechanism of conditional generative adversarial networks, abstract trend vectors are transformed into concrete makeup textures, realizing a direct conversion from digital trends to visual design. By using conditional control of the generation process, the output results are ensured to strictly follow the personalized trend guidance, providing a high-quality texture foundation for subsequent physical rendering and overcoming the technical limitations of traditional texture generation being disconnected from trends.
[0109] This step constructs a material parameter mapping network consisting of three fully connected layers, mapping a 256-dimensional personalized weighted trend vector to the material parameter space. First, the network calculates the probability distribution of material types, and then selects a base material template from a predefined physical material library based on the maximum probability value. The material library contains twelve base material types, each described by a bidirectional reflectance distribution function model. The network then further outputs specific material parameter adjustments, including continuous values for the surface roughness coefficient (0-1), scaled values for the metallicity parameter (0-1), and a preset range for the subsurface scattering coefficient based on skin optical properties. These parameters are linearly interpolated with the base material template to generate the final trend-adapted material model.
[0110] By using spatial mapping technology for material parameters, the precise conversion from trend semantics to physical properties is achieved. By combining probability distribution with parameter adjustment, the rationality of material selection and the precision of parameter optimization are ensured, providing accurate material descriptions for physical rendering and solving the problem of material and trend mismatch in traditional methods.
[0111] This step first constructs a real-time facial mesh model of the user using 468 3D coordinate points obtained through facial keypoint detection, and generates continuous facial surfaces using a parametric surface fitting algorithm. Then, adaptive UV mapping technology is employed to accurately project 2D makeup texture maps onto the 3D facial mesh, ensuring accurate texture fit in complex curved areas such as the nose wings and eye sockets. The physically based rendering engine receives color temperature and intensity data from an ambient light sensor to construct a physically based realistic lighting environment. The engine uses a Monte Carlo path tracing algorithm to simulate the reflection, refraction, and subsurface scattering of light on the makeup material surface, with precise calculations specifically for the anisotropic reflection of pearlescent materials and the Fresnel effect of lip gloss materials. Each pixel is sampled using 256 rays to balance computational efficiency and rendering quality.
[0112] By combining UV mapping and ray tracing, a natural transformation from 2D textures to 3D makeup effects is achieved; physical-based lighting calculations accurately simulate complex optical phenomena, ensuring consistency between virtual makeup effects and real-world lighting environments, significantly enhancing the realism and credibility of the makeup trial results.
[0113] This step employs multi-channel fusion technology to separate the color information of the texture map from the lighting information of the physically rendered result. First, the texture map undergoes color space conversion from RGB to HSV, preserving the hue and saturation channels. Simultaneously, the luminance channel and normal information are extracted from the physically rendered result. Then, a weighted fusion algorithm is used to reconstruct the hue and saturation of the texture map with the luminance information from the physically rendered result, ensuring color accuracy while maintaining realistic lighting. The fused image undergoes bilateral filtering for noise reduction, preserving edge details while smoothing color transitions. Finally, color correction is performed, adjusting the gamma value according to the display device's characteristic curve, and tone mapping is applied to convert the high dynamic range image to a standard dynamic range image, outputting the final makeup scheme image suitable for mobile device display.
[0114] Through multi-channel separation and fusion technology, a perfect combination of color accuracy and lighting realism is achieved; a professional post-processing workflow optimizes the visual effect, ensuring that the output image maintains the design intent and conforms to the laws of visual perception, ultimately providing a high-fidelity makeup try-on experience with photorealistic realism.
[0115] Specifically, the steps for performing interpretability analysis on the generated high-fidelity makeup scheme to generate a visual decision report containing recommendation reason weights are as follows:
[0116] Attribution analysis was used to calculate the contribution of each feature in the dynamic context vector and personalized weighted trend vector during the generation process to each pixel in the final makeup scheme.
[0117] The features with high contribution are clustered and summarized into core decision factors, which include color compatibility, trend fit and sentiment matching.
[0118] Assign a quantifiable weight percentage to each core decision factor and generate a corresponding natural language description fragment;
[0119] The weight percentages are integrated with natural language description fragments and overlaid on the sidebar of the high-fidelity makeup solution image to form a visual decision report.
[0120] It should be noted that,
[0121] This step uses an inter-layer correlation propagation algorithm for attribution analysis. First, dynamic context vectors and personalized weighted trend vectors are used as input features, and all forward propagation computation paths during the generation process are recorded. The algorithm then backpropagates from each pixel in the output layer, calculating the gradient contribution of each feature to the final result according to the chain rule. During backpropagation, specific rules are used to assign correlation scores to neurons in each layer, ensuring that the sum of all correlation scores equals the activation value of the output layer. Gradient-weighted activation mapping is used for convolutional neural network layers, and depth-enhanced interpretation is used for fully connected layers, ultimately generating a contribution heatmap with the same resolution as the output image. Each pixel in this heatmap contains a multi-dimensional vector representing the contribution of each dimension of the input features to that pixel.
[0122] By employing the precise attribution of the interlayer correlation propagation algorithm, the contribution of input features to output pixels is traced, transforming the black-box generation process into a quantifiable feature impact analysis. The multi-dimensional contribution mapping technology can accurately identify the degree of influence of different features on the final effect, providing a reliable data foundation for subsequent decision analysis and significantly improving the transparency and interpretability of the system's decision-making process.
[0123] This step first sets a contribution threshold to filter out feature dimensions that significantly impact the final result. The selected high-contribution features are then grouped into a feature vector set, and unsupervised clustering is performed using the K-means clustering algorithm. During clustering, the silhouette coefficient method is used to determine the optimal number of clusters, typically dividing the features into three to five core categories. Each cluster center represents a decision factor, named according to its feature composition, such as color compatibility, trend fit, and sentiment matching. The color compatibility category mainly includes feature dimensions related to skin tone harmony and color contrast; the trend fit category includes feature dimensions related to matching popular elements; and the sentiment matching category includes feature dimensions related to the resonance of the user's emotional state. The feature dimensions within each cluster are highly correlated, collectively influencing a specific aspect of the decision-making effect.
[0124] Cluster analysis is used to summarize the dispersed feature contributions into core decision factors with clear semantics, realizing the abstract transformation from low-level features to high-level semantics. Unsupervised learning methods are used to automatically discover the intrinsic relationships between features, ensuring that the summarized decision factors are both comprehensive and mutually exclusive, providing a clear logical framework for subsequent weight allocation and interpretation generation.
[0125] This step first calculates the total contribution of the features included in each core decision factor, and then normalizes the total contribution of each factor into a percentage weight. The weight calculation uses a softmax function to ensure that the sum of all weights is 100%. Simultaneously, a natural language generation template is constructed for each decision factor, containing the factor name, weight value, and a specific description of its impact. Color compatibility emphasizes the harmonious relationship between skin tone and makeup color; trend fit highlights how popular elements are combined with personal characteristics; and sentiment matching focuses on the resonance between makeup style and emotional state. The system automatically selects different adjectives and adverbs to strengthen the description based on the weight, using strong modifiers such as "significant" and "main" for high-weight factors and weak modifiers such as "moderate" and "auxiliary" for low-weight factors.
[0126] By combining weight quantification with natural language generation, abstract decision factors are transformed into intuitive numerical descriptions and textual explanations. The template-based language generation mechanism ensures the standardization and consistency of the explanations, while the description intensity is adaptively adjusted according to the weights, making the final explanations both accurate and easy to understand, effectively improving the credibility and acceptability of the recommendation results.
[0127] This step involves designing an interactive visualization interface, featuring a fixed-width decision report sidebar to the right of the makeup scheme image. The top of the sidebar displays the overall recommendation confidence score, while the bottom displays detailed information for each decision factor, sorted by weight in descending order. Each decision factor uses a horizontal bar chart to show its weight percentage, with bar colors matching the factor type, and a corresponding natural language description on the right. The interface employs a responsive layout to ensure all information is displayed fully on different screen sizes. Users can click on decision factor entries to view more detailed explanations, including the specific feature dimensions involved and their contribution distribution. All visual elements use a consistent color scheme and font style to ensure the overall aesthetics and readability of the interface.
[0128] By spatially integrating and visually presenting multimodal information, numerical weights, textual explanations, and visual solutions are organically combined. An interactive interface design empowers users to explore the details of decision-making, constructing a complete explanatory experience. This completely solves the black-box problem of recommendation systems, enabling users to fully understand the basis for recommendations and build trust in the system.
[0129] Specifically, the process of receiving user feedback based on visualized decision reports and dynamically adjusting personalized weighted trend vectors accordingly to iteratively optimize high-fidelity makeup solutions involves the following steps:
[0130] Analyze user feedback commands provided through natural language or interface sliders to the visual decision report, and translate the commands into intentions to adjust the weights of specific decision factors;
[0131] Based on the stated adjustment intention, the weight values of the corresponding trend elements in the personalized weighted trend vector are corrected in reverse.
[0132] The revised personalized weighted trend vector is then input back into the hybrid generation model to trigger a new round of high-fidelity makeup scheme generation.
[0133] The newly generated makeup scheme and the revised decision factors are then combined to form a visual decision report, which is presented to the user for confirmation or further optimization until the user is satisfied.
[0134] It should be noted that,
[0135] This step captures user feedback by integrating a natural language processing module and an interface event listener. For natural language input, the system employs a Transformer-based semantic understanding model. First, it segments and tags the user-input text, then extracts key intent features through an attention mechanism to identify the type of decision factor the user wishes to enhance or weaken, and the direction of adjustment. For interface slider operations, the system monitors slider position changes in real time, mapping continuous slider values to discrete decision factor adjustment instructions. Both input methods are ultimately unified into standardized adjustment intention tuples, containing three core fields: target decision factor identifier, adjustment direction (enhancement / weakening), and adjustment intensity. The system also records feedback timestamps and session context information, providing complete input data for subsequent weight adjustments.
[0136] By using a multimodal feedback parsing mechanism, user subjective preferences are transformed into standardized adjustment intentions that can be understood by machines, thus establishing a semantic bridge for human-computer interaction. The adoption of a unified intention representation format ensures the consistency of subsequent processing, laying a data foundation for precise personalized optimization and realizing the intelligent transformation from vague user feedback to explicit optimization instructions.
[0137] This step constructs a backpropagation correction network, taking the adjustment intention as input and calculating the correction amount for each element in the personalized weighted trend vector through a three-layer fully connected neural network. The network first locates the corresponding feature dimension group in the trend vector based on the target decision factor identifier in the adjustment intention; then, based on the adjustment direction and intensity, it calculates the required weight change for each dimension. The correction process employs the idea of gradient backpropagation, but uses manually designed propagation rules instead of automatic differentiation: for "enhancement" instructions, the corresponding dimension weights are increased proportionally; for "weakening" instructions, the weights are decreased inversely. Simultaneously, a weight normalization constraint is introduced to ensure that the sum of all dimension weights remains unchanged after correction, avoiding weight inflation. The corrected vector also undergoes smoothing filtering to eliminate abrupt changes during the adjustment process.
[0138] By using a neural network-based reverse correction mechanism, direct mapping from user feedback to model parameters is achieved. Constrained optimization methods are employed to ensure the rationality and stability of weight correction, fully respecting user intent while maintaining the mathematical rigor of the system, making personalized optimization both flexible and reliable.
[0139] This step restarts the entire workflow of the hybrid generative model, simultaneously inputting the corrected, personalized weighted trend vector into both the Conditional Generative Adversarial Network (CGAN) and the Physically Based Rendering (PBR) engine. The system first checks the validity of the input vector, verifying its dimensionality matching and numerical rationality. Then, it executes two generation branches in parallel: in the CGAN branch, the corrected vector influences each layer of the generator's feature map through a conditional injection mechanism, regenerating makeup texture maps adapted to the new weights; in the PBR branch, the corrected vector is remapped to the material parameter space, updating the configuration of physical properties such as gloss and metallicity. The outputs of both branches then enter the UV mapping and ray tracing process, generating a high-fidelity makeup scheme that matches the user's latest preferences based on the same facial model and ambient lighting data.
[0140] The end-to-end regeneration mechanism ensures that the optimized makeup solutions are consistent with user feedback in terms of texture design and physical performance; the parallel processing architecture improves iteration efficiency, enabling the system to quickly respond to user adjustment needs and achieve true real-time personalized optimization.
[0141] This step re-executes the interpretability analysis process, generating an updated visual decision report based on the revised personalized weighted trend vector and the new makeup scheme. The system first uses the same attribution analysis method to calculate the contribution of each decision factor in the new scheme, then generates the corresponding weight percentages and natural language descriptions. The updated decision report highlights the changed parts in the sidebar, using color coding and animated transitions to visually demonstrate the differences before and after optimization. The system provides two main entry points: confirmation and continued optimization, and records user dwell time and interaction behavior as satisfaction evaluation indicators. When the user chooses to continue optimization, the system retains the current optimization state and enters a new interaction loop; when the user confirms satisfaction, the system marks the current scheme as the final result.
[0142] Through a closed-loop iterative optimization mechanism, one-way recommendation is upgraded to a two-way collaborative design process; differential visualization technology is used to enhance users' understanding of the optimization effect, and the recommendation accuracy is continuously improved by combining multiple rounds of interaction records, ultimately realizing a personalized makeup co-creation experience with deep user participation.
[0143] Specifically, the finalized and optimized makeup scheme is stored, and the user profile model is updated based on this interaction to complete closed-loop learning. The specific steps are as follows:
[0144] The user-confirmed final makeup plan, the corresponding personalized weighted trend vector, and the user interaction log are packaged together to form a complete session record.
[0145] Differential privacy technology is used to add noise to the session record for desensitization, ensuring that it cannot be traced back to the individual's identity;
[0146] The anonymized session records are encrypted and transmitted to the federated learning server as training data to update the global user profile model.
[0147] After the federated learning aggregation update is completed on the server side, the updated model parameters are distributed to user terminals, thereby completing closed-loop learning without centrally collecting the original data.
[0148] It should be noted that,
[0149] This step initiates the data packaging engine, first retrieving the final makeup scheme image data confirmed by the user from the temporary storage area, including the high-resolution rendering result and the corresponding texture map file. The system then reads the personalized weighted trend vector generated for this session from memory, which contains 256 floating-point feature weight values. Simultaneously, it extracts the complete user operation record from the interaction log database, including the initial recommended scheme, adjustment instructions for each round, user feedback, and timestamp information. All data is serialized according to a predefined binary format; the makeup scheme image uses a lossless compression algorithm to reduce storage space; the vector data maintains its original precision; and the interaction log is converted to a structured text format. During the packaging process, a checksum is added to each data block to ensure data integrity, ultimately generating an independent session log file containing a header, data body, and checksum section. The header records the session identifier, time information, and data format version.
[0150] By using a structured data packaging mechanism, scattered multimodal interaction information is integrated into complete conversation units, providing a standardized data foundation for subsequent privacy processing and model training. The use of checksums and version control ensures the integrity and compatibility of data during transmission and processing, achieving comprehensive digital recording of user interaction behavior.
[0151] This step employs an epsilon-differential privacy protection mechanism. First, sensitive fields in the session logs are identified and classified, including facial feature data, environmental location information, and user preference tags. Based on preset privacy budget parameters, the system calculates the required Laplacian noise intensity. For numerical data, such as trend vector weights, random noise conforming to a Laplacian distribution is added, with the standard deviation of the noise proportional to the global sensitivity and privacy budget of the data. For biometric information in image data, local differential privacy technology is used, perturbing the feature point coordinates through a random response mechanism. Text-based interaction logs undergo generalization processing, removing direct identifiers and converting detailed timestamps into time intervals. All noise addition processes utilize cryptographically secure random number generators to ensure the unpredictability of the noise, and the processed data meets strict mathematical privacy protection definitions.
[0152] By ensuring the mathematical rigor of differential privacy, a precise balance is established between data availability and privacy protection; by adopting a multimodal differentiated de-identification strategy, the most appropriate protection measures are implemented for the characteristics of different types of data, fundamentally eliminating the possibility of reverse inference of user identity.
[0153] This step initiates a secure transmission protocol. First, the anonymized session log file is divided into blocks, each 1MB in size to accommodate network transmission requirements. A hybrid encryption mechanism is used: a symmetric encryption algorithm is employed to encrypt the data content, while the key is protected using an asymmetric encryption algorithm. The encrypted data is then transmitted to the federated learning server cluster via a secure channel established by the Hypertext Transfer Protocol (HTTP). Data compression is enabled during transmission to reduce bandwidth consumption. Upon receiving the data, the server first verifies its integrity and digital signature, then stores the valid session records in a temporary training database. The system organizes this anonymized data chronologically, preparing batch training samples for subsequent model updates while maintaining the distributed nature of the data source.
[0154] The end-to-end encrypted transmission system ensures the security of anonymized data during transmission; the use of block processing and compression optimization adapts to the characteristics of the mobile network environment, providing a stable and reliable data supply channel for federated learning.
[0155] This step executes a distributed model training process on the federated learning server. First, it extracts the gradient information needed for model updates from the anonymized data uploaded by each user terminal. The server uses a federated averaging algorithm to weight and aggregate partial model updates from millions of terminals, with weights dynamically adjusted based on the data quality and quantity from each terminal. Secure multi-party computation is used during the aggregation process to protect the update privacy of individual users, ensuring the server cannot infer the original data of a specific user. The updated global model parameters are quantized and compressed to reduce the amount of data transmitted, and then distributed to each user terminal via a content delivery network. After receiving the new parameters, the terminal device completes a hot model update locally, without affecting normal user operation, while preserving the user's personalized model components unaffected by the global update.
[0156] Through the distributed training architecture of federated learning, a perfect balance between global model optimization and personal data protection is achieved; a progressive model update mechanism is adopted to ensure that the system continues to evolve without interrupting service, ultimately building a virtuous cycle system that both protects privacy and has continuous learning capabilities.
[0157] This embodiment also provides an AI-powered personalized makeup generation system that combines user profiles with fashion trends, including:
[0158] The dynamic perception module acquires and fuses the user's multimodal context data to generate a structured dynamic context vector;
[0159] The trend modulation module, based on dynamic context vectors, performs personalized weighting on global popular trends to generate a personalized weighted trend vector;
[0160] The hybrid rendering module inputs personalized weighted trend vectors into the hybrid generative model to generate high-fidelity makeup schemes. The hybrid generative model simultaneously executes texture generation by the generative adversarial network and optical simulation by the physical rendering engine.
[0161] The decision analysis module performs interpretability analysis on the generated high-fidelity makeup solutions to generate a visual decision report that includes the weights of the recommendation reasons;
[0162] The interaction optimization module receives user feedback based on a visual decision report and dynamically adjusts the personalized weighted trend vector accordingly to iteratively optimize the high-fidelity makeup solution.
[0163] The closed-loop update module finalizes and stores the optimized makeup scheme, and updates the user profile model based on this interaction to complete the closed-loop learning.
[0164] This embodiment also provides a computer device applicable to the AI-powered personalized makeup generation method that combines user profiles and fashion trends, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the AI-powered personalized makeup generation method that combines user profiles and fashion trends as proposed in the above embodiment.
[0165] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0166] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the AI-powered personalized makeup generation method combining user profiles and trend information as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0167] In summary, this invention achieves significant improvements in recommendation accuracy and personalization by: constructing a collaborative mechanism between dynamic context vectors and personalized weighted trend vectors, effectively integrating real-time user status with macro-level trends; employing a hybrid model combining generative adversarial networks and physically based rendering engines, overcoming the technical bottlenecks of lighting distortion and unrealistic material representation in virtual makeup try-ons, and realizing high-fidelity makeup effect simulation; introducing an interpretability analysis module to generate visual decision reports, breaking the "black box" limitation of recommendation systems and enhancing user trust and participation in decision-making; and establishing a closed-loop update architecture based on federated learning and differential privacy, ensuring that sensitive data such as user biometrics remains on-device while balancing continuous model optimization with privacy security. This results in a complete technical solution that combines dynamic adaptability, realistic experience, decision transparency, and data security.
[0168] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for generating personalized AI-powered makeup products that combines user profiles with fashion trends, characterized in that: Includes the following steps: Acquire and fuse the user's multimodal context data to generate a structured dynamic context vector; Based on dynamic context vectors, global trends are individually weighted to generate a personalized weighted trend vector. Personalized weighted trend vectors are input into the hybrid generative model to generate high-fidelity makeup schemes. The hybrid generative model simultaneously executes texture generation by the generative adversarial network and optical simulation by the physical rendering engine. Perform interpretability analysis on the generated high-fidelity makeup solutions to generate a visual decision report that includes the weights of the reasons for recommendation; User feedback is received based on visualized decision reports, and personalized weighted trend vectors are dynamically adjusted accordingly to iteratively optimize high-fidelity makeup solutions. Finally, the optimized makeup scheme is generated and stored, and the user profile model is updated based on this interaction to complete the closed-loop learning.
2. The AI-powered personalized makeup generation method combining user profiles and trend analysis as described in claim 1, characterized in that: The specific steps for acquiring and fusing the user's multimodal context data to generate a structured dynamic context vector are as follows: The ambient light sensor of the mobile device collects the color temperature and intensity data of the ambient light, and at the same time calls the front camera to capture the real-time video stream of the user's face to extract the coordinates of key facial points and micro-expression action units. The collected raw ambient light data and facial video stream data are respectively input into a lightweight convolutional neural network for feature extraction to obtain quantized ambient light feature vectors and facial dynamic feature vectors. On the user terminal device, the ambient lighting feature vector, facial dynamic feature vector and locally stored user historical preference vector are aggregated by the federated averaging algorithm to form a preliminary local fusion vector. The local fusion vector is input into a recurrent neural network for time series modeling, and finally outputs a structured dynamic context vector representing the user's current integrated state.
3. The AI-powered personalized makeup generation method combining user profiles and trend analysis as described in claim 2, characterized in that: The process of generating a personalized weighted trend vector by applying personalized weights to the global trend based on dynamic context vectors involves the following steps: The global trend map, generated by analyzing global fashion data sources using graph neural networks, is obtained from the cloud. This map represents the popularity of different makeup attributes in vector form. The generated dynamic context vector is multiplied by the global trend vector to calculate the initial correlation score between the user profile and each trend element. Using an attention mechanism network, the initial relevance score is modulated a second time based on the intensity of the emotional state in the dynamic context vector, thereby enhancing the weight of trend elements that have a high degree of resonance with the user's current emotions. The weighted trend elements are weighted and fused with the original global trend vector to output a unique, personalized weighted trend vector.
4. The AI-powered personalized makeup generation method combining user profiles and trend analysis as described in claim 3, characterized in that: The process involves inputting a personalized weighted trend vector into a hybrid generative model to generate high-fidelity makeup schemes. The hybrid generative model simultaneously executes texture generation via a generative adversarial network and optical simulation via a physically based rendering engine. The specific steps are as follows: The personalized weighted trend vector is used as a conditional input and fed into a pre-trained conditional generative adversarial network generator to generate a preliminary two-dimensional makeup texture map. The personalized weighted trend vector is mapped to the material parameter space of the physical rendering engine, and a material model matching the trend is called from the predefined physical material library, including gloss, metallicity and subsurface scattering parameters. The generated 2D makeup texture map is aligned with the user's real-time facial model through UV mapping, and the physical rendering engine is driven to perform ray tracing-based lighting calculations based on the current ambient lighting data and material model to simulate the real optical performance of makeup on the curved surface of the face. The texture mapping and physically rendered results are fused and post-processed at the pixel level to output the final high-fidelity makeup solution image.
5. The AI-powered personalized makeup generation method combining user profiles and trend analysis as described in claim 4, characterized in that: The specific steps for performing interpretability analysis on the generated high-fidelity makeup scheme to generate a visual decision report containing the weights of the recommendation reasons are as follows: Attribution analysis was used to calculate the contribution of each feature in the dynamic context vector and personalized weighted trend vector during the generation process to each pixel in the final makeup scheme. The features with high contribution are clustered and summarized into core decision factors, which include color compatibility, trend fit and sentiment matching. Assign a quantifiable weight percentage to each core decision factor and generate a corresponding natural language description fragment; The weight percentages are integrated with natural language description fragments and overlaid on the sidebar of the high-fidelity makeup solution image to form a visual decision report.
6. The AI-powered personalized makeup generation method combining user profiles and trend analysis as described in claim 5, characterized in that: The process of receiving user feedback based on visualized decision reports and dynamically adjusting personalized weighted trend vectors to iteratively optimize high-fidelity makeup solutions involves the following steps: Analyze user feedback commands provided through natural language or interface sliders to the visual decision report, and translate the commands into intentions to adjust the weights of specific decision factors; Based on the stated adjustment intention, the weight values of the corresponding trend elements in the personalized weighted trend vector are corrected in reverse. The revised personalized weighted trend vector is then input back into the hybrid generation model to trigger a new round of high-fidelity makeup scheme generation. The newly generated makeup scheme and the revised decision factors are then combined to form a visual decision report, which is presented to the user for confirmation or further optimization until the user is satisfied.
7. The AI-powered personalized makeup generation method combining user profiles and trend analysis as described in claim 6, characterized in that: The finalized and optimized makeup scheme is stored, and the user profile model is updated based on this interaction to complete closed-loop learning. The specific steps are as follows: The user-confirmed final makeup plan, the corresponding personalized weighted trend vector, and the user interaction log are packaged together to form a complete session record. Differential privacy technology is used to add noise to the session record for desensitization, ensuring that it cannot be traced back to the individual's identity; The anonymized session records are encrypted and transmitted to the federated learning server as training data to update the global user profile model. After the federated learning aggregation update is completed on the server side, the updated model parameters are distributed to user terminals, thereby completing closed-loop learning without centrally collecting the original data.
8. An AI-powered personalized makeup generation system combining user profiles and fashion trends, based on the AI-powered personalized makeup generation method combining user profiles and fashion trends as described in any one of claims 1 to 7, characterized in that: include, The dynamic perception module acquires and fuses the user's multimodal context data to generate a structured dynamic context vector; The trend modulation module, based on dynamic context vectors, performs personalized weighting on global popular trends to generate a personalized weighted trend vector; The hybrid rendering module inputs personalized weighted trend vectors into the hybrid generative model to generate high-fidelity makeup schemes. The hybrid generative model simultaneously executes texture generation by the generative adversarial network and optical simulation by the physical rendering engine. The decision analysis module performs interpretability analysis on the generated high-fidelity makeup solutions to generate a visual decision report that includes the weights of the recommendation reasons; The interaction optimization module receives user feedback based on a visual decision report and dynamically adjusts the personalized weighted trend vector accordingly to iteratively optimize the high-fidelity makeup solution. The closed-loop update module finalizes and stores the optimized makeup scheme, and updates the user profile model based on this interaction to complete the closed-loop learning.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the AI-based personalized makeup generation method combining user profiles and trend information as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the AI-based personalized makeup generation method combining user profiles and trend information as described in any one of claims 1 to 7.