Intelligent makeup recommendation and makeup trying method and system based on multi-modal perception

By employing multimodal perception technology and dynamic correlation models, the problems of ambient lighting mismatch and insufficient user preference capture in intelligent makeup recommendation and try-on have been solved, achieving accurate try-on effects and self-optimization capabilities in real-world environments, thereby enhancing the system's intelligence and user experience.

CN121921079APending Publication Date: 2026-04-24FUJIAN LUYE FAIRY MIRROR SYSTEM INTEGRATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIAN LUYE FAIRY MIRROR SYSTEM INTEGRATION CO LTD
Filing Date
2025-11-21
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing intelligent makeup recommendation and try-on technologies cannot dynamically adapt to the lighting characteristics of users' real environment, resulting in a discrepancy between the try-on effect and the actual visual experience. Furthermore, relying on explicit feedback makes it impossible to capture users' subconscious aesthetic preferences, and the system lacks self-evolution capabilities, resulting in insufficient recommendation accuracy and try-on realism.

Method used

By using multimodal perception technology, high dynamic range cameras, ambient light sensors, depth-sensing cameras, and brain-computer interface devices are used to collect multimodal data. Combined with physical rendering and dynamic correlation models, makeup parameters are adjusted in real time to form a closed-loop verification and incremental learning to optimize recommendation decisions.

Benefits of technology

It achieves precise adaptation of virtual makeup try-on effects under real-world lighting conditions, enhancing the realism and credibility of the try-on results, accurately meeting users' deep aesthetic needs, and possessing self-evolution and continuous learning capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921079A_ABST
    Figure CN121921079A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent make-up recommendation and make-up trying method and system based on multi-modal perception, and relates to the technical field of intelligent beauty makeup, and the method comprises the steps: collecting a user face image, an ambient light parameter, three-dimensional geometric information and an electroencephalogram signal through synchronously triggering a multi-modal sensor; a digital light field is dynamically constructed by using ambient light parameters and three-dimensional information, and a makeup trying effect image fused with ambient light is generated based on a physical rendering technology; quantizing a subconsciousness preference score of the user by decoding the electroencephalogram signal; generating a color adjustment instruction through a dynamic association model in combination with the ambient light parameter and the preference score; adjusting makeup parameters in real time and feeding back the makeup parameters to the user to form closed loop verification; according to the method, the problems that in the prior art, the makeup trying effect is mismatched with environment illumination and depends on explicit feedback and model staticization are solved, and the makeup trying authenticity, recommendation accuracy and system self-evolution ability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent beauty technology, and in particular to an intelligent makeup recommendation and try-on method and system based on multimodal perception. Background Technology

[0002] In recent years, with the deep integration of augmented reality and computer vision technologies, intelligent makeup recommendation and virtual try-on technologies have seen significant development. Existing solutions mostly focus on achieving static application of makeup products through facial feature point detection and using machine learning algorithms to provide personalized recommendations based on visible features such as skin tone and facial contours. Related systems capture user images through mobile device cameras, thereby simulating and overlaying makeup effects such as lipstick and eyeshadow, which to some extent improves the user experience and reduces the trial-and-error costs of online shopping.

[0003] Existing technologies suffer from fundamental limitations. Their virtual makeup try-on effects largely rely on preset rendering models under standard lighting conditions, failing to dynamically adapt to the lighting characteristics of the user's real-world environment. This results in a significant discrepancy between the try-on effect and the actual visual experience—a disconnect between the "try" and "use" scenarios. Furthermore, the user feedback relied upon by recommendation systems is primarily explicit ratings or clicks, failing to capture users' subconscious aesthetic preferences, thus limiting recommendation accuracy. Finally, the system models are mostly static, lacking the ability to self-evolve based on continuous interaction, making it difficult to adapt to dynamic changes in user aesthetic trends. These shortcomings collectively contribute to the deficiencies of existing technologies in recommendation accuracy, makeup realism, and system intelligence. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a multimodal perception-based intelligent makeup recommendation and try-on method to address the problems of existing intelligent makeup recommendation and try-on methods, such as the mismatch between virtual makeup effects and ambient lighting leading to distorted try-on results, reliance on explicit feedback failing to reveal users' true preferences, and how to construct an intelligent try-on system that can continuously self-optimize and dynamically adapt to the environment and users' subconscious.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides an intelligent makeup recommendation and try-on method based on multimodal perception, characterized by comprising the following steps:

[0008] The system synchronously triggers a high dynamic range camera, an ambient light sensor, a depth sensing camera, and a brain-computer interface device to collect multimodal data packets containing user facial images, ambient light parameters, three-dimensional environmental geometric information, and raw electroencephalogram (EEG) signals.

[0009] A digital light field is dynamically constructed using ambient light parameters and 3D environmental geometric information, and virtual makeup is rendered onto the user's facial image in real time based on physical rendering technology, generating a makeup effect image that blends with the ambient light.

[0010] While users view the makeup trial effect images, the brain signals collected by the brain-computer interface device are decoded to quantify the user's subconscious preference score for the current makeup trial effect.

[0011] Ambient light parameters and subconscious preference scores are input into a dynamic correlation model for analysis, generating color adjustment parameter instructions for optimizing the makeup trial effect;

[0012] The makeup parameters in the trial makeup effect image are dynamically adjusted according to the color adjustment parameter instructions, and the optimized trial makeup effect image is immediately fed back to the user to form a closed loop verification.

[0013] Incremental learning is performed on the dynamic association model based on data generated from multiple interactions, so as to continuously evolve and optimize the model's recommendation decision-making ability.

[0014] As a preferred embodiment of the intelligent makeup recommendation and try-on method based on multimodal perception described in this invention, the system simultaneously triggers a high dynamic range camera, an ambient light sensor, a depth sensing camera, and a brain-computer interface device to collect a multimodal data packet containing user facial images, ambient light parameters, three-dimensional environmental geometric information, and raw electroencephalogram (EEG) signals. The specific steps are as follows:

[0015] The system generates a global hardware synchronization signal, which is simultaneously sent to all participating sensor control units via wired interrupt or wireless broadcast.

[0016] After receiving the synchronization signal, the high dynamic range camera performs an exposure, captures multiple frames of images including overexposed and underexposed areas, and combines them into a complete facial image. At the same time, the ambient light sensor reads the current ambient light intensity and color temperature values.

[0017] When triggered by a synchronization signal, the depth-sensing camera uses structured light or time-of-flight principles to project invisible light patterns onto the user's face and surrounding environment and receive their deformation. This allows it to calculate the depth value of each pixel in the scene and construct three-dimensional point cloud data.

[0018] The brain-computer interface device uses this synchronization signal as the zero point of time and begins to continuously record the user's raw electroencephalogram (EEG) signals from the prefrontal cortex at a sampling rate of no less than 256 Hz. It also packages the data acquired by all sensors within the same clock cycle into a multimodal data packet with a unified timestamp.

[0019] As a preferred embodiment of the intelligent makeup recommendation and try-on method based on multimodal perception described in this invention, the method involves: dynamically constructing a digital light field using ambient light parameters and three-dimensional environmental geometric information, and rendering virtual makeup onto the user's facial image in real time using physically based rendering technology to generate a try-on effect image that blends with the ambient lighting. The specific steps are as follows:

[0020] The rendering engine receives ambient light parameters and 3D point cloud data from the data packet, converts the point cloud data into a 3D mesh model with normal vectors, and places a hemispherical ambient light with a corresponding color temperature in the virtual scene based on the light intensity and color temperature values, and determines the incident direction and intensity of the main light source based on the 3D mesh analysis.

[0021] The engine precisely maps the texture of the user's facial image onto the corresponding 3D mesh model and defines the material properties of the virtual makeup product, including base color, metallicity, roughness, and subsurface scattering coefficient for skin characteristics.

[0022] The physically based lighting shader is invoked, which calculates the diffuse reflection of ambient hemispherical light at various points on the facial surface and simulates specular highlights under the combined effects of metallicity and roughness. For foundation products, it also calculates the scattering of light beneath the skin surface to simulate a soft, translucent effect.

[0023] The facial area after color calculation is blended with the unmodified background area in real time, and a makeup effect image with correct light and shadow relationship under the current real ambient lighting conditions is output to the display device.

[0024] As a preferred embodiment of the intelligent makeup recommendation and try-on method based on multimodal perception described in this invention, the step of decoding the EEG signals collected by the brain-computer interface device when the user views the try-on effect image to quantify the user's subconscious preference score for the current try-on effect includes the following steps:

[0025] The system uses the precise moment when the makeup trial effect image is displayed on the screen as an event marker, and adds an event tag to the continuous EEG signal stream;

[0026] The EEG signal processing module takes the event marker as the starting point, extracts the EEG signal segment within a specific time window, and uses a blind source separation algorithm to filter out physiological artifacts in the signal caused by blinking, eye movement and muscle activity.

[0027] Temporal locked-means analysis was performed on the processed pure EEG signals to extract event-related potential components associated with cognitive assessment and emotional response, and their average amplitude within a specific time window was measured.

[0028] The measured event-related potential amplitudes, especially the amplitudes of the P300 component, are converted into a scalar value between 0 and 1 through a pre-calibrated linear mapping function. This value is the score that characterizes the intensity of the user's subconscious preference.

[0029] As a preferred embodiment of the intelligent makeup recommendation and try-on method based on multimodal perception described in this invention, the step of inputting ambient light parameters and subconscious preference scores into a dynamic correlation model for analysis to generate color adjustment parameter instructions for optimizing the try-on effect includes the following specific steps:

[0030] The dynamic association model receives the current set of ambient light parameters and the calculated subconscious preference scores, and normalizes the ambient light parameters into a standardized feature space.

[0031] The core of the model is a pre-trained tensor decomposition network, which performs tensor multiplication operations on the normalized ambient light features and subconscious preference scores to uncover the deep nonlinear relationship between ambient lighting and user aesthetic preferences.

[0032] The network output layer generates an optimization vector for the current makeup color based on the current association mode. This vector indicates the direction and magnitude of adjustment to the brightness, red-green axis, and yellow-blue axis components of the color parameters in the CIELAB color space.

[0033] The system encapsulates the optimization vector into a machine-readable color adjustment parameter instruction, which explicitly specifies how to modify the shader parameters of the current virtual makeup.

[0034] As a preferred embodiment of the intelligent makeup recommendation and try-on method based on multimodal perception described in this invention, the steps of dynamically adjusting the makeup parameters in the try-on effect image according to color adjustment parameter instructions and immediately feeding back the optimized try-on effect image to the user to form a closed-loop verification are as follows:

[0035] The parametric adjustment engine receives color adjustment parameter instructions and parses out the color space adjustment vectors contained therein;

[0036] The engine directly accesses the global variables of the currently used makeup shader and modifies the parameters that affect color performance in real time based on the adjustment vector, such as increasing the brightness component of the base color or decreasing the saturation component.

[0037] The changes take effect immediately. The physical rendering engine uses the updated shader parameters to quickly re-render the facial area of ​​the current frame. This process skips the time-consuming light field reconstruction and only updates the color and lighting locally.

[0038] The optimized makeup effect image is presented to the user in a very short delay, replacing the previous frame, thus completing a rapid closed-loop verification from analysis to execution, and is ready to receive new neural feedback from the user.

[0039] As a preferred embodiment of the intelligent makeup recommendation and try-on method based on multimodal perception described in this invention, the step of incrementally learning the dynamic association model based on data generated from multiple interactions to continuously evolve and optimize the model's recommendation decision-making ability includes the following steps:

[0040] The system encapsulates each complete interaction loop, including ambient light parameters, color adjustments performed, and the resulting changes in the user's subconscious preference scores, into a timestamped training sample and stores it in the local cache.

[0041] When the number of cached samples reaches a preset threshold, the system initiates a secure federated learning client process, which uses these new samples on the local device to calculate the update gradient of the model parameters.

[0042] The calculated gradients are uploaded to a central server, which aggregates gradient updates from multiple anonymous users and uses the aggregated gradients to perform an iterative optimization of the shared dynamic correlation model without uploading any original user data.

[0043] The system periodically downloads the latest version of model parameters from the server and completes local updates, enabling the model to continuously adapt to the ever-changing aesthetic trends and personal preferences of the user group, thereby evolving its decision-making capabilities.

[0044] Secondly, this invention provides an intelligent makeup recommendation and try-on system based on multimodal perception, comprising:

[0045] The synchronous acquisition module allows the system to synchronously trigger a high dynamic range camera, an ambient light sensor, a depth sensing camera, and a brain-computer interface device to acquire multimodal data packets containing user facial images, ambient light parameters, three-dimensional environmental geometric information, and raw electroencephalogram (EEG) signals.

[0046] The light field rendering module dynamically constructs a digital light field using ambient light parameters and 3D environmental geometric information, and renders virtual makeup onto the user's facial image in real time based on physical rendering technology, generating a makeup effect image that blends with the ambient light.

[0047] The neural decoding module decodes the electroencephalogram (EEG) signals collected by the brain-computer interface device when the user views the makeup trial effect image, in order to quantify the user's subconscious preference score for the current makeup trial effect.

[0048] The intelligent decision-making module inputs ambient light parameters and subconscious preference scores into a dynamic correlation model for analysis, generating color adjustment parameter instructions to optimize the makeup trial effect;

[0049] The closed-loop optimization module dynamically adjusts the makeup parameters in the makeup trial effect image according to the color adjustment parameter instructions, and immediately feeds back the optimized makeup trial effect image to the user to form a closed-loop verification.

[0050] The continuous learning module performs incremental learning on the dynamic association model based on data generated from multiple interactions, so as to continuously evolve and optimize the model's recommendation decision-making capabilities.

[0051] Thirdly, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the intelligent makeup recommendation and try-on method based on multimodal perception as described in the first aspect of the present invention.

[0052] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the intelligent makeup recommendation and try-on method based on multimodal perception as described in the first aspect of the present invention.

[0053] The beneficial effects of this invention are as follows: By constructing a deep fusion mechanism of ambient light field and neural feedback, the virtual makeup try-on effect is accurately adapted and dynamically optimized under real ambient lighting, significantly improving the realism and credibility of the try-on results; by using brain-computer interface to decode users' subconscious preferences, the subjectivity and limitations of traditional explicit feedback are overcome, making makeup recommendations more accurately match users' deep aesthetic needs; by establishing a dynamic association model and forming a closed-loop system of "perception-decision-optimization-verification", it can respond to changes in users and the environment in real time and achieve personalized adaptive adjustments; further, by using a federated learning framework for incremental learning, while fully protecting user data privacy, the system's generalization ability and intelligence level are continuously improved, ultimately constructing an intelligent makeup recommendation and try-on ecosystem with self-evolution capabilities. Attached Figure Description

[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is a flowchart of the intelligent makeup recommendation and try-on method based on multimodal perception in Example 1. Detailed Implementation

[0056] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0057] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0058] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0059] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides an intelligent makeup recommendation and try-on method based on multimodal perception, characterized by including the following steps:

[0060] The system synchronously triggers a high dynamic range camera, an ambient light sensor, a depth sensing camera, and a brain-computer interface device to collect multimodal data packets containing user facial images, ambient light parameters, three-dimensional environmental geometric information, and raw electroencephalogram (EEG) signals.

[0061] A digital light field is dynamically constructed using ambient light parameters and 3D environmental geometric information, and virtual makeup is rendered onto the user's facial image in real time based on physical rendering technology, generating a makeup effect image that blends with the ambient light.

[0062] While users view the makeup trial effect images, the brain signals collected by the brain-computer interface device are decoded to quantify the user's subconscious preference score for the current makeup trial effect.

[0063] Ambient light parameters and subconscious preference scores are input into a dynamic correlation model for analysis, generating color adjustment parameter instructions for optimizing the makeup trial effect;

[0064] The makeup parameters in the trial makeup effect image are dynamically adjusted according to the color adjustment parameter instructions, and the optimized trial makeup effect image is immediately fed back to the user to form a closed loop verification.

[0065] Incremental learning is performed on the dynamic association model based on data generated from multiple interactions, so as to continuously evolve and optimize the model's recommendation decision-making ability.

[0066] It should be noted that the system first generates a unified hardware synchronization signal, which is simultaneously sent to four sensing units—a high dynamic range camera, an ambient light sensor, a depth-sensing camera, and a brain-computer interface—via the device's internal bus or wireless commands. Upon receiving the synchronization signal, the high dynamic range camera immediately performs multi-exposure sequence acquisition, synthesizing image frames of different exposure durations into a complete dynamic range facial image. The ambient light sensor simultaneously captures ambient light intensity and color temperature values. The depth-sensing camera uses a structured light array to project a specific coded pattern, constructing a millimeter-precision 3D point cloud model by calculating light reflection deformation. The brain-computer interface device continuously acquires prefrontal cortex EEG signals at a sampling rate of no less than 256Hz, using the synchronization signal as a reference. All sensor data is timestamped and then encapsulated into a unified format multimodal data packet.

[0067] The hardware-level synchronization mechanism solves the problem of temporal misalignment of multi-source data, providing a data foundation for precise alignment for subsequent cross-modal analysis and ensuring the spatiotemporal consistency of ambient lighting, facial geometry and neural responses.

[0068] After receiving the data packet, the rendering engine first converts the 3D point cloud data into a mesh surface with normal vectors, and then constructs a hemispherical ambient lighting model in virtual space by combining ambient light parameters. The ray tracing engine calculates the incident angle and intensity distribution of the main light source, accurately mapping the user's facial image onto the 3D mesh, and then assigns physical material properties (including base hue, metallic reflectivity, surface roughness, and skin subsurface scattering coefficient) to the virtual makeup. A physically based lighting shader calculates the diffuse reflection, specular reflection, and transmission effects of ambient light on the surface pixel by pixel, and finally blends the rendered result with the real background in real time.

[0069] By using physically realistic light field modeling and material simulation, the distortion problem of traditional AR makeup try-on under complex lighting conditions is solved, enabling virtual makeup to maintain visual consistency in different environments.

[0070] When the makeup trial image is displayed, the system immediately marks the event trigger point in the EEG signal stream. The signal processing module uses an adaptive filtering algorithm to remove eye movement and electromyography artifacts, and extracts pure cognitive EEG signals through independent component analysis. Event-related potentials are subjected to time-domain locked averaging, and the amplitude characteristics of the P300 component and late positive potentials are accurately measured within a specific time window. Finally, the neuroelectric activity is quantified into preference intensity values ​​in the 0-1 interval using a pre-calibrated psychophysical mapping function.

[0071] Breaking through the limitations of traditional explicit feedback, this technology directly captures users' subconscious aesthetic responses through neural decoding, providing an objective and reliable source of preference data for recommendation systems.

[0072] The dynamic correlation model employs a tensor decomposition architecture, performing multi-dimensional tensor operations on normalized ambient light parameters and subconscious preference scores. The model analyzes the implicit correlation between ambient lighting characteristics and aesthetic preferences through a nonlinear mapping layer, generating adjustment vectors for brightness, red-green axis, and yellow-blue axis components within the CIELAB color space, ultimately outputting a machine-readable color correction instruction set.

[0073] A deep correlation model between environmental factors and neural feedback is established to achieve intelligent color adaptation based on multimodal perception, which significantly improves the contextual adaptability of recommendation decisions.

[0074] After parsing the color instructions, the parametric adjustment engine directly modifies the color space transformation matrix in the shader. The rendering pipeline skips the geometry processing stage and only performs real-time re-rendering of the makeup area based on the new parameters, completing the image update within milliseconds of latency. The optimized makeup effect is immediately presented to the user, while a new round of neural signal acquisition is initiated to verify the optimization effect.

[0075] This forms a complete closed loop of "perception-decision-execution-verification," continuously improving the user experience through rapid iterative optimization and achieving dynamic collaborative adaptation between the system and the user.

[0076] The system encapsulates the environmental parameters, adjustment commands, and preference score changes of each interaction into training samples, and uses a federated learning framework to calculate the model gradient locally. The central model parameters are updated through a secure multi-party aggregation mechanism, and the optimized model is periodically deployed to terminal devices.

[0077] While protecting user privacy, the model can continuously evolve, enabling the system to have long-term learning capabilities to adapt to dynamic changes in aesthetic trends.

[0078] Specifically, the system synchronously triggers a high dynamic range camera, an ambient light sensor, a depth-sensing camera, and a brain-computer interface device to collect multimodal data packets containing user facial images, ambient light parameters, three-dimensional environmental geometric information, and raw electroencephalogram (EEG) signals. The specific steps are as follows:

[0079] The system generates a global hardware synchronization signal, which is simultaneously sent to all participating sensor control units via wired interrupt or wireless broadcast.

[0080] After receiving the synchronization signal, the high dynamic range camera performs an exposure, captures multiple frames of images including overexposed and underexposed areas, and combines them into a complete facial image. At the same time, the ambient light sensor reads the current ambient light intensity and color temperature values.

[0081] When triggered by a synchronization signal, the depth-sensing camera uses structured light or time-of-flight principles to project invisible light patterns onto the user's face and surrounding environment and receive their deformation. This allows it to calculate the depth value of each pixel in the scene and construct three-dimensional point cloud data.

[0082] The brain-computer interface device uses this synchronization signal as the zero point of time and begins to continuously record the user's raw electroencephalogram (EEG) signals from the prefrontal cortex at a sampling rate of no less than 256 Hz. It also packages the data acquired by all sensors within the same clock cycle into a multimodal data packet with a unified timestamp.

[0083] It should be noted that the system's core processor generates precise synchronization pulse signals through a built-in hardware clock module. These signals are transmitted to each sensor control unit via the device's internal bus or a dedicated wireless communication protocol (such as ultra-wideband technology). In wired interrupt mode, the synchronization signal directly triggers the interrupt service routine of each sensor via an independent GPIO pin. In wireless broadcast mode, the synchronization signal, after data packet processing, is sent to all sensor nodes via the radio frequency module in the form of a broadcast address. To ensure timing consistency, the system pre-calculates the delay time of each transmission path before signal transmission and incorporates time compensation information into the signal encoding, enabling all sensors to synchronously initiate the data acquisition process within a microsecond-level error range.

[0084] By establishing a precise hardware-level synchronization mechanism, the timing drift problem during multi-source heterogeneous sensor data acquisition is fundamentally solved, providing a strictly aligned time reference for subsequent multimodal data fusion and significantly improving the accuracy of cross-modal data analysis.

[0085] Upon receiving the synchronization signal, the high dynamic range camera immediately initiates a rapid exposure sequence acquisition, continuously capturing three sets of image frames with different exposure durations within milliseconds: short exposure frames preserve highlight details, normal exposure frames record midtones, and long exposure frames capture shadow information. The image processing chip then aligns and corrects these three sets of raw images, using a tone mapping algorithm to perform pixel-level fusion of the optimal exposure areas from each frame, generating a detailed facial image. Simultaneously, the ambient light sensor measures the spectral distribution of ambient light through a built-in photodiode array, transmitting the quantized light intensity and color temperature values ​​to the main processor via an I2C interface.

[0086] The multi-exposure fusion technology overcomes the problem of missing image information under single exposure, ensuring the complete presentation of facial features under various lighting conditions, while providing accurate ambient lighting parameters for virtual rendering.

[0087] At the moment of synchronization signal triggering, the depth-sensing camera projects an encoded structured light pattern (such as Gray code or sinusoidal stripes) or emits a modulated laser pulse onto the target area via an infrared laser. When the light strikes the curved surface of the face, it deforms, and the infrared camera simultaneously captures the deformed light spot pattern. For the structured light scheme, the processor calculates depth information pixel-by-pixel by solving the phase difference between the original pattern and the deformed pattern; for the time-of-flight scheme, the distance is calculated by measuring the time difference between the laser's round trip. The final result is a 3D point cloud dataset containing hundreds of thousands of spatial coordinate points, each containing 3D coordinates and reflection intensity information.

[0088] High-precision 3D reconstruction of the face and environment was achieved through active optical ranging technology, providing an accurate geometric basis for subsequent light field modeling and overcoming the limitation of traditional 2D images lacking depth information.

[0089] Upon detecting the edge of the synchronization signal, the brain-computer interface device immediately activates the multi-channel bioelectrical signal acquisition system. Using a dry electrode array attached to the forehead, it synchronously records raw EEG signals from eight channels at a sampling rate of at least 256 Hz. The signal conditioning circuit amplifies and filters the acquired microvolt-level signals, and the analog-to-digital converter converts the analog signals into digital sequences. The data encapsulation module then aligns the EEG signals with data from the vision, environment, and depth sensors along a unified timeline, adds millisecond-accurate timestamps, and packages them into a standard-format multimodal data packet for transmission to the central processing unit.

[0090] It achieves precise spatiotemporal alignment of neural signals and multimodal sensing data, providing a reliable data foundation for studying the causal relationship between environmental stimuli and neural responses, and breaking through the technical bottleneck of isolated operation of each sensing module in traditional multimodal systems.

[0091] Specifically, the process involves dynamically constructing a digital light field using ambient light parameters and 3D environmental geometric information, and then rendering virtual makeup onto the user's facial image in real time using physically based rendering technology to generate a makeup effect image that blends with the ambient lighting. The specific steps are as follows:

[0092] The rendering engine receives ambient light parameters and 3D point cloud data from the data packet, converts the point cloud data into a 3D mesh model with normal vectors, and places a hemispherical ambient light with a corresponding color temperature in the virtual scene based on the light intensity and color temperature values, and determines the incident direction and intensity of the main light source based on the 3D mesh analysis.

[0093] The engine precisely maps the texture of the user's facial image onto the corresponding 3D mesh model and defines the material properties of the virtual makeup product, including base color, metallicity, roughness, and subsurface scattering coefficient for skin characteristics.

[0094] The physically based lighting shader is invoked, which calculates the diffuse reflection of ambient hemispherical light at various points on the facial surface and simulates specular highlights under the combined effects of metallicity and roughness. For foundation products, it also calculates the scattering of light beneath the skin surface to simulate a soft, translucent effect.

[0095] The facial area after color calculation is blended with the unmodified background area in real time, and a makeup effect image with correct light and shadow relationship under the current real ambient lighting conditions is output to the display device.

[0096] It should be noted that after receiving ambient light parameters and 3D point cloud data from the preprocessing flow via the data interface, the rendering engine first initiates the point cloud data processing flow. It uses the Poisson disk sampling algorithm to denoise and resample the original point cloud, and then uses the moving cube algorithm to convert the processed point cloud data into a 3D mesh model with continuous curved surfaces. During mesh generation, the system calculates the normal vector for each vertex, with the normal direction determined by the weighted average of adjacent triangular faces. After completing the geometric modeling, the rendering engine constructs a hemispherical ambient light mask covering the entire rendering area in the virtual scene based on the illuminance and color temperature values ​​provided by the ambient light sensor. The radiation intensity distribution on the inner surface of this light mask strictly follows the lighting characteristics of the real environment. Simultaneously, by analyzing the normal vector distribution characteristics of the 3D mesh model and combining it with ambient light intensity data, the system uses principal component analysis to calculate the incident angle and relative intensity of the main lighting directions in the scene, establishing a complete physical light source model for subsequent lighting calculations.

[0097] By accurately mapping the lighting characteristics of the real environment to the virtual rendering space, a lighting environment model that is completely consistent with the actual situation is established, providing an accurate lighting basis for subsequent physical rendering and effectively solving the problem of visual distortion caused by mismatched lighting conditions in traditional virtual makeup try-on.

[0098] After completing the 3D mesh modeling, the rendering engine initiates the texture mapping process. First, a feature point matching algorithm precisely aligns the 2D facial image with the 3D mesh model. Perspective transformation is used to correct image distortion. Then, UV unwrapping technology is employed to accurately project the facial texture onto the 3D mesh surface. During texture mapping, the system uses bilinear interpolation to ensure the complete preservation of texture details, while performing local optimization on specific areas such as the eye area and lips. After completing the basic texture mapping, the system creates a complete material description file for each virtual makeup product. This file contains basic color parameters based on physically based rendering, values ​​for metallic reflectivity, surface roughness levels, and subsurface scattering parameters specifically designed for the optical properties of human skin. These parameters collectively define the absorption and scattering behavior of light within skin tissue.

[0099] Through high-precision texture mapping and physically based optical parameter settings, a seamless integration of virtual makeup with the real face is achieved. In particular, by simulating the skin's unique subsurface scattering effect, the realism and naturalness of the makeup on the skin are significantly improved.

[0100] Once the material parameters are set, the system invokes the physically based rendering lighting shader. This shader first calculates the diffuse light intensity of each pixel using Lambert's cosine law, based on the ambient light intensity and direction of the hemispherical environment and the normal vector information of the mesh vertices. For makeup products with metallic properties, the shader, based on microsurface theory, calculates the impact of surface roughness on specular highlights using the GGX distribution function to generate realistic reflection effects. For products that cover the skin surface, such as foundation, the shader additionally activates a subsurface scattering calculation module. This module uses a diffusion approximation algorithm to simulate the multiple scattering process of light in the stratum corneum, epidermis, and dermis, calculating the propagation path and energy attenuation of light within the skin tissue, ultimately synthesizing a translucent texture effect unique to the skin.

[0101] By accurately simulating the physical interaction between light and skin and makeup materials, a qualitative leap has been achieved from simple color superposition to realistic optical performance. In particular, the simulation of skin translucency has enabled virtual makeup to achieve a visual fidelity close to that of real makeup application.

[0102] After completing the color calculations for the facial areas, the system activates the image compositing engine. First, it uses an edge detection algorithm to accurately identify the boundaries of the processed facial areas and employs feathering technology to eliminate jagged edges in the composite. During the image fusion stage, the system performs color matching and brightness adjustments on the rendered facial areas based on the lighting characteristics of the original background image, ensuring the overall consistency of the composite image. Simultaneously, the system uses depth buffer testing to ensure the correct spatial relationship between the virtual makeup and the real-world scene, avoiding any noticeable discrepancies. The final generated makeup effect image, after color space conversion, is output to the display device as a video stream. The entire process is completed within milliseconds of latency, enabling real-time user interaction.

[0103] Through precise image fusion technology and a real-time rendering pipeline, a seamless integration of virtual makeup with the real environment is achieved, ensuring visual consistency of the makeup try-on effect in the user's actual usage scenario, and greatly enhancing the practical value and user experience of the virtual makeup try-on system.

[0104] Specifically, when a user views a makeup trial image, the brain-computer interface device decodes the electroencephalogram (EEG) signals to quantify the user's subconscious preference score for the current makeup trial effect. The specific steps are as follows:

[0105] The system uses the precise moment when the makeup trial effect image is displayed on the screen as an event marker, and adds an event tag to the continuous EEG signal stream;

[0106] The EEG signal processing module takes the event marker as the starting point, extracts the EEG signal segment within a specific time window, and uses a blind source separation algorithm to filter out physiological artifacts in the signal caused by blinking, eye movement and muscle activity.

[0107] Temporal locked-means analysis was performed on the processed pure EEG signals to extract event-related potential components associated with cognitive assessment and emotional response, and their average amplitude within a specific time window was measured.

[0108] The measured event-related potential amplitudes, especially the amplitudes of the P300 component, are converted into a scalar value between 0 and 1 through a pre-calibrated linear mapping function. This value is the score that characterizes the intensity of the user's subconscious preference.

[0109] It should be noted that the moment the makeup trial image completes frame buffer refresh on the display device and begins to be presented to the user, the system's graphics processing unit generates a vertical synchronization signal. This signal is captured in real time and used as a precise time reference point. The system's event tagging module then creates a uniquely identified event descriptor at this point in time, containing information such as a timestamp, the makeup trial image number, and the event type. Simultaneously, the EEG signal acquisition system continuously records the raw EEG data stream at a high sampling rate. Upon receiving an event tagging command, it immediately inserts the event tag at the corresponding position in the data stream. To ensure time accuracy, the system uses hardware-level interrupts to process event tagging, directly transmitting the event signal to the event channel of the EEG acquisition device through a dedicated digital input / output interface. The time error throughout the process is controlled within milliseconds. The system also records the pipeline delay from image rendering completion to actual display and performs corresponding compensation and correction on the event timestamp.

[0110] By establishing a precise temporal correspondence between visual stimuli and neural responses, a reliable time reference is provided for subsequent EEG signal analysis, ensuring that specific EEG components generated when users view makeup trial images can be accurately captured.

[0111] The signal processing module first extracts multi-channel EEG signal segments from a continuous EEG data stream, covering a specific time period from before to after the event, based on the event's marked time point. Next, the system uses independent component analysis (ICA), a blind source separation technique, to process the extracted signal segments. By calculating the higher-order statistical properties between the signals from each channel, the mixed EEG signals are decomposed into multiple statistically independent components. Then, the system uses a pre-trained classification model to automatically identify physiological artifacts within these independent components, particularly those feature patterns related to blinking, eye movements, and head muscle activity. The identified artifacts are subtracted from the original signal, while retaining the neural signal components related to cognitive activity. The entire process employs a sliding window technique for real-time processing, ensuring that artifact removal of the current segment is completed before the next makeup trial image is displayed.

[0112] Advanced signal separation technology effectively removes various interfering components from EEG signals, significantly improving the signal-to-noise ratio of neural signals and providing a clean data foundation for subsequent cognitive response analysis.

[0113] After obtaining clean EEG signals free of artifacts, the system initiates an event-related potential (ERP) analysis process. First, EEG signal segments of the same type from multiple trials are aligned according to the event occurrence time. These segments are then superimposed and averaged to enhance event-related EEG activity while reducing random noise. Next, the system focuses on analyzing event-related potential components occurring within specific timeframes after stimulus presentation, particularly the P300 component related to attention allocation and decision evaluation, and late-stage positive potential components related to emotional processing. For each target component, the system calculates the average voltage amplitude within its specific time window; for example, the P300 component is typically measured within a 300-500 ms time window after stimulus presentation. Simultaneously, the system calculates characteristic parameters such as latency and peak-to-peak value of these components, forming a complete component characteristic description.

[0114] By using time-domain locked averaging technology, neurophysiological indicators related to the aesthetic cognition process can be effectively extracted, providing objective biological markers for quantifying users' subconscious reactions to makeup trial effects.

[0115] The system inputs the measured event-related potential (ERP) characteristic parameters into a pre-established standardized transformation model. This model, trained on a large amount of user experimental data, establishes a correspondence between EEG component amplitudes and subjective preference scores through regression analysis. Specifically, the system first standardizes the raw EEP amplitudes to eliminate individual differences in baseline EEG activity. Then, these standardized amplitudes are input into a linear transformation function whose parameters are obtained through supervised learning, mapping neurophysiological indicators to a uniform preference intensity scale. Finally, the system applies a sigmoid function to constrain the output value to a range of 0 to 1, where 0 represents no preference, 1 represents strong preference, and intermediate values ​​represent different degrees of preference intensity. This final score is stored in the user database along with the makeup trial image information.

[0116] By establishing a quantitative mapping relationship between neural signals and subjective preferences, an objective measurement of users' subconscious aesthetic responses is achieved, providing a more reliable and sensitive preference indicator for makeup recommendation systems than traditional scoring.

[0117] Specifically, the step of inputting ambient light parameters and subconscious preference scores into a dynamic correlation model for analysis to generate color adjustment parameter instructions for optimizing the makeup trial effect involves the following steps:

[0118] The dynamic association model receives the current set of ambient light parameters and the calculated subconscious preference scores, and normalizes the ambient light parameters into a standardized feature space.

[0119] The core of the model is a pre-trained tensor decomposition network, which performs tensor multiplication operations on the normalized ambient light features and subconscious preference scores to uncover the deep nonlinear relationship between ambient lighting and user aesthetic preferences.

[0120] The network output layer generates an optimization vector for the current makeup color based on the current association mode. This vector indicates the direction and magnitude of adjustment to the brightness, red-green axis, and yellow-blue axis components of the color parameters in the CIELAB color space.

[0121] The system encapsulates the optimization vector into a machine-readable color adjustment parameter instruction, which explicitly specifies how to modify the shader parameters of the current virtual makeup.

[0122] It should be noted that the dynamic correlation model receives a set of ambient light parameters from the preceding processing flow via a data interface. This set includes environmental feature data across multiple dimensions, such as light intensity, color temperature, and the direction angle of the main light source, while also receiving subconscious preference scores obtained through neural signal decoding. The model first initiates a data preprocessing process to standardize the ambient light parameters, using a min-max normalization method to transform environmental parameters of different dimensions into a unified numerical range. Specifically, the system performs a linear transformation on each environmental feature dimension based on preset upper and lower threshold values, ensuring that its values ​​are distributed within a standard range of zero to one. For the color temperature parameter, the system uses a logarithmic transformation of the Kelvin scale to conform to the linear characteristics of human visual perception. After normalization, all ambient light parameters are integrated into a standardized multidimensional feature vector, which serves as the input data for subsequent tensor operations. The entire normalization process uses a sliding window statistical method to dynamically update the parameter range, ensuring that the system can adapt to changes in lighting conditions under different usage environments.

[0123] By mapping multi-dimensional environmental parameters to a unified standard feature space, the interference of different physical dimensions on model analysis is eliminated, a standardized input format is established for subsequent multimodal data fusion, and the stability and convergence efficiency of model training are significantly improved.

[0124] Tensor decomposition networks employ a Tucker decomposition architecture, whose core structure comprises three dimensions: ambient light features, subconscious preferences, and hidden factors. The network first combines the normalized ambient light feature vector with the subconscious preference score into a second-order tensor, projecting it onto a high-dimensional feature space through tensor multiplication. Internally, multiple decomposition kernel functions simultaneously extract multi-dimensional features from the input tensor, solving for the factor matrix under each modality using an alternating least squares algorithm. During training, the network learns the complex mapping relationship between ambient lighting conditions and aesthetic preferences. During inference, forward propagation decomposes the input environment-preference combination into a weighted sum of a series of rank-tensors. Each rank-tensor represents a specific environment-preference association pattern. The network's attention mechanism dynamically adjusts the weight coefficients of each pattern based on the current input, thereby capturing the nonlinear interaction between ambient lighting characteristics and user aesthetic decisions.

[0125] By using tensor decomposition networks to perform deep feature mining on multimodal data, we can discover complex, non-obvious correlations between ambient lighting conditions and user aesthetic preferences, providing a more accurate basis for color adaptive adjustment.

[0126] After the forward propagation process of the network is completed, the output layer receives activation signals from the hidden layers and maps these signals into a three-dimensional optimization vector through a fully connected layer. The three components of this vector correspond to the lightness (L-axis), red-green (a-axis), and yellow-blue (b-axis) of the CIELAB color space, respectively. The system uses a hyperbolic tangent activation function to constrain the output value of each component to a range of -1 to +1, where a positive value indicates that the color dimension parameter needs to be increased, a negative value indicates that the parameter needs to be decreased, and the absolute value indicates the intensity of the adjustment. Specifically, the adjustment of the lightness component considers the influence of ambient light intensity on color perception, the red-green axis component mainly responds to the color balance requirements brought about by color temperature changes, and the yellow-blue axis component compensates for the spectral characteristics of ambient light. The generation process of the optimization vector also introduces an attention mechanism, dynamically adjusting the optimization priority of each color dimension according to the importance of the current ambient light parameters.

[0127] By generating precise optimization vectors based on color science theory, complex aesthetic preferences are transformed into specific and executable color adjustment parameters, realizing the quantitative transformation from qualitative description of subjective preferences to objective color parameters.

[0128] The system combines and encapsulates 3D optimization vectors with metadata information to generate standardized color adjustment parameter instructions. The instruction data structure consists of three parts: an instruction header, a parameter body, and a checksum. The instruction header indicates the instruction type and timestamp; the parameter body records the floating-point values ​​of the three color dimension adjustments and the adjustment priority identifier; and the checksum ensures the integrity of data transmission. The system uses JSON format for instruction serialization to ensure compatibility across different rendering modules. The instructions also include shader parameter mapping rules, explicitly specifying the shader uniform variable name and adjustment formula corresponding to each optimization vector component. During instruction distribution, the system sends color adjustment instructions to the rendering pipeline via a message queue mechanism, while retaining a copy of the instructions for subsequent analysis and learning. The entire encapsulation process employs a lightweight data compression algorithm to ensure the real-time requirements of instruction transmission.

[0129] A reliable communication mechanism between intelligent decision-making and rendering execution was established through standardized instruction encapsulation, ensuring that color optimization intentions can be accurately transmitted to the rendering system, and achieving seamless integration from data analysis to visual presentation.

[0130] Specifically, the steps involve dynamically adjusting the makeup parameters in the try-on effect image according to color adjustment parameter instructions, and immediately feeding back the optimized try-on effect image to the user to form a closed-loop verification.

[0131] The parametric adjustment engine receives color adjustment parameter instructions and parses out the color space adjustment vectors contained therein;

[0132] The engine directly accesses the global variables of the currently used makeup shader and modifies the parameters that affect color performance in real time based on the adjustment vector, such as increasing the brightness component of the base color or decreasing the saturation component.

[0133] The changes take effect immediately. The physical rendering engine uses the updated shader parameters to quickly re-render the facial area of ​​the current frame. This process skips the time-consuming light field reconstruction and only updates the color and lighting locally.

[0134] The optimized makeup effect image is presented to the user in a very short delay, replacing the previous frame, thus completing a rapid closed-loop verification from analysis to execution, and is ready to receive new neural feedback from the user.

[0135] It should be noted that the parametric adjustment engine receives color adjustment parameter instructions from the intelligent decision-making module through an internal message queue. These instructions are encapsulated in a lightweight data exchange format, containing complete color space adjustment vectors and related metadata. The engine first decrypts and deserializes the instructions to verify their integrity and validity, ensuring that the data has not been damaged or tampered with during transmission. Subsequently, the parsing module extracts the color space adjustment vector from the instructions. This vector exists as a three-dimensional floating-point array, corresponding to adjustment parameters in different dimensions of the color space. During parsing, the system checks whether the values ​​of each parameter are within predetermined thresholds and handles abnormal values ​​securely. After parsing, the engine converts the adjustment vector into an internal data structure, while recording the received timestamp and instruction sequence number, providing data support for subsequent performance analysis and error tracking. The entire parsing process employs a pipelined architecture, ensuring instruction processing is completed within milliseconds.

[0136] The standardized instruction parsing process ensures the accuracy and reliability of color adjustment parameters transmission between various modules of the system, providing precise data input for subsequent real-time rendering and effectively avoiding rendering anomalies caused by data parsing errors.

[0137] After successfully parsing the color adjustment parameters, the parametric adjustment engine directly accesses the currently running makeup shader program through the graphics application programming interface (API). The engine first obtains the addresses of the shader program's global variables, and then dynamically modifies key parameters affecting color performance based on the specific values ​​of the adjustment vectors. Specifically, the engine adjusts global variables such as basic hue, saturation, and brightness in the shader. For makeup products with special optical properties, it also adjusts advanced material parameters such as metallic reflectivity and surface roughness. During parameter modification, the engine uses atomic operations to ensure the integrity of data updates, avoiding inconsistencies during rendering. Simultaneously, the system records the modification history of all parameters, supporting parameter rollback and version management. The entire parameter update process is completed in the shared memory of the graphics processor, ensuring that modifications take effect immediately.

[0138] By directly manipulating global shader variables, the makeup parameters can be dynamically adjusted in real time, avoiding the performance overhead caused by recompiling the shader program in the past, and significantly improving the system's response speed and interactive experience.

[0139] Once the shader parameters are updated, the physically based rendering engine immediately initiates a fast re-rendering process. This process first identifies the facial regions that need updating, determining the rendering range through stencil and depth testing to avoid re-rendering background areas that don't require updating. During rendering, the system fully utilizes the light field data and geometric information already constructed in previous steps, recalculating only pixels affected by color parameter changes. For local updates to color and lighting, the rendering engine employs incremental rendering technology, calculating only the differences caused by parameter changes based on the rendering results of the previous frame. Simultaneously, the engine applies temporal reprojection technology, reusing the rendering results of unaffected pixels from the previous frame to further optimize rendering performance. The entire re-rendering process is executed in the parallel computing units of the graphics processor, ensuring that the image update is completed in a very short time.

[0140] By optimizing the rendering process and adopting an incremental update strategy, rendering efficiency has been significantly improved while ensuring visual effects. This enables real-time adjustment and instant feedback of makeup parameters, providing users with a smooth interactive experience.

[0141] After rapid re-rendering, the system immediately sends the optimized makeup trial image to the display pipeline. The image compositing module first seamlessly blends the newly rendered facial area with the original background, applying edge smoothing algorithms to eliminate visual imperfections at the seams. Subsequently, the composite image undergoes color space conversion and gamma correction to ensure consistent visual effects across different display devices. The system outputs the final image to the screen via the display controller, simultaneously recording the exact timestamp of the image display. While the image is presented to the user, the system sends a trigger signal to the neural signal acquisition module, initiating a new round of EEG signal acquisition to prepare data for subsequent preference analysis. The entire display process employs vertical synchronization technology to avoid screen tearing and ensure the stability and smoothness of the visual output.

[0142] By establishing a complete closed-loop verification mechanism, a complete cycle from parameter adjustment to visual feedback and then to neural signal acquisition was achieved, providing real-time data support for the continuous optimization of the system and significantly improving the accuracy and personalization of recommendations.

[0143] Specifically, the incremental learning of the dynamic association model based on data generated from multiple interactions to continuously evolve and optimize the model's recommendation decision-making ability involves the following steps:

[0144] The system encapsulates each complete interaction loop, including ambient light parameters, color adjustments performed, and the resulting changes in the user's subconscious preference scores, into a timestamped training sample and stores it in the local cache.

[0145] When the number of cached samples reaches a preset threshold, the system initiates a secure federated learning client process, which uses these new samples on the local device to calculate the update gradient of the model parameters.

[0146] The calculated gradients are uploaded to a central server, which aggregates gradient updates from multiple anonymous users and uses the aggregated gradients to perform an iterative optimization of the shared dynamic correlation model without uploading any original user data.

[0147] The system periodically downloads the latest version of model parameters from the server and completes local updates, enabling the model to continuously adapt to the ever-changing aesthetic trends and personal preferences of the user group, thereby evolving its decision-making capabilities.

[0148] It should be noted that after each complete interaction loop, the system automatically initiates a data encapsulation process, integrating key data such as ambient light parameters, executed color adjustment parameters, and corresponding changes in user subconscious preference scores generated during the current interaction into a structured training sample. This training sample is organized using a specific data format, comprising three parts: sample header information, feature data area, and label data area. The sample header information records a timestamp accurate to milliseconds, a session identifier, and a device identifier; the feature data area stores normalized ambient light parameter vectors and color adjustment parameter vectors; and the label data area stores the relative changes in user subconscious preference scores. The system uses a lightweight database engine to store these training samples in chronological order in a local cache area, while simultaneously establishing an indexing mechanism to support efficient subsequent queries. During storage, the system performs data integrity checks to ensure the integrity and accuracy of each training sample. The cache management system periodically cleans up expired data and employs data compression algorithms to optimize storage space utilization.

[0149] Through systematic data encapsulation and storage management, a complete interactive history record system has been established, providing a high-quality source of training data for incremental model learning, while ensuring the timeliness and integrity of data use.

[0150] When the number of training samples in the local cache reaches a preset threshold, the system automatically triggers the federated learning client process. This process first preprocesses the training samples in the cache, including data cleaning, feature standardization, and sample shuffle operations to improve training performance. Then, the client initializes the model training environment on the local device, loading the current dynamically correlated model parameters as the starting point for training. The training process uses the stochastic gradient descent algorithm, performing multiple iterations on the local dataset to calculate the updated gradients of the model parameters. Throughout the training process, the system implements a differential privacy protection mechanism, adding calibrated random noise to prevent information leakage from the original data. After training is complete, the client prunes and quantizes the calculated gradients to ensure they meet the receiving specifications of the federated learning server, and generates a quality report of the training process.

[0151] By completing the key computational steps of model training on local devices, local processing of user data is achieved, fundamentally avoiding the risk of leakage of raw privacy data and providing technical protection for secure distributed machine learning.

[0152] After completing local gradient calculations, the system uploads the processed gradient data to the central server via an encrypted communication channel. Uploaded data packets undergo digital signature and encryption to ensure the security and integrity of the transmission process. Upon receiving gradient updates from multiple clients, the server first performs authentication and data decryption, then uses a secure multi-party aggregation algorithm to fuse these gradients. During aggregation, the server calculates a weighted average of the gradients based on each client's data quality metrics and contribution, while simultaneously detecting and filtering for possible anomalous gradient updates. After aggregation, the server uses the aggregated gradients to perform an iterative optimization of the shared dynamic correlation model, updating the model's weight parameters. The entire optimization process employs momentum acceleration and adaptive learning rate techniques to ensure the stability and efficiency of model convergence.

[0153] The secure gradient aggregation mechanism effectively integrates knowledge from multiple users, significantly improving the model's generalization ability and robustness while protecting individual privacy, and overcoming the limitations of limited data from a single user.

[0154] The system periodically sends version query requests to the central server according to a preset time strategy to check if updated model parameters are available for download. When a new model version is detected, the system initiates a secure download process, verifying the model file's digital signature to ensure its authenticity and integrity. After downloading, the system creates a new model runtime environment on the local device and gradually deploys the updated model parameters to the inference engine. During model switching, the system employs a smooth transition strategy, setting a transition period between the old and new versions and verifying the new model's performance through A / B testing. Simultaneously, the system retains a backup of the old model version for rapid recovery in case of performance regression. After the model update is complete, the system automatically cleans up temporary files and updates the version history, ensuring the reliability and traceability of the entire update process.

[0155] By establishing a systematic model update mechanism, the recommendation system has achieved continuous evolution and self-optimization, enabling it to dynamically adapt to ever-changing user aesthetic trends and personalized needs, thus maintaining long-term service quality.

[0156] This embodiment also provides an intelligent makeup recommendation and try-on system based on multimodal perception, including:

[0157] The synchronous acquisition module allows the system to synchronously trigger a high dynamic range camera, an ambient light sensor, a depth sensing camera, and a brain-computer interface device to acquire multimodal data packets containing user facial images, ambient light parameters, three-dimensional environmental geometric information, and raw electroencephalogram (EEG) signals.

[0158] The light field rendering module dynamically constructs a digital light field using ambient light parameters and 3D environmental geometric information, and renders virtual makeup onto the user's facial image in real time based on physical rendering technology, generating a makeup effect image that blends with the ambient light.

[0159] The neural decoding module decodes the electroencephalogram (EEG) signals collected by the brain-computer interface device when the user views the makeup trial effect image, in order to quantify the user's subconscious preference score for the current makeup trial effect.

[0160] The intelligent decision-making module inputs ambient light parameters and subconscious preference scores into a dynamic correlation model for analysis, generating color adjustment parameter instructions to optimize the makeup trial effect;

[0161] The closed-loop optimization module dynamically adjusts the makeup parameters in the makeup trial effect image according to the color adjustment parameter instructions, and immediately feeds back the optimized makeup trial effect image to the user to form a closed-loop verification.

[0162] The continuous learning module performs incremental learning on the dynamic association model based on data generated from multiple interactions, so as to continuously evolve and optimize the model's recommendation decision-making capabilities.

[0163] This embodiment also provides a computer device applicable to the intelligent makeup recommendation and try-on method based on multimodal perception, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to realize the intelligent makeup recommendation and try-on method based on multimodal perception as proposed in the above embodiment.

[0164] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0165] This embodiment also provides a storage medium storing a computer program. When executed by a processor, the program implements the intelligent makeup recommendation and try-on method based on multimodal perception as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0166] In summary, this invention achieves precise adaptation and dynamic optimization of virtual makeup try-on effects under real-world lighting by constructing a deep fusion mechanism of ambient light field and neural feedback, significantly enhancing the realism and credibility of the try-on results; it overcomes the subjectivity and limitations of traditional explicit feedback by decoding users' subconscious preferences through brain-computer interfaces, making makeup recommendations more accurately match users' deep aesthetic needs; by establishing a dynamic association model and forming a closed-loop system of "perception-decision-optimization-verification," it can respond to changes in users and the environment in real time, achieving personalized adaptive adjustments; furthermore, by leveraging a federated learning framework for incremental learning, it continuously improves the system's generalization ability and intelligence level while fully protecting user data privacy, ultimately constructing an intelligent makeup recommendation and try-on ecosystem with self-evolutionary capabilities.

[0167] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for intelligent makeup recommendation and try-on based on multimodal perception, characterized in that, Includes the following steps: The system synchronously triggers a high dynamic range camera, an ambient light sensor, a depth sensing camera, and a brain-computer interface device to collect multimodal data packets containing user facial images, ambient light parameters, three-dimensional environmental geometric information, and raw electroencephalogram (EEG) signals. A digital light field is dynamically constructed using ambient light parameters and 3D environmental geometric information, and virtual makeup is rendered onto the user's facial image in real time based on physical rendering technology, generating a makeup effect image that blends with the ambient light. While users view the makeup trial effect images, the brain signals collected by the brain-computer interface device are decoded to quantify the user's subconscious preference score for the current makeup trial effect. Ambient light parameters and subconscious preference scores are input into a dynamic correlation model for analysis, generating color adjustment parameter instructions for optimizing the makeup trial effect; The makeup parameters in the trial makeup effect image are dynamically adjusted according to the color adjustment parameter instructions, and the optimized trial makeup effect image is immediately fed back to the user to form a closed loop verification. Incremental learning is performed on the dynamic association model based on data generated from multiple interactions, so as to continuously evolve and optimize the model's recommendation decision-making ability.

2. The intelligent makeup recommendation and try-on method based on multimodal perception as described in claim 1, characterized in that: The system acquires multimodal data packets containing user facial images, ambient light parameters, three-dimensional environmental geometric information, and raw electroencephalogram (EEG) signals by synchronously triggering a high dynamic range camera, an ambient light sensor, a depth-sensing camera, and a brain-computer interface device. The specific steps are as follows: The system generates a global hardware synchronization signal, which is simultaneously sent to all participating sensor control units via wired interrupt or wireless broadcast. After receiving the synchronization signal, the high dynamic range camera performs an exposure, captures multiple frames of images including overexposed and underexposed areas, and combines them into a complete facial image. At the same time, the ambient light sensor reads the current ambient light intensity and color temperature values. When triggered by a synchronization signal, the depth-sensing camera uses structured light or time-of-flight principles to project invisible light patterns onto the user's face and surrounding environment and receive their deformation. This allows it to calculate the depth value of each pixel in the scene and construct three-dimensional point cloud data. The brain-computer interface device uses this synchronization signal as the zero point of time and begins to continuously record the user's raw electroencephalogram (EEG) signals from the prefrontal cortex at a sampling rate of no less than 256 Hz. It also packages the data acquired by all sensors within the same clock cycle into a multimodal data packet with a unified timestamp.

3. The intelligent makeup recommendation and try-on method based on multimodal perception as described in claim 2, characterized in that: The process involves dynamically constructing a digital light field using ambient light parameters and 3D environmental geometric information, and then rendering virtual makeup onto the user's facial image in real time using physically based rendering technology to generate a makeup effect image that blends with the ambient lighting. The specific steps are as follows: The rendering engine receives ambient light parameters and 3D point cloud data from the data packet, converts the point cloud data into a 3D mesh model with normal vectors, and places a hemispherical ambient light with a corresponding color temperature in the virtual scene based on the light intensity and color temperature values, and determines the incident direction and intensity of the main light source based on the 3D mesh analysis. The engine precisely maps the texture of the user's facial image onto the corresponding 3D mesh model and defines the material properties of the virtual makeup product, including base color, metallicity, roughness, and subsurface scattering coefficient for skin characteristics. The physically based lighting shader is invoked, which calculates the diffuse reflection of ambient hemispherical light at various points on the facial surface and simulates specular highlights under the combined effects of metallicity and roughness. For foundation products, it also calculates the scattering of light beneath the skin surface to simulate a soft, translucent effect. The facial area after color calculation is blended with the unmodified background area in real time, and a makeup effect image with correct light and shadow relationship under the current real ambient lighting conditions is output to the display device.

4. The intelligent makeup recommendation and try-on method based on multimodal perception as described in claim 3, characterized in that: The process involves decoding the electroencephalogram (EEG) signals collected by the brain-computer interface device while the user views the makeup trial effect image to quantify the user's subconscious preference score for the current makeup trial effect. The specific steps are as follows: The system uses the precise moment when the makeup trial effect image is displayed on the screen as an event marker, and adds an event tag to the continuous EEG signal stream; The EEG signal processing module takes the event marker as the starting point, extracts the EEG signal segment within a specific time window, and uses a blind source separation algorithm to filter out physiological artifacts in the signal caused by blinking, eye movement and muscle activity. Temporal locked-means analysis was performed on the processed pure EEG signals to extract event-related potential components associated with cognitive assessment and emotional response, and their average amplitude within a specific time window was measured. The measured event-related potential amplitudes, especially the amplitudes of the P300 component, are converted into a scalar value between 0 and 1 through a pre-calibrated linear mapping function. This value is the score that characterizes the intensity of the user's subconscious preference.

5. The intelligent makeup recommendation and try-on method based on multimodal perception as described in claim 4, characterized in that: The process involves inputting ambient light parameters and subconscious preference scores into a dynamic correlation model for analysis, generating color adjustment parameter instructions to optimize the makeup trial effect. The specific steps are as follows: The dynamic association model receives the current set of ambient light parameters and the calculated subconscious preference scores, and normalizes the ambient light parameters into a standardized feature space. The core of the model is a pre-trained tensor decomposition network, which performs tensor multiplication operations on the normalized ambient light features and subconscious preference scores to uncover the deep nonlinear relationship between ambient lighting and user aesthetic preferences. The network output layer generates an optimization vector for the current makeup color based on the current association mode. This vector indicates the direction and magnitude of adjustment to the brightness, red-green axis, and yellow-blue axis components of the color parameters in the CIELAB color space. The system encapsulates the optimization vector into a machine-readable color adjustment parameter instruction, which explicitly specifies how to modify the shader parameters of the current virtual makeup.

6. The intelligent makeup recommendation and try-on method based on multimodal perception as described in claim 5, characterized in that: The process involves dynamically adjusting the makeup parameters in the virtual makeup trial image based on color adjustment parameter instructions, and immediately feeding back the optimized virtual makeup trial image to the user to form a closed-loop verification. The specific steps are as follows: The parametric adjustment engine receives color adjustment parameter instructions and parses out the color space adjustment vectors contained therein; The engine directly accesses the global variables of the currently used makeup shader and modifies the parameters that affect color performance in real time based on the adjustment vector, such as increasing the brightness component of the base color or decreasing the saturation component. The changes take effect immediately. The physical rendering engine uses the updated shader parameters to quickly re-render the facial area of ​​the current frame. This process skips the time-consuming light field reconstruction and only updates the color and lighting locally. The optimized makeup effect image is presented to the user in a very short delay, replacing the previous frame, thus completing a rapid closed-loop verification from analysis to execution, and is ready to receive new neural feedback from the user.

7. The intelligent makeup recommendation and try-on method based on multimodal perception as described in claim 6, characterized in that: The incremental learning of the dynamic association model based on data generated from multiple interactions, in order to continuously evolve and optimize the model's recommendation decision-making ability, involves the following steps: The system encapsulates each complete interaction loop, including ambient light parameters, color adjustments performed, and the resulting changes in the user's subconscious preference scores, into a timestamped training sample and stores it in the local cache. When the number of cached samples reaches a preset threshold, the system initiates a secure federated learning client process, which uses these new samples on the local device to calculate the update gradient of the model parameters. The calculated gradients are uploaded to a central server, which aggregates gradient updates from multiple anonymous users and uses the aggregated gradients to perform an iterative optimization of the shared dynamic correlation model without uploading any original user data. The system periodically downloads the latest version of model parameters from the server and completes local updates, enabling the model to continuously adapt to the ever-changing aesthetic trends and personal preferences of the user group, thereby evolving its decision-making capabilities.

8. A multimodal perception-based intelligent makeup recommendation and try-on system, based on the multimodal perception-based intelligent makeup recommendation and try-on method according to any one of claims 1 to 7, characterized in that: include, The synchronous acquisition module allows the system to synchronously trigger a high dynamic range camera, an ambient light sensor, a depth sensing camera, and a brain-computer interface device to acquire multimodal data packets containing user facial images, ambient light parameters, three-dimensional environmental geometric information, and raw electroencephalogram (EEG) signals. The light field rendering module dynamically constructs a digital light field using ambient light parameters and 3D environmental geometric information, and renders virtual makeup onto the user's facial image in real time based on physical rendering technology, generating a makeup effect image that blends with the ambient light. The neural decoding module decodes the electroencephalogram (EEG) signals collected by the brain-computer interface device when the user views the makeup trial effect image, in order to quantify the user's subconscious preference score for the current makeup trial effect. The intelligent decision-making module inputs ambient light parameters and subconscious preference scores into a dynamic correlation model for analysis, generating color adjustment parameter instructions to optimize the makeup trial effect; The closed-loop optimization module dynamically adjusts the makeup parameters in the makeup trial effect image according to the color adjustment parameter instructions, and immediately feeds back the optimized makeup trial effect image to the user to form a closed-loop verification. The continuous learning module performs incremental learning on the dynamic association model based on data generated from multiple interactions, so as to continuously evolve and optimize the model's recommendation decision-making capabilities.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the intelligent makeup recommendation and try-on method based on multimodal perception as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the intelligent makeup recommendation and try-on method based on multimodal perception as described in any one of claims 1 to 7.