A method, device and medium for generating content based on user images

By collecting user facial images and environmental features and establishing visual image personality vectors, the problems of poor emotion perception accuracy and weak visual preference modeling in existing technologies are solved. The generated images are more personalized and emotionally consistent, improving the accuracy of user state perception and the adaptability of generated content.

CN120374785BActive Publication Date: 2025-09-19BEIJING CENT BIOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510875894.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-09-19
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Existing content generation methods based on user images have poor emotion perception accuracy, weak visual preference modeling capabilities, and low personalization of generated content, making it difficult to accurately capture the user's current status and preferences.

Method used

By collecting user facial images and environmental features, extracting facial key points, hairstyle features, expression intensity parameters and posture features, and combining them with color stimulus image sequences, we establish facial expression response curves and color physiological response curves, generate visual image personality vectors, and drive the image generation model through the image style parameter mapping network to output personalized images.

Benefits of technology

It achieves highly accurate perception of the user's subjective state, and the generated images are more personalized, adaptable, and emotionally consistent, enhancing the personalization and emotional matching of the generated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374785B_ABST
    Figure CN120374785B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image generation, and in particular to a method, device and medium for generating content based on user images. The method comprises: collecting user facial images and environmental features, extracting the user's facial key points, hairstyle features, expression intensity parameters, and posture features in the facial images, and generating an image interaction state vector in combination with the environmental features; presenting a sequence of color stimulation images containing different tones to the user, collecting the user's facial micro-expressions, establishing a facial expression response curve, and further generating a color physiological response curve; extracting specified feature parameters based on the color physiological response curve, and constructing a visual image personality vector in combination with the image interaction state vector; inputting the visual image personality vector into an image style parameter mapping network to obtain an image style control parameter set, and an image generation and rendering module processing a three-dimensional reconstructed model of the user's face according to the image style control parameter set, and outputting a static image or a dynamic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image generation, and in particular to a method, device, and medium for generating content based on user images. Background Art

[0002] With the rapid development of artificial intelligence and image processing technologies, personalized content generation has gradually become a research hotspot in the fields of image synthesis and human-computer interaction. Within the technical framework related to image generation and visual expression, user personality modeling, image style control, and personalized synthesis models have become key technical paths. In particular, in the field of computer vision image generation (G06T), technologies that extract semantic information from user image data and drive image generation are widely used in scenarios such as virtual avatar construction, digital human generation, and immersive interaction.

[0003] Existing technologies for generating personalized image content often rely on static user profiles or user-selected style templates, making it difficult to accurately capture the user's current, real-time state and underlying preferences. For example, some methods achieve image style generation through facial recognition and style matching recommendations. While these methods can provide a certain degree of visual style transformation, they struggle to extract deeper psychological or physiological characteristics from multi-dimensional features such as facial micro-expressions, posture changes, and emotional reactions, resulting in insufficient personalization and matching of generated content. Furthermore, existing image style transfer or generation methods often construct style parameters based on static images or preset labels, lacking modeling of user reactions under varying environmental perceptual conditions (such as lighting and color temperature), making it difficult to construct dynamic content generation mechanisms that are closely linked to the user's current state. While recent advances in emotion computing and micro-expression recognition have provided new technical support for capturing users' true emotions and psychological preferences, this information remains underutilized in existing image generation systems.

[0004] For example, some solutions, such as patent CN112164135A (Main Classification Number G06T), propose a device and method for constructing a virtual character. However, these methods suffer from significant deficiencies in emotion perception, visual preference modeling, and personalized content generation. This solution lacks an emotion recognition mechanism, making it unable to accurately capture a user's emotional state; it relies solely on subjective descriptions to select facial features, lacking the ability to model a user's visual preferences; and the generated content is fixed, unable to dynamically adjust the results based on the user's image and emotions. This results in a low level of personalization and struggles to meet the diverse needs of users.

[0005] Another approach, patent CN116433800B (main classification number G06T), proposes an image generation method based on the joint guidance of user preferences and text in social scenarios. This method combines the image generation methods of user preferences and text in social scenarios, but it lacks the real-time and precision of emotion perception and personalized content generation. Its user preference modeling relies on social relationships and image interaction history, making it difficult to accurately capture the user's current emotional state and lacking the ability to perceive real-time user images and facial emotional features. In addition, the generated content is limited by social network data and pre-trained models, making it difficult to achieve high-precision, in-depth personalized expression based on specific user images. Summary of the Invention

[0006] (1) Technical issues to be resolved

[0007] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present application provides a content generation method, device and medium based on user images, which solves the technical problems of the existing content generation methods based on user images, such as poor emotion perception accuracy, weak visual preference modeling capabilities and low degree of personalization of generated content.

[0008] (2) Technical solution

[0009] In order to achieve the above objectives, the main technical solutions adopted in this application include:

[0010] In a first aspect, an embodiment of the present application provides a method for generating content based on a user image, comprising the following steps:

[0011] S1. Capturing a user's facial image through an image acquisition device and collecting environmental features of the user through an environmental perception module, extracting the user's facial key points, hairstyle features, expression intensity parameters, and posture features from the facial image, and generating an image interaction state vector based on the environmental features;

[0012] S2. Presenting a sequence of color stimulus images of different tones to the user, and synchronously collecting the user's facial image data during the presentation process, extracting the user's facial micro-expressions from the facial image data, establishing a facial expression response curve, and further generating a color physiological response curve;

[0013] S3. extracting specified feature parameters based on the color physiological response curve, and fusing them with the image interaction state vector to generate a visual image personality vector for controlling the image style;

[0014] S4. Inputting the visual image personality vector into an image style parameter mapping network to obtain an image style control parameter set. The image generation and rendering module processes the pre-acquired 3D reconstruction model of the user's face according to the image style control parameter set according to the image processing process, and outputs the generated personalized static image or dynamic image.

[0015] The image processing process includes: color remapping and brightness adjustment, texture detection and refinement generation, edge enhancement and anti-aliasing processing, lighting simulation and shadow rendering, expression and posture animation processing.

[0016] Preferably, the S1 specifically includes:

[0017] S11, using an image acquisition device to capture a user's facial image;

[0018] S12, analyzing the collected user facial image through a posture recognition algorithm to extract the user's posture features;

[0019] The posture features include: head posture angle and body posture information;

[0020] S13, using a facial key point detection model based on a convolutional neural network to identify the locations of facial key points in the user's facial image;

[0021] The facial key points include: facial contour, eyes, nose, and mouth;

[0022] S14, using an image segmentation model to identify the hair area in the user's facial image and obtain a hairstyle contour feature vector;

[0023] S15, using an expression recognition model to identify expression action units contained in the user's facial image, and recording the intensity value of each expression action unit as an expression intensity parameter;

[0024] S16, aligning the user's posture features, facial key points, hairstyle contour feature vector, expression action unit intensity parameters, and environmental features according to a unified timestamp, inputting them into a feature embedding encoder for standardization, and generating a low-dimensional joint feature vector through a feature fusion network as an image interaction state vector representing the current user interaction state;

[0025] The environmental characteristics include: light intensity and color temperature.

[0026] Preferably, the S14 specifically includes:

[0027] S141, using an image segmentation model to process the collected user facial image, accurately extract the hair area, and generate a high-resolution hair mask image;

[0028] The image segmentation model is a multi-scale semantic segmentation network; S142, based on the hair mask image, a contour detection algorithm is used to extract the outer edge line of the hairstyle; S143, a geometric feature analysis is performed on the extracted hairstyle edge curve to extract contour parameters representing the hairstyle morphology, wherein the contour parameters include:

[0029] Contour length: the total length of the edge curve of the hairstyle;

[0030] Curvature: the curvature change value of the edge curve within a unit length, which is used to reflect the degree of tortuosity of the edge curve;

[0031] Closure: The degree of connectivity between the beginning and end of the edge curve, used to evaluate the possibility of forming a closed figure;

[0032] S144 , normalizing or standardizing the contour parameters to eliminate scale differences, and further combining the standardized contour parameters to form a hairstyle contour feature vector for comprehensively describing the shape category and structural features of the user's hairstyle.

[0033] Preferably, the facial expression action units include: an inner eyebrow raising action unit, an outer eyebrow raising action unit, a frowning action unit, an upper eyelid lifting action unit, a zygomatic muscle lifting action unit, an eyelid closing action unit, a nose wrinkling action unit, an upper lip lifting action unit, a mouth corner raising action unit, and a mouth corner lowering action unit;

[0034] The expression recognition model is a convolutional neural network model.

[0035] Preferably, the S2 specifically includes:

[0036] S21, continuously presenting a sequence of color stimulation images containing different hues to the user, and during the presentation of the color stimulation images, using an image acquisition device to collect facial image data of the user in real time;

[0037] S22, inputting the collected facial image data into an expression recognition model to extract the user's facial micro-expression features, wherein the facial micro-expression features include a plurality of recognized expression action units and their corresponding intensity values;

[0038] S23, establishing a set of facial expression response curves for the users based on the intensity values ​​of the expression action units at the corresponding time points and the frames of the color stimulus image sequence;

[0039] In the process of establishing the user's facial expression response curve, the intensity value of each expression action unit is used as the vertical axis, and the time sequence of the color stimulus image is used as the horizontal axis, and a facial expression response curve is established for each expression action unit separately;

[0040] S24. Based on the facial expression response curve corresponding to each expression action unit, a color physiological response curve reflecting the user under color stimulation conditions of different tones is constructed.

[0041] Preferably, the color physiological response curve is obtained by weighted fusion of the response curves of all expression action units according to preset weights.

[0042] Preferably, the S3 specifically includes:

[0043] S31, extracting specified characteristic parameters from the color physiological response curve;

[0044] The specified characteristic parameters include: response start time, maximum response intensity, and high level maintenance time;

[0045] The response start time is: with the time axis on the curve as the horizontal axis, the time point corresponding to when the intensity value exceeds the preset threshold for the first time;

[0046] Maximum response intensity is: the maximum intensity value in the color physiological response curve;

[0047] High maintenance time refers to the length of time during which the intensity value remains above the high intensity threshold;

[0048] The high intensity threshold is 80% of the maximum response intensity;

[0049] S32. Fusing the specified feature parameters with the image interaction state vector to form a combined feature vector, and using the combined feature vector as the visual image personality vector.

[0050] Preferably, the S4 specifically includes:

[0051] S41, inputting the visual image personality vector into an image style parameter mapping network to map it into an image style control parameter set;

[0052] S42: Inputting the image style control parameter set into an image generation and rendering module, wherein the image generation and rendering module processes the pre-acquired 3D reconstruction model of the user's face according to an image processing flow using an image generation model and the image style control parameter set, and outputs a generated personalized static image or dynamic image;

[0053] Among them, the image style parameter mapping network is a deep neural network, which is used to map the visual image personality vector into the style control parameters required for image generation, and the image generation model is a deep generation model based on a generative adversarial network or a diffusion model.

[0054] In a second aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, wherein:

[0055] The memory is used to store computer programs;

[0056] The processor is configured to execute the computer program to implement the above-mentioned method for generating content based on user images.

[0057] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is used to implement the above-mentioned method for generating content based on user images.

[0058] (3) Beneficial effects

[0059] The embodiment of the present application provides a content generation method based on user images. By collecting user facial images and environmental features, it can achieve comprehensive and synchronous collection and fusion of user posture, facial key points, expression intensity, hairstyle contour and environmental factors (such as lighting, color temperature, etc.), and then construct a unified expression vector of the current user state, effectively improving the accuracy and stability of the perception of the user's subjective state.

[0060] Furthermore, the present embodiment of the application uses color stimulation and facial micro-expression linkage to capture the user's facial micro-expressions in images of different tones to establish a detailed facial expression response curve and color physiological response curve, thereby capturing the user's physiological response characteristics to color stimulation, breaking through the limitations of traditional static expression recognition. By constructing a visual image personality vector that integrates interaction state and color response, the present embodiment of the application can accurately reflect the user's subjective preferences, perceptual characteristics, and visual emotional characteristics, and map them as input control parameters to the deep style generation network, making the generated image more personalized, adaptive, and emotionally consistent.

[0061] Furthermore, the method described in this application extracts hairstyle contours through a multi-scale image segmentation model and generates a structured description vector after normalization based on geometric features (such as curvature and closure), significantly improving the system's ability to express hairstyle patterns and individual characteristics. During the image generation phase, the visual image personality vector is converted into control parameters of the image generation model through an image style parameter mapping network. This can flexibly drive the automatic generation of highly customized static and dynamic images based on the image generation model. Even in situations where the user's state is complex, their expressions are subtle, or the environment is changing, the generated content can still be consistent with the user's actual emotions and preferences, ensuring a personalized, emotional, and highly compatible overall interactive experience.

[0062] In summary, the embodiments of the present application establish a standardized and expressive visual personality modeling and generation path by integrating multi-dimensional data such as user images, environmental perception and color response, which not only enhances the subjective fit and aesthetic relevance of the generated images. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1Schematic diagram of a flow chart of a method for generating content based on a user image according to an embodiment of the present application;

[0064] Figure 2 Flowchart of a process for generating an image interaction state vector according to one embodiment of the present application;

[0065] Figure 3 Flowchart of a process for generating a color physiological response curve according to one embodiment of the present application;

[0066] Figure 4 A flowchart of a process for constructing a visual image personality vector according to one embodiment of the present application;

[0067] Figure 5 The figure is a flowchart of a process for generating a corresponding static image or dynamic image according to one embodiment of the present application. DETAILED DESCRIPTION

[0068] In order to better explain the present application and facilitate understanding, the present application is described in detail below with reference to the accompanying drawings through specific implementation methods.

[0069] In the current field of image generation and personalized content presentation, especially in the task of generating content based on user images, existing technologies mainly have the following typical problems:

[0070] Traditional image style generation methods typically rely on fixed input image features (such as style and content images) for image synthesis. Some technologies use user images for avatar generation or cartoonization (such as CN113470147A, main classification number G06T). However, these methods lack in-depth modeling of individual user subjective characteristics such as emotions, physiology, and aesthetic preferences. For example, physiological and psychological feedback such as micro-expression changes and color preferences generated by users when presented with images of different tones, textures, and contexts is not effectively perceived or utilized.

[0071] Current methods often use facial landmark location and expression recognition to obtain static information of user facial images, but often ignore more style-discriminative information dimensions such as hairstyle outline and posture characteristics, resulting in the generated content lacking detailed personality in visual expression and making it difficult to accurately express the user's overall visual image.

[0072] The lack of a dynamic vector modeling mechanism for the "interaction state" formed by the combination of environmental features (such as lighting and color temperature) and the user's current posture and expression results in a disconnect between the generated image and the user's current real scene, lacking contextual consistency and interactive immersion. In image generation applications involving user emotional participation, emotional states are often modeled based on preset labels (such as "happy" and "sad") or text descriptions. This fails to build a dynamic physiological response model based on the relationship between "user facial micro-expressions and color stimulus images," making it impossible to truly portray the user's physiological response path to different color tones and form a personalized "visual personality." Because the user's facial key points, expression intensity, hairstyle geometry, posture, ambient lighting, and other features belong to different dimensions and semantic spaces, existing methods lack effective encoding structures and control strategies for unified embedding and style parameter mapping, resulting in an unstable style mapping process, insufficient individual differentiation in the generated results, and poor visual continuity.

[0073] In the existing field of image generation and recommendation, traditional methods mainly rely on user selection or preset parameters for content presentation, ignoring the implicit preferences reflected by users' physiological and emotional responses during actual interactions, resulting in limited personalization of generated content and low user acceptance.

[0074] To address these issues, this application proposes a method for generating content based on user images. By deeply integrating image processing, facial feature modeling, color physiological response modeling, and image style control, this method constructs a user modeling and image content generation system centered on a "visual image personality vector." This method falls within the field of image data processing (G06T) and specifically includes the following key technical steps:

[0075] The system uses an image acquisition device to capture user facial image data. This is combined with an environmental perception module to collect environmental factors such as light intensity and color temperature. A convolutional neural network is then used to extract visual features such as facial key points (e.g., eyes, nose, mouth), expression intensity, posture angle, and hairstyle outline. Through unified timestamp alignment and embedding network processing, a low-dimensional vector (i.e., image interaction state vector) is generated that reflects the user's current interaction state. This stage simultaneously addresses the synchronization of multi-source heterogeneous image features (expression, hairstyle, and posture) in the temporal domain, effectively modeling the user's current visual state.

[0076] The user is presented with a sequence of dynamic color stimulus images. During this process, the user's facial micro-expressions are simultaneously captured, and multiple action units (AUs) such as inner eyebrow raises and mouth corner raises are identified. A temporal intensity curve, known as the facial expression response curve, is then constructed. Furthermore, by weightedly fusing the responses of different AUs, a color physiological response curve is generated that reflects the user's physiological feedback to the color stimuli. This step establishes a technical chain for "user-color preference" modeling in the field of image data processing.

[0077] Based on image data feature processing, the system extracts specific characteristic parameters (such as response onset time, maximum intensity, and duration of high-level maintenance) from the color physiological response curve and fuses them with the aforementioned image interaction state vector to form a visual image personality vector that incorporates emotional responses. This vector, a compact representation of the user's content preferences in the current scenario, helps improve the accuracy of style matching in image generation tasks.

[0078] The generated visual image personality vector serves as input and is converted into style control parameters through a nonlinear mapping network. This is used to drive image generation models (such as StyleGAN and Diffusion) to produce static or dynamic images that meet the user's emotional preferences and personalized needs. This stage focuses on transforming the abstract user perception modeling results into concrete visual content, establishing an end-to-end closed-loop path from "image recognition" to "image generation."

[0079] To better understand the above technical solutions, exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a clearer and more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0080] Example 1

[0081] Figure 1 FIG. 1 is a flow chart of a method for generating content based on user images according to an embodiment of the present application. Figure 1 As shown, the content generation method based on user images includes:

[0082] S1. Capturing a user's facial image through an image acquisition device and collecting environmental features of the user through an environmental perception module, extracting the user's facial key points, hairstyle features, expression intensity parameters, and posture features from the facial image, and generating an image interaction state vector based on the environmental features;

[0083] Specifically, this embodiment can make subsequent content more suitable for the current psychological and physiological state of the user through the image interaction state vector.

[0084] See also Figure 2 In the practical application of this embodiment, S1 specifically includes:

[0085] S11, using an image acquisition device to capture a user's facial image;

[0086] S12, analyzing the collected user facial image through a posture recognition algorithm to extract the user's posture features;

[0087] The posture features include: head posture angle and body posture information;

[0088] Specifically, step S12 uses a sophisticated posture recognition algorithm to extract posture features such as the user's head pitch, yaw, and roll angles. This effectively determines the user's attention, mental state, and willingness to participate. For example, a significant degree of facial deviation may indicate a distracted user, while lowering the head or leaning back may indicate interaction fatigue. This data helps enhance dynamic perception of the user's state.

[0089] S13, using a facial key point detection model based on a convolutional neural network to identify the locations of facial key points in the user's facial image;

[0090] The facial key points include: facial contour, eyes, nose, and mouth;

[0091] This embodiment uses a deep learning model based on the CNN (Convolutional Neural Network) architecture to automatically extract two-dimensional or three-dimensional position information of key areas such as facial contours, eyes, nose, and mouth, thereby improving the accuracy and robustness of detection.

[0092] Facial landmark detection models based on convolutional neural networks (CNNs) are the most widely used and highly accurate facial structure recognition technology in computer vision. The core task of these models is to automatically identify and locate multiple semantically significant facial landmarks, such as the corners of the eyes, the ends of the eyebrows, the tip of the nose, and the corners of the mouth, from an input face image. These models are often used in applications such as constructing facial geometry, analyzing facial expressions, creating virtual filters, and recognizing emotions.

[0093] S14, using an image segmentation model to identify the hair area in the user's facial image and obtain a hairstyle contour feature vector;

[0094] The user's hairstyle is an important component of their visual style. Step S14 uses an image segmentation model (such as U-Net or DeepLab) to accurately segment the hair region in the image and extract its contour feature vector. This not only enhances comprehensive modeling of the user's appearance but also provides data support for maintaining hairstyle consistency in the generated results, ensuring that the generated image faithfully reproduces the user's original appearance.

[0095] S15, using an expression recognition model to identify expression action units contained in the user's facial image, and recording the intensity value of each expression action unit as an expression intensity parameter;

[0096] Specifically, step S15 identifies and quantifies the intensity of expression units (AUs) in the Facial Action Coding System (FACS), helping to accurately depict subtle changes in the user's current emotional state. Compared to traditional coarse classification methods (such as distinguishing only between joy, anger, sadness, and happiness), this method has higher emotional resolution and can effectively capture complex facial expressions. It is particularly suitable for micro-expression recognition scenarios, greatly improving sensitivity and responsiveness to emotional state.

[0097] It should be noted that expression recognition models are a type of artificial intelligence model used to automatically analyze facial expressions in facial images. They are widely used in fields such as emotional computing, human-computer interaction, virtual reality, and psychological analysis. In this technical solution, the expression recognition model identifies the action units (AUs) contained in the user's facial image and further extracts the intensity value of each AU to construct a quantitative representation of the user's current facial emotional state—the expression intensity parameter. The core basis of this expression recognition model is the Facial Action Coding System (FACS), which decomposes human facial expressions into multiple independent muscle action units, such as raising eyebrows, stretching the corners of the mouth, and squinting the eyes. Each AU reflects the movement state of a specific facial muscle group, and different combinations represent different emotions.

[0098] The facial expression action units include: an inner eyebrow raising action unit, an outer eyebrow raising action unit, a frowning action unit, an upper eyelid lifting action unit, a zygomatic muscle lifting action unit, an eyelid closing action unit, a nose wrinkling action unit, an upper lip lifting action unit, a mouth corner raising action unit, and a mouth corner lowering action unit; the facial expression recognition model is a convolutional neural network model.

[0099] For example, facial expression recognition models generally use deep convolutional neural networks (CNNs) as their foundational framework. By training on large amounts of data labeled with action units, they automatically identify each action unit in an image. For example, the model outputs the activation level values ​​corresponding to action units AU01 (inner eyebrow raise), AU06 (zygomatic muscle contraction), and AU12 (mouth corner raise) as "expression intensity parameters."

[0100] The typical expression recognition process includes the following steps:

[0101] Face detection and alignment: First, detect the faces in the image and standardize their orientation and size through methods such as affine transformation to ensure the accuracy of subsequent recognition results.

[0102] Use pre-trained CNN models (such as ResNet, MobileNet, EfficientNet) to process facial images, extract deep semantic features, and capture subtle differences in facial muscle movements.

[0103] Each action unit is identified and its strength assessed using a multi-label classification or regression network. Some systems use attention mechanisms or graph convolutional neural networks (GCNs) to simulate the interconnectedness between muscles, thereby improving recognition accuracy.

[0104] Finally, the numbers of multiple action units and their corresponding activation intensities are output to form a set of high-dimensional numerical features as an objective expression of the user's current facial expression state.

[0105] S16, aligning the user's posture features, facial key points, hairstyle contour feature vector, expression action unit intensity parameters, and environmental features according to a unified timestamp, inputting them into a feature embedding encoder for standardization, and generating a low-dimensional joint feature vector through a feature fusion network as an image interaction state vector representing the current user interaction state;

[0106] The environmental characteristics include: light intensity and color temperature.

[0107] Specifically, step S16 aligns multi-source heterogeneous features through a time synchronization mechanism to eliminate interference from data time delays on analysis results. Furthermore, a feature embedding encoder standardizes and vectorizes data of different dimensions and types, improving the effectiveness of data fusion. Finally, a feature fusion network (such as a Transformer or multi-layer perceptron) is used to generate a highly expressive and compact image interaction state vector, serving as an abstract representation of the user's current psychological and behavioral state. This vector offers advantages such as strong real-time performance, high information density, and good generalizability.

[0108] S2. Presenting a sequence of color stimulus images of different tones to the user, and synchronously collecting the user's facial image data during the presentation process, extracting the user's facial micro-expressions from the facial image data, establishing a facial expression response curve, and further generating a color physiological response curve;

[0109] In this embodiment, presenting a sequence of color stimulus images with varying hues to the user and simultaneously capturing their facial micro-expressions during viewing facilitates non-invasive quantification of their physiological responses to different colors. Facial micro-expressions are highly unconscious and instantaneous, and compared to traditional questionnaires or subjective scoring methods, they can more realistically reflect a user's underlying emotional preferences for color. By analyzing the dynamic trends of facial micro-expressions as they change with color, establishing a user facial expression response curve, and further extracting a color physiological response curve representing the user's visual emotional changes, this effectively enhances insight into the user's underlying psychological state.

[0110] Preferably, in some embodiments of the present application, see Figure 3 , the S2 specifically includes:

[0111] S21, continuously presenting a sequence of color stimulation images containing different hues to the user, and during the presentation of the color stimulation images, using an image acquisition device to collect facial image data of the user in real time;

[0112] By playing a sequence of images with diverse tones in a controlled environment, users' natural reactions to different colors are stimulated. Simultaneously, high-frame-rate image acquisition equipment is used to record users' facial image data in real time, ensuring that even brief facial muscle movements are captured. Unlike traditional psychological surveys or subjective feedback methods, this method collects information through real physiological reactions, which are less susceptible to conscious interference or masking by the user, resulting in higher accuracy and ecological validity.

[0113] S22, inputting the collected facial image data into an expression recognition model to extract the user's facial micro-expression features, wherein the facial micro-expression features include a plurality of recognized expression action units and their corresponding intensity values;

[0114] A high-precision facial expression recognition model analyzes each facial image frame, outputting multiple facial action units (AUs) and their intensity values. Because facial micro-expressions typically last between 0.1 and 0.5 seconds, are fleeting and difficult to control, a deep learning model is used for rapid, fine-grained analysis, effectively capturing the user's true emotional shifts. For example, when presented with a cold image, the model might detect uncomfortable movements such as eyebrow droop or orbicularis oculi muscle contraction, reflecting the user's underlying negative emotions.

[0115] S23, establishing a set of facial expression response curves for the users based on the intensity values ​​of the expression action units at the corresponding time points and the frames of the color stimulus image sequence;

[0116] In the process of establishing the user's facial expression response curve, the intensity value of each expression action unit is used as the vertical axis, and the time sequence of the color stimulus image is used as the horizontal axis, and a facial expression response curve is established for each expression action unit separately;

[0117] By matching the hue of each frame with the intensity of the facial action unit at that point in time, a response curve is constructed for each action unit, showing how its intensity changes over time. The curve, with time as the horizontal axis and action unit intensity as the vertical axis, reflects the impact of different hue images on facial muscle activity. This response curve not only provides a tool for visualizing the user's physiological responses, but also makes subsequent data analysis more consistent and interpretable.

[0118] S24. Based on the facial expression response curve corresponding to each expression action unit, a color physiological response curve reflecting the user under color stimulation conditions of different tones is constructed.

[0119] The color physiological response curve is obtained by weighted fusion of the response curves of all expression action units according to preset weights.

[0120] By weightedly merging multiple action unit response curves, a comprehensive and representative color physiological response curve is formed. Preset weights can be set based on the contribution of each action unit in expressing a specific emotion (such as joy, surprise, and disgust). For example, a raised corner of the mouth (AU12) is highly correlated with pleasure, so its weight can be relatively high. The resulting color physiological response curve not only reveals the user's emotional response patterns to different hues.

[0121] S3. extracting specified feature parameters based on the color physiological response curve and combining them with the image interaction state vector to construct a visual image personality vector that integrates emotional responses;

[0122] Based on the aforementioned color physiological response curve, specific feature parameters are extracted and combined with the image interaction state vector to construct a visual image personality vector that integrates emotional response characteristics and environmental state characteristics. This visual image personality vector not only incorporates the user's long-term preferences (such as color sensitivity and style preferences) but also encompasses the user's instantaneous emotional state in the current environment, achieving a bidirectional "personality + context" user modeling approach. The visual image personality vector possesses excellent expressiveness and generalizability, allowing it to be used for both immediate image style adjustment and to accumulate a user's long-term style profile, enhancing personalized service capabilities.

[0123] Preferably, in some embodiments of the present application, see Figure 4 , the S3 specifically includes:

[0124] S31, extracting specified characteristic parameters from the color physiological response curve;

[0125] The specified characteristic parameters include: response start time, maximum response intensity, and high level maintenance time;

[0126] The response start time is: with the time axis on the curve as the horizontal axis, the time point corresponding to when the intensity value exceeds the preset threshold for the first time;

[0127] In this embodiment, the response start time reflects the user's ability to react immediately to a certain color stimulus, representing their sensitivity. For example, the shorter the time required for red to trigger a change in facial expression, the more sensitive the user is to red.

[0128] Maximum response intensity is: the maximum intensity value in the color physiological response curve;

[0129] The maximum response intensity reflects the extreme value of emotional fluctuations and represents the user's maximum emotional reaction to a certain color tone. It can be used to distinguish the ability of stimulus colors to drive user emotions.

[0130] High maintenance time refers to the length of time during which the intensity value remains above the high intensity threshold;

[0131] The high intensity threshold is 80% of the maximum response intensity;

[0132] High maintenance time measures the persistence of users' high-intensity emotional responses, reflecting the emotion regulation mechanism or preference stability. For example, a long period of high response may indicate strong preference or aversion.

[0133] The specified feature parameters constitute an accurate characterization of the dynamic reaction process of the user's facial micro-expressions, which is far superior to the simple method of using only peak or average values. Therefore, it can significantly enhance the accuracy of user emotion modeling and the recognizability of personalized features.

[0134] S32. Fusing the specified feature parameters with the image interaction state vector to form a combined feature vector, and using the combined feature vector as the visual image personality vector.

[0135] In this embodiment, step S3 integrates emotional response features with interaction context, bridging the gap from "physiology-emotion" to "personality-behavior," effectively enhancing the depth and breadth of user modeling. The resulting visual image personality vector is more adaptable.

[0136] S4. Inputting the visual image personality vector into an image style parameter mapping network to obtain an image style control parameter set. The image generation and rendering module processes the pre-acquired 3D reconstruction model of the user's face according to the image style control parameter set according to the image processing process, and outputs the generated personalized static image or dynamic image.

[0137] The image processing process includes: color remapping and brightness adjustment, texture detection and refinement generation, edge enhancement and anti-aliasing processing, lighting simulation and shadow rendering, expression and posture animation processing.

[0138] Through the above processing, we can achieve deep matching and efficient mapping from the user's multi-dimensional image data to image style control parameters. The generated images are closer to the user's psychological needs in terms of color style, content elements, etc., significantly improving the quality of personalized generation. In summary, this embodiment combines image processing technology, computer vision technology, and user psychological state modeling to build an intelligent image generation system driven by user state.

[0139] Preferably, in some embodiments of the present application, see Figure 5 , the S4 specifically includes:

[0140] S41, inputting the visual image personality vector into an image style parameter mapping network to map it into an image style control parameter set (i.e., a set of style control parameters);

[0141] In step S41, the constructed visual image personality vector is input into the image style parameter mapping network for decoding. The image style parameter mapping network adopts a deep neural network architecture, and its typical structure may include fully connected layers, residual blocks, and a multi-head attention mechanism. It is used to learn the complex nonlinear mapping relationship between the visual image personality vector and the image style control parameters.

[0142] The introduction of the Image Style Parameter Mapping Network addresses the semantic gap in mapping psychological features to image control parameters. It not only decodes high-dimensional psychological feature vectors but also learns structural patterns in the style space through training, enabling the generation of highly personalized yet style-coordinated control parameters. This significantly improves the accuracy and interpretability of personalized image generation.

[0143] The image style parameter mapping network is a deep neural network, which is used to map the visual image personality vector into the style control parameters required for image generation.

[0144] The Image Style Parameter Mapping Network (ISPN) is a key module based on deep neural networks. Its core function is to convert the input visual image personality vector into the style control parameters required for image generation. The ISPN combines the user's long-term visual preferences (such as color sensitivity and style orientation) with their current interactive emotional state (such as calmness, excitement, and anxiety), and is rich in semantic information and psychological cues. The ISPN uses a multi-layer, fully connected neural network structure and nonlinear activation functions (such as ReLU and Leaky ReLU) to deeply decode these complex features, outputting a set of structured style parameters. These parameters include color preference, texture control, light and dark contrast, and emotional tonal distribution, enabling fine-grained tuning of the image generation model's behavior in terms of style expression.

[0145] In practical applications, the image style parameter mapping network can be flexibly integrated with various image generation frameworks. For example, in generative adversarial networks (GANs), the image style parameter mapping network receives a personality vector for a visual image and outputs multiple "style layer control vectors" that are used to adjust the distribution of feature maps within each layer of the generator, thereby achieving detailed style injection. In diffusion models, the image style parameter mapping network can provide conditional guidance on emotions or preferences during the feature fusion stage of the diffusion process. This structure not only improves the personalization of image generation but also enhances the model's adaptability to different style requirements.

[0146] For example, if a user's visual image personality vector indicates a preference for "cool blues" and a "mild anxiety" interaction state, the image style parameter mapping network will decode a style parameter set with a high blue weight, enhanced texture detail, and slightly reduced overall brightness. The image generation model then uses these parameters to generate an image with cool colors, a gentle rhythm, and an overall atmosphere that matches the user's current psychological state. This approach effectively achieves a direct translation from "psychological state" to "image expression," significantly improving the experience and quality of personalized generation.

[0147] For example, a typical image style parameter mapping network consists of several key components that work together to efficiently map the visual image personality vector to image style control parameters. First, the input layer receives the preprocessed visual image personality vector. This vector typically has a dimension of 64 to 512 and incorporates semantic features such as the user's long-term style preferences and current emotional state. The input can be a one-dimensional vector or a high-dimensional tensor that incorporates positional information or emotional modulation to adapt to different style control requirements and contextual scenarios.

[0148] Next comes the nonlinear transformation layer, the core component of the image style parameter mapping network. This layer typically consists of 4 to 8 fully connected layers, each followed by a nonlinear activation function such as ReLU, Leaky ReLU, or GELU, to enhance the network's ability to fit and represent complex features. To improve network training stability and convergence speed, some architectures also incorporate layer normalization or batch normalization to ensure consistent feature distribution during training and prevent vanishing or exploding gradients.

[0149] Finally, after multiple layers of nonlinear transformations, the information enters the style control decoding layer, which maps high-dimensional features into style control parameters that can be used for image generation. The output can take various forms, including a set of style code vectors (Style Codes) used to drive multi-layer feature modulation in generators like StyleGAN; a modulation parameter tensor used to control the morphology of feature maps in certain convolutional layers of the generative model; or a set of attention weights used to adjust the weight of different style features in the generation process. This output determines the final generated image's color tendency, texture detail, light and dark contrast, and shape composition, ensuring that the image style closely matches the user's psychological characteristics.

[0150] S42: Inputting the image style control parameter set into an image generation and rendering module, wherein the image generation and rendering module processes the pre-acquired 3D reconstruction model of the user's face according to an image processing flow using an image generation model and the image style control parameter set, and outputs a generated personalized static image or dynamic image;

[0151] The image generation model is a deep generation model based on a generative adversarial network or a diffusion model.

[0152] Specifically, the image generation model in this embodiment is composed of multiple functional modules that work together to complete the entire process from input vector to image generation.

[0153] The first is the input module, which is used to receive style control parameters. These parameters are derived from the mapping results of the user's visual image personality vector and determine the style tendency, texture structure and visual features of the final image.

[0154] Next comes the feature mapping module, which typically consists of a multi-layer fully connected network, an embedding layer, or a lightweight encoder. This module converts the input style control parameters into style feature vectors in the latent space. These latent representations are further modulated within the generator to generate various hierarchical features of the generated image, achieving style-driven content shaping.

[0155] The core of image generation lies in the generator. This network structure typically consists of a series of convolutional layers, residual blocks, and upsampling modules, generating image details layer by layer and transforming low-resolution latent features into high-resolution images. For generative models that support style control (such as StyleGAN), a style modulation module (StyleModulation) is incorporated into the generator. This module injects style control parameters into the image generation process through layer-specific normalization operations (such as AdaIN and StyleNorm), dynamically adjusting visual attributes such as color, texture, and structure to generate image content that better suits user preferences.

[0156] Models based on generative adversarial networks (GANs) also feature a discriminator to determine the authenticity of generated images. Through adversarial training, the generator and discriminator continuously iteratively optimize each other, making the generated images more realistic. In generative architectures based on diffusion models, image generation is achieved through multiple rounds of "denoising." Compared to traditional GAN ​​models, diffusion models offer significant advantages in generation stability and detail expression, making them particularly suitable for complex image content and high-resolution output.

[0157] Furthermore, to ensure the quality and style consistency of generated images, the entire generation system relies on a set of carefully designed loss functions for training. Common loss functions include adversarial loss (for improving image realism), perceptual loss (for capturing semantic hierarchy), reconstruction loss (for ensuring content fidelity), and style loss (for maintaining style consistency). Together, they constrain the model output to conform to user preferences and aesthetic expectations.

[0158] The image generation model in this embodiment can be implemented using the StyleGAN2 model (Style-based Generative Adversarial Network v2). StyleGAN2 is a high-performance image generation model based on a generative adversarial network (GAN), offering excellent image synthesis quality and style controllability. Compared to traditional generative networks, StyleGAN2 introduces an image style parameter mapping network and a layer-by-layer modulation mechanism. This allows the image generation process to not only express rich textures and details, but also achieve refined style adjustment through externally input style control parameters.

[0159] In StyleGAN2, the personality vector of a visual image is first converted into a style code in latent space through an image style parameter mapping network. This code is then injected into multiple layers of the generator, where adaptive normalization (such as AdaIN) modulates the distribution of feature maps at each layer, achieving precise control over visual attributes such as image color, shape, structure, and texture. Through this mechanism, StyleGAN2 can synthesize stylistically diverse and distinctively personalized still images, dynamic images, or virtual character appearances based on the visual personality traits of different users.

[0160] The StyleGAN2 model is highly expressive and controllable, generating images with a resolution of up to 1024×1024, rich in detail, and with a coherent style. It is widely used in fields such as avatar modeling, anime style transfer, face generation, and digital art creation. The image generation system built based on StyleGAN2 in this embodiment can significantly improve the matching of generated results with user preferences, meeting diverse personalized image content needs.

[0161] Optionally, in some embodiments of the present application, the image generation and rendering module adopts an image generation model based on a generative adversarial network or a diffusion model, and performs color remapping and brightness adjustment, texture feature refinement, edge structure optimization and anti-aliasing, lighting and shadow mapping simulation, and facial expression and posture-driven rendering in a multi-layer perception channel in sequence for the image style control parameter set, thereby achieving personalized image synthesis.

[0162] Optionally, in some embodiments of the present application, the S14 specifically includes:

[0163] S141, using an image segmentation model to process the collected user facial image, accurately extract the hair area, and generate a high-resolution hair mask image;

[0164] The image segmentation model is a multi-scale semantic segmentation network;

[0165] In this embodiment, in step S141, an image segmentation model is used to process the collected facial image of the user, accurately separate the hair area in the image, and generate a corresponding high-resolution hair mask image. The image segmentation model can use a deep neural network with multi-scale semantic modeling capabilities, such as U-Net, DeepLab or HRNet. These models use encoder-decoder structures, void convolution, jump connection and other technologies to effectively enhance the ability to segment complex textures and edge details in the hair area. Achieving high-quality hairstyle mask extraction lays a clear and accurate image foundation for subsequent edge extraction and geometric analysis, and significantly improves the ability to restore the structure of the user's original hairstyle. S142. Based on the hair mask image, a contour detection algorithm is used to extract the outer edge lines of the hairstyle;

[0166] Based on the generated hair mask image, a contour detection algorithm (such as Canny, Sobel, or edge tracking algorithms) is applied to extract the outer edge lines of the hairstyle area. This contour line represents the overall morphological boundary of the user's current hairstyle, including key information such as hair distribution, hairline contour, and hair tip trend. Through precise edge extraction, this step can construct a geometric representation framework for the hairstyle, making subsequent modeling more structured and comparable. S143: Perform geometric feature analysis on the extracted hairstyle edge curve to extract contour parameters that characterize the hairstyle morphology. The contour parameters include:

[0167] Outline length: the total length of the edge curve of the hairstyle, used to measure the degree of expansion of the hairstyle and reflect the overall coverage of the hair;

[0168] Curvature: The curvature change value per unit length of the edge curve is used to reflect the degree of curvature of the edge curve. It describes the curvature change per unit length of the edge curve and is used to quantify the curvature of the hair edge, such as the distinction between straight hair and curly hair.

[0169] Closure: The degree of connectivity between the beginning and end of the edge curve is used to evaluate the possibility of it forming a closed figure. The degree of connection between the beginning and end of the edge curve is used to evaluate whether the hairstyle forms a closed structure, such as whether the bangs surround the forehead area.

[0170] Converting the complex visual information in the original image into quantifiable and comparable structural features helps to express the hairstyle morphology in the feature space.

[0171] S144 , normalizing or standardizing the contour parameters to eliminate scale differences, and further combining the standardized contour parameters to form a hairstyle contour feature vector for comprehensively describing the shape category and structural features of the user's hairstyle.

[0172] In step S144, the extracted profile parameters are normalized or standardized to eliminate scale differences between the parameters and ensure numerical stability and expression consistency during subsequent model learning. After processing, these normalized parameters are combined in multiple dimensions to generate a hairstyle profile feature vector. This vector, as part of the user's appearance feature vector, is integrated with the image style parameter mapping network and the image generation model to control the expression and consistency of the hairstyle in the generated image.

[0173] This application discloses a personalized content generation method based on user image data processing, which involves the fields of computer vision, image recognition and synthesis, and belongs to the technical category of International Patent Classification G06T (image data processing or generation, especially image analysis or recognition performed by a computer).

[0174] This method collects and processes images of the user's face and the environment in which he is located, and uses image processing algorithms and multimodal deep learning models to model and identify the user's current interaction state, thereby achieving context-awareness-driven content generation optimization. It has multiple technical innovations and significant application effects.

[0175] First, the solution collects user facial images and environmental features, then combines them with multiple models for collaborative analysis, including gesture recognition, expression recognition, hairstyle segmentation, and key point detection, to construct an image interaction state vector. This vector accurately reflects the user's psychological and physiological state in real time, enhancing the personalization and adaptability of generated content. Compared to traditional methods, which often suffer from a cursory understanding of user status and delayed response, this solution achieves precise recognition of user attention, emotion, and engagement through high-dimensional information fusion and deep feature modeling.

[0176] Secondly, the solution innovatively incorporates color stimulus image sequences and real-time facial micro-expression acquisition to automatically establish user color response curves. This overcomes the limitations of traditional subjective questionnaires, which struggle to accurately reflect user preferences, and enables personalized, emotion-driven content generation. Furthermore, the use of deep learning technologies such as CNN ensures recognition accuracy and robustness, demonstrating strong practicality and scalability.

[0177] In summary, the embodiments of the present application significantly enhance the emotional perception ability of human-computer interaction, breaking through the bottlenecks of existing technologies such as the single dimension of user status recognition and insufficient personalization capabilities.

[0178] In addition, an embodiment of the present application further proposes a computer device, which includes a processor and a memory, wherein the processor is used to execute instructions stored in the memory so that the computer device executes the content generation method based on user images described in the above embodiment.

[0179] Finally, an embodiment of the present application also proposes a computer-readable storage medium, including computer program instructions, which, when executed by a processor, implements the content generation method based on user images described in the above embodiment.

[0180] Example 2

[0181] This embodiment discloses a method for generating content based on user images, aiming to create a highly personalized static avatar or virtual character appearance for the user. This method not only captures the user's image but also incorporates their facial expressions, head posture, hairstyle, and emotional response to color stimuli, ultimately generating a virtual avatar that matches the user's style and emotional characteristics.

[0182] The first step is to capture the user's facial image through the front camera, and at the same time use the environmental perception module to record the current ambient light intensity and color temperature to obtain accurate image input conditions. In the image processing stage, the following analysis operations are performed in sequence:

[0183] Use a convolutional neural network (CNN)-based model (such as MediaPipe FaceMesh) to identify 68 facial key points, including the corners of the eyes, mouth, and nose, providing a positioning basis for subsequent action recognition;

[0184] Use posture recognition algorithms such as OpenPose to calculate the user's head angle (pitch, yaw, roll) and shoulder posture;

[0185] Separate the hair area through a multi-scale semantic segmentation network and generate a hair mask map;

[0186] Based on the hair mask image, a contour detection algorithm is used to extract the hair edge, and its length, curvature, closure and other indicators are further calculated and normalized to form a numerical hair contour feature.

[0187] Apply the CNN expression recognition model to identify facial expression action units such as AU1 (inner eyebrow lift), AU6 (zygomatic muscle lift), and AU12 (mouth corner stretch) and their intensity values;

[0188] After the above feature data are aligned according to the timestamp, they are encoded into feature embedding vectors and input into the fusion network to output a unified image interaction state vector.

[0189] In the second step, to further explore the user's emotional response to visual stimuli, a set of images with different tones (such as cold blue, warm orange, light green, dark red, etc.) are shown to the user, with each frame lasting about 1 second, and 10 frames are played continuously. During this process:

[0190] The camera records the user's facial response to each frame of color image in real time, identifying facial micro-expressions such as raised eyebrows and raised corners of the mouth, as well as their intensity;

[0191] Taking each frame of the image as the time reference point, the intensity of the extracted facial action units is plotted as multiple time-intensity response curves;

[0192] Each response curve is fused according to the preset weights (such as 0.5 for raising the corner of the mouth, 0.3 for frowning, and 0.2 for lifting the zygomatic muscle) to form a physiological response curve that represents the user's color emotional preference.

[0193] The third step is to further extract key parameters based on the above color physiological response curve:

[0194] Response start time: If a certain action (such as raising the corner of the mouth) exceeds the response threshold for the first time in the third frame image;

[0195] Maximum response intensity: records the peak intensity of key facial expressions such as AU12;

[0196] High-level maintenance time: Statistics show the time interval during which a certain intensity (such as ≥0.68) lasts, such as 1.2 seconds.

[0197] These parameters describing the user's emotional response pattern are concatenated and fused with the previously generated image interaction state vector to generate a low-dimensional visual image personality vector. This vector summarizes the user's emotional preferences, style characteristics, and appearance tendencies in the current state.

[0198] The fourth step is to input the visual image personality vector into an image style parameter mapping network (e.g., a fully connected neural network with BatchNorm and ReLU activation units). The network outputs multi-dimensional style control parameters, including but not limited to:

[0199] Color tendency: For example, the preference value for cool colors is 0.3, and the preference value for warm colors is 0.7;

[0200] Expression brightness: For example, "high" represents a clear and bright expression style;

[0201] Hairstyle complexity: If the curvature of the hair outline is high, it indicates a preference for curly or fluffy hairstyles.

[0202] These style parameters are fed into an image generation model (such as StyleGAN2 or Stable Diffusion, among other advanced image generators) to synthesize a realistic virtual avatar image that matches the user's style. For example, if a user exhibits a strong positive reaction to warm orange tones, the generated avatar might feature a warm background, naturally curly hair, and a slightly upturned smile, fully reflecting the user's personality.

[0203] This embodiment uses real-life user images, combining facial micro-expressions, physiological responses, and environmental characteristics to comprehensively model the user's appearance and emotional preferences, driving the image generation process. The resulting avatar is highly personalized, faithfully reflecting the user's physical features while also incorporating their underlying style and emotional expressions.

[0204] In another embodiment, after obtaining the visual image personality vector, the visual image personality vector is input into an image style parameter mapping network. This mapping network is a deep neural network model that performs feature extraction and nonlinear mapping on the input personality vector, thereby outputting a set of control parameters representing the image style, referred to as the image style control parameter set.

[0205] The visual image personality vector incorporates key facial features (such as hairstyle, expression, and posture), environmental characteristics, and the user's physiological response to color stimuli. Therefore, it can reflect multidimensional information such as the user's personality, emotions, and style preferences. By modeling this high-dimensional semantic information, the mapping network learns the mapping relationship between different users' image style preferences, thereby generating a set of parameters that can drive image style changes.

[0206] The training data of the image style parameter mapping network may include a large number of user image samples and their corresponding style labels. A supervised or self-supervised learning mechanism is adopted, and the loss function constraint is used to enable the network to stably output controllable and personalized style parameters.

[0207] The image style control parameter set is then input into the image generation and rendering module. This module, based on a deep generative model, uses the parameter set as style guidance information to perform image synthesis and stylization on a pre-acquired 3D reconstruction model of the user's face.

[0208] Specifically, the image generation and rendering module includes an image generation model, which can be one of the following types:

[0209] Generative Adversarial Network (GAN): Through adversarial training, the generator can learn to generate realistic and personalized images under the supervision of the discriminator;

[0210] Diffusion Model: This model gradually models the image denoising and de-noising process to achieve high-quality, high-fidelity image synthesis.

[0211] During the image generation and rendering process, the following image processing steps are performed in sequence for the user's 3D reconstructed model:

[0212] Color remapping and brightness adjustment: adjust the overall color tone and lighting atmosphere according to style parameters;

[0213] Texture detection and refinement generation: Generate detailed textures for areas such as faces, hairstyles, and clothing;

[0214] Edge enhancement and anti-aliasing processing: improve image clarity and naturalness;

[0215] Lighting simulation and shadow rendering: enhance the three-dimensional sense and realism of space;

[0216] Expression and gesture animation processing: Dynamically restore the user's emotional expression or head movement in dynamic images.

[0217] The final output is a static or dynamic image with highly personalized features. The image not only retains the user's own structural information, but also reflects his or her unique preferences in visual style and emotion.

[0218] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0219] In this application, unless otherwise specified or limited, the terms "mounted," "connected," "connect," "fixed," etc. should be understood broadly. For example, they can refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.

[0220] In this application, unless otherwise expressly specified or limited, when a first feature is “on” or “below” a second feature, it may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Moreover, when a first feature is “above”, “above”, or “above” a second feature, it may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is “below”, “below”, or “below” a second feature, it may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is at a lower level than the second feature.

[0221] In the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiments", "examples", "specific examples" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.

[0222] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A content generation method based on user images, characterized in that: The following steps are involved: S1. Capturing a user's facial image through an image acquisition device and collecting environmental features of the user through an environmental perception module, extracting the user's facial key points, hairstyle features, expression intensity parameters, and posture features from the facial image, and generating an image interaction state vector based on the environmental features; S2. Presenting a sequence of color stimulus images of different tones to the user, and synchronously collecting the user's facial image data during the presentation process, extracting the user's facial micro-expressions from the facial image data, establishing a facial expression response curve, and further generating a color physiological response curve; The S2 specifically includes: S21, continuously presenting a sequence of color stimulation images containing different tones to the user, and during the presentation of the color stimulation images, using an image acquisition device to collect the user's facial image data in real time; S22, inputting the collected facial image data into an expression recognition model to extract the user's facial micro-expression features, wherein the facial micro-expression features include multiple recognized expression action units and their corresponding intensity values; S23, establishing a set of facial expression response curves of the user based on the intensity values ​​of the expression action units at corresponding time points and each frame image in the color stimulation image sequence; S24, constructing a color physiological response curve reflecting the user under color stimulation conditions of different tones based on the facial expression response curve corresponding to each expression action unit; wherein the color physiological response curve is obtained by weighted fusion of the response curves of all the expression action units according to preset weights; S3. extracting specified feature parameters based on the color physiological response curve, and fusing them with the image interaction state vector to generate a visual image personality vector for controlling the image style; S4. Inputting the visual image personality vector into an image style parameter mapping network to obtain an image style control parameter set. The image generation and rendering module processes the pre-acquired 3D reconstruction model of the user's face according to the image style control parameter set according to the image processing process, and outputs the generated personalized static image or dynamic image. The image processing process includes: color remapping and brightness adjustment, texture detection and refinement generation, edge enhancement and anti-aliasing processing, lighting simulation and shadow rendering, expression and posture animation processing.

2. The method for generating content based on user images according to claim 1, characterized in that: Said S1 specifically includes: S11, using an image acquisition device to capture a user's facial image; S12, analyzing the collected user facial image through a posture recognition algorithm to extract the user's posture features; The posture features include: head posture angle and body posture information; S13, using a facial key point detection model based on a convolutional neural network to identify the locations of facial key points in the user's facial image; The facial key points include: facial contour, eyes, nose, and mouth; S14, using an image segmentation model to identify the hair area in the user's facial image and obtain a hairstyle contour feature vector; S15, using an expression recognition model to identify expression action units contained in the user's facial image, and recording the intensity value of each expression action unit as an expression intensity parameter; S16, aligning the user's posture features, facial key points, hairstyle contour feature vector, expression action unit intensity parameters, and environmental features according to a unified timestamp, inputting them into a feature embedding encoder for standardization, and generating a low-dimensional joint feature vector through a feature fusion network as an image interaction state vector representing the current user interaction state; The environmental characteristics include: light intensity and color temperature.

3. The method for generating content based on user images according to claim 2, characterized in that: The S14 specifically includes: S141, using an image segmentation model to process the collected user facial image, accurately extract the hair area, and generate a high-resolution hair mask image; The image segmentation model is a multi-scale semantic segmentation network; S142, extracting the outer edge line of the hairstyle using a contour detection algorithm based on the hair mask image; S143. Performing geometric feature analysis on the extracted hairstyle edge curve to extract contour parameters representing the hairstyle morphology. The contour parameters include: Contour length: the total length of the edge curve of the hairstyle; Curvature: the curvature change value of the edge curve within a unit length, which is used to reflect the degree of tortuosity of the edge curve; Closure: The degree of connectivity between the beginning and end of the edge curve, used to evaluate the possibility of forming a closed figure; S144 , normalizing or standardizing the contour parameters to eliminate scale differences, and further combining the standardized contour parameters to form a hairstyle contour feature vector for comprehensively describing the shape category and structural features of the user's hairstyle.

4. The method for generating content based on user images according to claim 3, characterized in that: The facial expression action units include: inner eyebrow raising action unit, outer eyebrow raising action unit, frowning action unit, upper eyelid lifting action unit, zygomatic muscle lifting action unit, eyelid closing action unit, nose wrinkling action unit, upper lip lifting action unit, mouth corner raising action unit, and mouth corner lowering action unit; The expression recognition model is a convolutional neural network model.

5. The method for generating content based on user images according to claim 4, characterized in that: In the process of establishing the user's facial expression response curve, the intensity value of each expression action unit is used as the vertical axis, and the time sequence of the color stimulus image is used as the horizontal axis. A facial expression response curve is established for each expression action unit separately.

6. The method for generating content based on user images according to claim 5, characterized in that: The S3 specifically includes: S31, extracting specified characteristic parameters from the color physiological response curve; The specified characteristic parameters include: response start time, maximum response intensity, and high level maintenance time; The response start time is: with the time axis on the curve as the horizontal axis, the time point corresponding to when the intensity value exceeds the preset threshold for the first time; Maximum response intensity is: the maximum intensity value in the color physiological response curve; High maintenance time refers to the length of time during which the intensity value remains above the high intensity threshold; The high intensity threshold is 80% of the maximum response intensity; S32. Fusing the specified feature parameters with the image interaction state vector to form a combined feature vector, and using the combined feature vector as the visual image personality vector.

7. The method for generating content based on user images according to claim 6, characterized in that: The S4 specifically includes: S41, inputting the visual image personality vector into an image style parameter mapping network to map it into an image style control parameter set; S42: Inputting the image style control parameter set into an image generation and rendering module, wherein the image generation and rendering module processes the pre-acquired 3D reconstruction model of the user's face according to an image processing flow using an image generation model and the image style control parameter set, and outputs a generated personalized static image or dynamic image; Among them, the image style parameter mapping network is a deep neural network, which is used to map the visual image personality vector into the style control parameters required for image generation, and the image generation model is a deep generation model based on a generative adversarial network or a diffusion model.

8. An electronic device, characterized in that: comprising a memory and a processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program to implement the content generation method based on user images according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is used to implement the method for generating content based on a user image according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Virtual character image construction device and method

    CN112164135A

  • Image cartoonalization method and device

    CN113470147A

  • Three-dimensional digital human generation and interaction method and system

    CN117496072A

  • Emotional engagement detection method based on positive emotional perception

    US20250014320A1