Content generation method and device based on user image and medium

By collecting user facial images and environmental features, combining color stimulation image sequences, establishing expressions and color response curves, and generating visual image personality vectors, the problems of poor emotional perception accuracy and low degree of personalization in the prior art are solved, and personalized image generation is achieved.

CN120374785AActive Publication Date: 2025-07-25BEIJING CENT BIOLOGY CO LTD

Patent Information

Application Number
CN202510875894.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Existing content generation methods based on user images are difficult to accurately capture the user's real-time state and potential preferences of users in terms of poor emotional perception accuracy, weak visual preference modeling ability, and low personalization of content generated.

Method used

By collecting user facial images and environmental features, facial key points, hairstyle features, expression intensity parameters and posture features are extracted, facial micro-expressions are collected in combination with color stimulation image sequences, expression response curves and color physiological response curves are established, visual image personality vectors are generated, and personalized images are generated through image style parameter mapping networks.

Benefits of technology

It realizes high accuracy perception of the user's subjective state, and the generated images are more personalized, adaptable and emotionally consistent, enhancing the degree of personalization of the generated content and the matching of the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374785A_ABST
    Figure CN120374785A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image generation, in particular to a content generation method and device based on a user image and a medium, and the method comprises the steps: collecting a user face image and environment features, and extracting face key points, hair style features, expression intensity parameters and posture features of a user in the face image; generating an image interaction state vector in combination with the environment features; presenting a stimulus image sequence containing different tones and colors to the user, collecting facial micro-expressions of the user, establishing a facial expression response curve, and further generating a color physiological response curve; extracting specified characteristic parameters based on the color physiological response curve, and constructing a visual image personality vector in combination with the image interaction state vector; and inputting the personality vector of the visual image into an image style parameter mapping network to obtain an image style control parameter set, processing the three-dimensional reconstruction model of the user face by an image generation and rendering module according to the image style control parameter set, and outputting a static image or a dynamic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image generation, and particularly to a method, device and medium for generating content based on user images. Background Art

[0002] With the rapid development of artificial intelligence and image processing technologies, personalized content generation has gradually become a research hotspot in the fields of image synthesis and human-computer interaction. In the technical system related to image generation and visual expression, user personality feature modeling, image style control, and personalized synthesis models have become key technical paths. Especially in the field of computer vision image generation (G06T), the technology of extracting semantic information from user image data and driving image generation is widely used in scenarios such as virtual image construction, digital human generation, and immersive interaction.

[0003] In the prior art, the generation of personalized image content mostly relies on static user portraits or style templates selected by users, and it is difficult to accurately capture the user's current real-time state and potential preferences. For example, some methods achieve image style generation through face recognition and style matching recommendations. Although they can provide a certain degree of visual style transformation, it is difficult to extract deep psychological or physiological features from multi-dimensional features such as the user's facial micro-expressions, posture changes, and emotional reactions, resulting in insufficient personalization and matching of the generated content. At the same time, in existing image style conversion or generation methods, style parameters are mostly constructed based on static images or preset labels, lacking modeling of the user's reactions under different environmental perception conditions (such as lighting and color temperature), and it is difficult to construct a dynamic content generation mechanism closely related to the user's current state. In recent years, the development of emotion computing and micro-expression recognition has provided new technical support for obtaining the user's real emotions and psychological preferences, but such information has not been fully utilized in existing image generation systems.

[0004] For example, some solutions, such as Patent CN112164135A (main classification number G06T), propose a device and method for constructing a virtual character image, but it has obvious defects in terms of emotion perception, visual preference modeling, and personalized content generation. This solution does not set up an emotion recognition mechanism and cannot accurately obtain the user's emotional state; it only relies on subjective descriptions to select facial features and lacks the ability to model the user's visual preferences; the generated content is fixed and cannot dynamically adjust the generation result according to the user's image and emotion, with low personalization and difficulty in meeting the user's differentiated needs.

[0005] Another type of solution, Patent CN116433800B (main classification number G06T) proposes an image generation method based on the joint guidance of user preferences and text in social scenarios. This method combines user preferences and text in social scenarios for image generation, but it has deficiencies in terms of emotional perception and the real-time performance and accuracy of individualized content generation. Its user preference modeling relies on social relationships and image interaction history, making it difficult to accurately capture the user's current emotional state and lacking the ability to perceive real-time user image and facial emotion features. In addition, the generated content is limited by social network data and pre-trained models, making it difficult to achieve high-precision and deep-level personalized expression based on specific user images. Summary of the Invention

[0006] (I) Technical Problems to be Solved

[0007] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present application provides a content generation method, device, and medium based on user images, which solve the technical problems of poor emotional perception accuracy, weak visual preference modeling ability, and low personalization degree of the existing content generation method based on user images.

[0008] (II) Technical Solutions

[0009] To achieve the above object, the main technical solutions adopted in the present application include:

[0010] In a first aspect, an embodiment of the present application provides a content generation method based on user images, including the following steps:

[0011] S1. Collect a user's facial image through an image acquisition device, and collect the environmental characteristics of the user's location through an environmental perception module. Extract the facial key points, hairstyle features, expression intensity parameters, and posture features of the user in the facial image, and generate an image interaction state vector in combination with the environmental characteristics;

[0012] S2. Present a sequence of color-stimulating images with different hues to the user, and synchronously collect the user's facial image data during the presentation. Extract the user's facial micro-expressions in the facial image data, establish a facial expression response curve, and further generate a color physiological response curve;

[0013] S3. Extract specified feature parameters based on the color physiological response curve, and fuse them with the image interaction state vector to generate a visual image personality vector for controlling the image style;

[0014] S4. Input the visual image personality vector into an image style parameter mapping network to map and obtain an image style control parameter set. The image generation and rendering module processes the three-dimensional reconstruction model of the user's face obtained in advance according to the image processing process based on the image style control parameter set, and outputs the generated personalized static image or dynamic image.

[0015] The described image processing process sequentially includes: color remapping and brightness adjustment, texture detection and refinement generation, edge enhancement and anti-aliasing processing, lighting simulation and shadow rendering, and expression and pose animation processing.

[0016] Preferably, the S1 specifically includes:

[0017] S11. Use an image acquisition device to collect the user's facial image;

[0018] S12. Analyze the collected user's facial image through a pose recognition algorithm to extract the user's pose features;

[0019] The pose features include: head pose angle and body pose information;

[0020] S13. Use a facial key point detection model based on a convolutional neural network to identify the positions of facial key points in the user's facial image;

[0021] The facial key points include: facial contour, eyes, nose, mouth;

[0022] S14. Use an image segmentation model to identify the hair region in the user's facial image and obtain a hairstyle contour feature vector;

[0023] S15. Use an expression recognition model to identify the expression action units included in the user's facial image, and record the intensity values of each expression action unit as expression intensity parameters;

[0024] S16. Synchronize and align the user's pose features, facial key points, hairstyle contour feature vector, expression action unit intensity parameters, and environmental features at the same timestamp, respectively input them into a feature embedding encoder for standardization processing, and generate a low-dimensional joint feature vector through a feature fusion network as an image interaction state vector representing the current user interaction state.

[0025] The environmental features include: light intensity, color temperature.

[0026] Preferably, the S14 specifically includes:

[0027] S141. Use an image segmentation model to process the collected user's facial image, accurately extract the hair region, and generate a high-resolution hair mask image;

[0028] The image segmentation model is a multi-scale semantic segmentation network; S142. Based on the hair mask image, use a contour detection algorithm to extract the outer edge line of the hairstyle; S143. Conduct geometric feature analysis on the extracted hairstyle edge curve, and extract contour parameters characterizing the hairstyle morphology. The contour parameters include:

[0029] Contour length: the total length of the hairstyle edge curve;

[0030] Curvature: the curvature change value within the unit length of the edge curve, used to reflect the tortuous degree of the edge curve;

[0031] Closure: the connectivity degree between the head and the tail of the edge curve, used to evaluate the possibility of forming a closed figure;

[0032] S144. Normalize or standardize the contour parameters to eliminate the scale difference, and further combine the standardized contour parameters to form a hairstyle contour feature vector for comprehensively describing the hairstyle shape category and structural characteristics of the user.

[0033] Preferably, the expression action units include: inner eyebrow raising action unit, outer eyebrow raising action unit, frowning action unit, upper eyelid lifting action unit, zygomaticus major muscle lifting action unit, eyelid closing action unit, nose wrinkling action unit, upper lip lifting action unit, mouth corner raising action unit, mouth corner lowering action unit;

[0034] The expression recognition model is a convolutional neural network model.

[0035] Preferably, the S2 specifically includes:

[0036] S21. Continuously present a sequence of color stimulus images with different hues to the user, and during the presentation of the color stimulus images, use an image acquisition device to collect the facial image data of the user in real time;

[0037] S22. Input the collected facial image data into the expression recognition model to extract the facial micro-expression features of the user, where the facial micro-expression features include multiple recognized expression action units and their corresponding intensity values;

[0038] S23. Establish a set of facial expression response curves of the user according to the intensity values of the expression action units in each frame image of the color stimulus image sequence and the corresponding time points;

[0039] During the establishment of the facial expression response curves of the user, use the intensity value of each expression action unit as the vertical axis and the time sequence of the color stimulus images as the horizontal axis, and separately establish a facial expression response curve for each expression action unit;

[0040] S24. Based on the facial expression response curves corresponding to each expression action unit, construct a color physiological response curve reflecting the user's color physiological response under different hue color stimulus conditions.

[0041] Preferably, the color physiological response curve is obtained by weighted fusion of the response curves of all expression action units according to preset weights.

[0042] Preferably, S3 specifically includes:

[0043] S31. Extract specified feature parameters from the color physiological response curve;

[0044] The specified feature parameters include: response start time, maximum response intensity, and high-level maintenance time;

[0045] Among them, the response start time is: with the time axis on the curve as the horizontal axis, the time point corresponding to when the intensity value first exceeds a preset threshold;

[0046] The maximum response intensity is: the maximum intensity value in the color physiological response curve;

[0047] The high-level maintenance time is the length of the time period during which the intensity value continuously remains above the high-intensity threshold;

[0048] The high-intensity threshold is 80% of the maximum response intensity;

[0049] S32. Fuse the specified feature parameters with the image interaction state vector to form a combined feature vector, and use the combined feature vector as the visual image personality vector.

[0050] Preferably, S4 specifically includes:

[0051] S41. Input the visual image personality vector into the image style parameter mapping network to map it into an image style control parameter set;

[0052] S42. Input the image style control parameter set into the image generation and rendering module. The image generation and rendering module uses an image generation model through the image style control parameter set to process the three-dimensional reconstruction model of the user's face obtained in advance according to the image processing process, and outputs the generated personalized static image or dynamic image;

[0053] Among them, the image style parameter mapping network is a deep neural network for mapping the visual image personality vector into the style control parameters required for image generation, and the image generation model is a deep generation model based on a generative adversarial network or a diffusion model.

[0054] In a second aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, wherein:

[0055] The memory is used to store a computer program;

[0056] The processor is used to execute the computer program to implement the method for generating content based on a user image as described above.

[0057] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to implement the method for generating content based on a user image as described above.

[0058] (III) Advantageous Effects

[0059] The method for generating content based on a user image provided by the embodiment of the present application can comprehensively and synchronously collect and fuse user postures, facial key points, expression intensities, hairstyle contours, and environmental factors (such as light, color temperature, etc.) by collecting user facial images and environmental features, and then construct a unified expression vector of the current user state, effectively improving the perception accuracy and stability of the user's subjective state.

[0060] Furthermore, the embodiment of the present application captures the physiological response characteristics of the user to color stimuli by collecting color stimuli in conjunction with facial micro-expressions and combining the facial micro-expressions of the user in images of different color tones, establishing a fine facial expression response curve and a color physiological response curve, thus breaking through the limitations of traditional static expression recognition. By constructing a visual image personality vector that fuses the interaction state and color response, the embodiment of the present application can accurately reflect the user's subjective preferences, perception characteristics, and visual emotion characteristics, and map them as input control parameters to the deep style generation network, making the generated images more personalized, adaptable, and emotionally consistent.

[0061] In addition, the method of the present application extracts the hairstyle contour through a multi-scale image segmentation model, and generates a structured description vector after standardizing in combination with geometric features (such as curvature, closure, etc.), significantly improving the system's expression ability for hairstyle patterns and individual characteristics. In the image generation stage, the visual image personality vector is converted into control parameters of the image generation model through an image style parameter mapping network, which can flexibly drive the image generation model based on, and realize the automatic generation of highly customized static images and dynamic images. Even in the case of complex user states, subtle expressions, or changing environments, it can still maintain the consistency between the generated content and the user's actual emotions and preferences, ensuring the personalization, emotionalization, and high matching degree of the overall interaction experience.

[0062] In summary, the embodiment of the present application establishes a standardized and richly expressive visual personality modeling and generation path by fusing multi-dimensional data such as user images, environmental perception, and color response, not only enhancing the subjective fit and aesthetic relevance of the generated images. Description of the Drawings

[0063] Figure 1Schematic flowchart of a content generation method based on user images according to an embodiment of the present application;

[0064] Figure 2 Flowchart of the process for generating an image interaction state vector according to an embodiment of the present application;

[0065] Figure 3 Flowchart of the process for generating a color physiological response curve according to an embodiment of the present application;

[0066] Figure 4 Flowchart of the process for constructing a visual image personality vector according to an embodiment of the present application;

[0067] Figure 5 Flowchart of the process for generating corresponding static or dynamic images according to an embodiment of the present application. Detailed implementation manners

[0068] In order to better explain the present application for easy understanding, the present application will be described in detail below with reference to the accompanying drawings through specific implementation manners.

[0069] In the current field of image generation and personalized content presentation, especially in the task of content generation based on user images, the following several typical problems mainly exist in the prior art:

[0070] Traditional image style generation methods usually rely on fixed input image features (such as style images and content images) for image synthesis. Some technologies use user images for avatar generation or cartoonization (such as CN113470147A, the main classification number is G06T), but lack in-depth modeling of subjective level features of users' individuals in aspects such as emotions, physiology, and aesthetic preferences. For example, physiological and psychological feedback information such as micro-expression changes and color preferences generated when users face images with different hues, textures, and scenarios are not effectively perceived or utilized.

[0071] In current methods, means such as facial key point positioning and expression recognition are often used to obtain static information of users' facial images, but information dimensions with more discriminative power for style such as hairstyle contours and posture features are often ignored, resulting in the generated content lacking delicate individuality in visual performance and being difficult to accurately express the overall visual image of users.

[0072] There is no dynamic vector modeling mechanism for the "interaction state" composed of environmental characteristics (such as light, color temperature) and the user's current posture and expression combination, resulting in the disconnection between the generated image and the user's current real scene, lacking situational consistency and interaction immersion. In image generation applications involving user emotion participation, it often relies on preset tags (such as "happy", "sad") or text descriptions to model the emotion state, fails to construct a dynamic physiological response model based on the "user's facial micro-expression and color-stimulated image", cannot truly depict the physiological response path of the user to different hues, and cannot form a personalized "visual personality". Since the user's facial key points, expression intensity, hairstyle geometry, posture, environmental light, etc. are features in different dimensions and semantic spaces, the existing methods lack effective coding structures and control strategies in unified embedding and style parameter mapping, resulting in an unstable style mapping process, insufficient individual differences in the generated results, and poor visual continuity.

[0073] In the existing image generation and recommendation fields, traditional methods mainly rely on user selection or preset parameters for content presentation, ignoring the implicit preferences reflected by the user's physiological and emotional responses during the actual interaction process, resulting in limited personalization of the generated content and low user acceptance.

[0074] To solve the above problems, this application proposes a content generation method based on user images, which deeply integrates image processing, facial feature modeling, color physiological response modeling, and image style control to construct a user modeling and image content generation system with the "visual image personality vector" as the core. This method belongs to the field of image data processing (G06T) and specifically includes the following key technical steps:

[0075] Use an image acquisition device to obtain the user's facial image data, combine with the environmental perception module to collect environmental factors such as light intensity and color temperature, extract visual features such as facial key points (such as eyes, nose, mouth, etc.), expression intensity, posture angle, and hairstyle contour through a convolutional neural network, and generate a low-dimensional vector (i.e., image interaction state vector) that can reflect the user's current interaction state through unified timestamp alignment and embedding network processing. This stage also solves the synchronization problem of multi-source heterogeneous image features (expression, hairstyle, posture) in the time domain and realizes effective modeling of the user's current visual state.

[0076] Present a sequence of dynamic color-stimulated images to the user, synchronously collect the changes in the user's facial micro-expressions during the process, identify multiple action units (AUs) of expressions including inner eyebrow raising, mouth corner raising, etc., and establish the intensity change curve in the time dimension, that is, the facial expression response curve. Further, generate a color physiological response curve that reflects the user's physiological feedback to color stimuli by weighted fusion of the responses of different action units (AUs) of expressions. This step constructs the technical chain of "user-color preference" modeling in the field of image data processing.

[0077] Based on the feature processing of image data, the system extracts specified feature parameters (such as response start time, maximum intensity, and high-level maintenance time) from the color physiological response curve, and fuses them with the aforementioned image interaction state vector to form a visual image personality vector that integrates emotional responses. This vector, as a compact expression of the user's content preference in the current scene, helps improve the accuracy of style matching in image generation tasks.

[0078] The generated visual image personality vector is used as input and converted into style control parameters through a nonlinear mapping network to drive image generation models (such as StyleGAN, Diffusion, etc.) to generate static or dynamic images that meet user emotional preferences and personalized needs. The focus of this stage is to implement the abstract user perception modeling results into specific visual content, opening up an end-to-end closed-loop path from "image recognition" to "image generation."

[0079] In order to better understand the above technical solution, exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a clearer and more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0080] Embodiment 1

[0081] Figure 1 FIG. 1 is a flow chart of a method for generating content based on user images according to an embodiment of the present application. Figure 1 As shown, the content generation method based on user images includes:

[0082] S1, collecting a user's facial image through an image acquisition device, and collecting the user's environmental features through an environmental perception module, extracting the user's facial key points, hairstyle features, expression intensity parameters, and posture features in the facial image, and generating an image interaction state vector in combination with the environmental features;

[0083] Specifically, this embodiment can make subsequent content more suitable for the current psychological and physiological state of the user through the image interaction state vector.

[0084] See also Figure 2 In the practical application of this embodiment, S1 specifically includes:

[0085] S11, using an image acquisition device to acquire a user's facial image;

[0086] S12, analyzing the collected user facial image through a posture recognition algorithm to extract the user's posture features;

[0087] The pose features include: head pose angle and body posture information;

[0088] Specifically, in step S12, mature pose recognition algorithms are used to extract pose features such as the pitch angle, yaw angle, and roll angle of the user's head, which can effectively judge the user's attention direction, mental state, and willingness to participate. For example, a large degree of facial deviation may indicate that the user's attention is scattered, and lowering the head or leaning back may indicate their state of interaction fatigue. Through these data, it helps to enhance the dynamic perception ability of the user's state.

[0089] S13. Use a facial key point detection model based on a convolutional neural network to identify the positions of facial key points in the user's facial image;

[0090] The facial key points include: facial contour, eyes, nose, and mouth;

[0091] In this embodiment, a deep learning model based on the CNN (Convolutional Neural Network) architecture is used to automatically extract two-dimensional or three-dimensional position information of key regions such as the facial contour, eyes, nose, and mouth, improving the accuracy and robustness of the detection.

[0092] The facial key point detection model based on a convolutional neural network (CNN, Convolutional Neural Network) is one of the most widely used and most accurate face structure recognition technologies in the current field of computer vision. The core task of this type of model is to automatically identify and locate multiple semantically meaningful facial key points from the input face image, such as the corners of the eyes, the ends of the eyebrows, the tip of the nose, the corners of the mouth, etc., and is commonly used in application scenarios such as constructing facial geometric structures, expression analysis, virtual filters, and emotion recognition.

[0093] S14. Use an image segmentation model to identify the hair region in the user's facial image and obtain a hairstyle contour feature vector;

[0094] The user's hairstyle is an important part of their visual style. In this step S14, an image segmentation model (such as U-Net or DeepLab) is used to accurately segment the hair region in the image and extract its contour feature vector. This not only enhances the comprehensive modeling ability of the user's appearance but also provides data support for maintaining the consistency of the hairstyle in the generated results, thus ensuring the faithful restoration of the generated image to the user's original image.

[0095] S15. Use an expression recognition model to identify the expression action units contained in the user's facial image and record the intensity values of each expression action unit as expression intensity parameters;

[0096] Specifically, step S15 involves identifying the action units (AUs) in the Facial Action Coding System (FACS) and quantifying their intensities, which helps accurately depict the subtle changes in the user's current emotional state. Compared with traditional coarse classification methods (such as only distinguishing between happiness, anger, sorrow, and joy), this method has higher emotional resolution, can effectively capture complex expression responses, and is particularly suitable for micro-expression recognition scenarios, greatly enhancing the sensitivity and response ability to emotional situations.

[0097] It should be noted that the expression recognition model is a type of artificial intelligence model used to automatically analyze the facial expression states in face images and is widely applied in fields such as emotion computing, human-computer interaction, virtual reality, and psychological analysis. In this technical solution, the role of the expression recognition model is to identify the action units (Action Units, AUs) contained in the user's facial image and further extract the intensity values of each action unit, thereby constructing a quantitative expression of the user's current facial emotional state - namely, the expression intensity parameter. The core basis of this expression recognition model is the "Facial Action Coding System" (FACS), which decomposes human facial expressions into multiple independent muscle action units, such as raising the inner eyebrows, stretching the corners of the mouth, and narrowing the eyes. Each action unit reflects the movement state of specific facial muscle groups, and different combinations represent different emotions.

[0098] The action units include: inner eyebrow raising action unit, outer eyebrow raising action unit, frowning action unit, upper eyelid lifting action unit, zygomaticus major muscle lifting action unit, eyelid closure action unit, nose wrinkling action unit, upper lip lifting action unit, corner of the mouth raising action unit, corner of the mouth lowering action unit; the expression recognition model is a convolutional neural network model.

[0099] For example, the expression recognition model generally uses a deep convolutional neural network (CNN) as the basic framework and automatically identifies each action unit in the image by training and learning a large amount of data with action unit annotations. For example, the model will output the activation degree values corresponding to action unit AU01 (inner eyebrow raising), action unit AU06 (zygomaticus major muscle contraction), action unit AU12 (corner of the mouth raising), etc. as the "expression intensity parameters".

[0100] A typical expression recognition process includes the following steps:

[0101] Face detection and alignment: First, detect the face in the image and standardize its orientation and size through affine transformation and other methods to ensure the accuracy of subsequent recognition results.

[0102] Process the face images using pre-trained CNN models (such as ResNet, MobileNet, EfficientNet) to extract deep semantic features and capture the subtle differences in facial muscle movements.

[0103] Identify and evaluate the intensity of each action unit through a multi-label classification or regression network. Some systems will adopt an attention mechanism or a graph convolutional neural network (GCN) to simulate the correlation between muscles, thereby improving the recognition accuracy.

[0104] Finally, output the numbers of multiple action units and their corresponding activation intensities, forming a set of high-dimensional numerical features as an objective expression of the user's current facial expression state.

[0105] S16. Synchronize and align the user's pose features, facial key points, hairstyle contour feature vectors, expression action unit intensity parameters, and environmental features at the same timestamp, input them into the feature embedding encoder for standardization processing respectively, and generate a low-dimensional joint feature vector through the feature fusion network as the image interaction state vector representing the current user interaction state;

[0106] The environmental features include: light intensity, color temperature.

[0107] Specifically, in step S16, a time synchronization mechanism is used to align multi-source heterogeneous features to eliminate the interference of data time delay on the analysis results; in addition, the feature embedding encoder is used to complete the standardization and vectorization mapping of data with different dimensions and types, improving the effectiveness of data fusion; finally, a feature fusion network (such as Transformer or multi-layer perceptron) is used to generate an image interaction state vector with strong expressiveness and compactness as an abstract representation of the user's current mental and behavioral state. This vector has the advantages of strong real-time performance, high information density, and good generalization ability.

[0108] S2. Present a sequence of color stimulus images with different hues to the user, and synchronously collect the user's face image data during the presentation, extract the user's facial micro-expressions in the face image data, establish a facial expression response curve, and further generate a color physiological response curve;

[0109] In this embodiment, a sequence of color stimulus images with different tones is presented to the user, and the user's facial micro-expressions during the viewing process are simultaneously collected, which helps to quantify the user's physiological response to different colors in a non-invasive manner. Facial micro-expressions are highly unconscious and instantaneous, and can more realistically reflect the user's potential emotional preferences for colors compared to traditional questionnaires or subjective scoring methods. By analyzing the dynamic trend of facial micro-expressions changing with color, establishing a user's facial expression response curve, and further extracting a color physiological response curve representing the user's visual emotional changes, the ability to effectively enhance the insight into the user's potential psychological state can be effectively enhanced.

[0110] Preferably, in some embodiments of the present application, see Figure 3 , the S2 specifically includes:

[0111] S21, continuously presenting a sequence of color stimulation images containing different tones to the user, and during the presentation of the color stimulation images, using an image acquisition device to collect facial image data of the user in real time;

[0112] By playing a sequence of images with various tones in a controlled environment, the user's natural response to different colors is stimulated. At the same time, a high-frame rate image acquisition device is used to record the user's facial image data in real time to ensure that brief facial muscle movements are captured. Unlike traditional psychological surveys or subjective feedback methods, this method collects information through real physiological reactions, which is not easily interfered or concealed by the user's consciousness, thus having higher accuracy and ecological validity.

[0113] S22, inputting the collected facial image data into an expression recognition model to extract the user's facial micro-expression features, wherein the facial micro-expression features include a plurality of recognized expression action units and their corresponding intensity values;

[0114] A high-precision expression recognition model is used to analyze each frame of facial images, and multiple facial expression action units (AUs) and their intensity values are output. Since facial micro-expressions are usually completed between 0.1 and 0.5 seconds, and are short-lived and difficult to control, the use of deep learning models for fast and fine-grained analysis can effectively capture the user's true emotional changes. For example, when a cold-colored image is presented, the model may identify uncomfortable movements such as eyebrow depression and orbicularis oculi contraction, thereby reflecting the user's potential negative emotions.

[0115] S23, establishing a group of user facial expression response curves according to the intensity values of the expression action units at the corresponding time points in each frame image in the color stimulus image sequence;

[0116] In the process of establishing the user's facial expression response curve, for each facial expression action unit, the intensity value is used as the vertical axis, and the chronological order of the color stimulus images is used as the horizontal axis. A separate facial expression response curve is established for each facial expression action unit.

[0117] By matching the hue information of each frame of the image with the intensity value of the facial expression action unit at that time point, a response curve of intensity changing with time is constructed for each facial expression action unit. The curve takes time as the horizontal axis and the intensity of the facial expression action unit as the vertical axis, reflecting the influence trend of different hue images on facial muscle activities. The establishment of this response curve not only provides a visualization tool for the user's physiological response but also makes subsequent data analysis more continuous and interpretable.

[0118] S24. Based on the facial expression response curve corresponding to each facial expression action unit, construct a color physiological response curve reflecting the user's color physiological response under different hue color stimulus conditions.

[0119] Among them, the color physiological response curve is obtained by weighted fusion of the response curves of all facial expression action units according to a preset weight.

[0120] By weighted fusion of multiple facial expression action unit response curves, a color physiological response curve with strong comprehensiveness and high representativeness is formed. The preset weight can be set according to the contribution degree of each facial expression action unit in expressing specific emotions (such as joy, surprise, disgust). For example, the upward curvature of the mouth corners (AU12) is highly correlated with the sense of pleasure, so its weight can be relatively high. The finally obtained color physiological response curve not only reveals the user's emotional response pattern to different hues.

[0121] S3. Extract specified feature parameters based on the color physiological response curve, and combine with the image interaction state vector to construct a visual image personality vector integrating emotional responses;

[0122] Extract specified feature parameters based on the above color physiological response curve, and combine with the image interaction state vector to construct a visual image personality vector integrating emotional response features and environmental state features. This visual image personality vector not only contains the user's long-term preferences (such as color sensitivity, style preference, etc.) but also contains the user's instantaneous emotional state in the current environment, realizing a user modeling method of "personality + context" two-way integration. The visual image personality vector has good expression ability and generalization, and can be used not only for immediate image style adjustment but also for accumulating to form the user's long-term style profile, improving the personalized service ability.

[0123] Preferably, in some embodiments of the present application, refer to Figure 4 , where S3 specifically includes:

[0124] S31. Extract the specified feature parameters from the color physiological response curve;

[0125] The specified feature parameters include: response start time, maximum response intensity, and high-level maintenance time;

[0126] Among them, the response start time is: with the time axis on the curve as the horizontal axis, the time point corresponding to when the intensity value first exceeds a preset threshold;

[0127] In this embodiment, the response start time reflects the user's first-time response ability to a certain hue stimulus and represents their sensitivity. For example, the shorter the time required for a red stimulus to trigger an expression change, the more sensitive the user is to red;

[0128] The maximum response intensity is: the maximum intensity value in the color physiological response curve;

[0129] The maximum response intensity reflects the extreme value of emotional fluctuations and represents the user's maximum emotional response degree to a certain hue, which can be used to distinguish the driving ability of the stimulating color on the user's emotions;

[0130] The high-level maintenance time is the length of the time period during which the intensity value continuously remains above the high-intensity threshold;

[0131] The high-intensity threshold is 80% of the maximum response intensity;

[0132] The high-level maintenance time measures the persistence of the user's high-intensity emotional response and reflects the emotional regulation mechanism or preference stability. For example, a long-term high response may indicate a strong preference or aversion.

[0133] The specified feature parameters constitute an accurate description of the user's facial micro-expression dynamic response process, far superior to the simple method of only using the peak or average value, so it can significantly enhance the accuracy of user emotion modeling and the recognizability of personalized features.

[0134] S32. Integrate the specified feature parameters with the image interaction state vector to form a combined feature vector, and use the combined feature vector as the visual image personality vector.

[0135] In this embodiment, step S3 realizes the bridging from "physiology-emotion" to "personality-behavior" by fusing emotion response features and interaction contexts, effectively improving the depth and breadth of user modeling. The constructed visual image personality vector is more adaptable.

[0136] S4. Input the visual image personality vector into the image style parameter mapping network to map and obtain an image style control parameter set. The image generation and rendering module processes the pre-acquired three-dimensional reconstruction model of the user's face according to the image style control parameter set and outputs the generated personalized static image or dynamic image;

[0137] The described image processing process sequentially includes: color remapping and brightness adjustment, texture detection and refinement generation, edge enhancement and anti-aliasing processing, lighting simulation and shadow rendering, and expression and pose animation processing.

[0138] Through the above processing, the deep matching and efficient mapping from user multi-dimensional image data to image style control parameters can be realized. The generated images are closer to the user's psychological needs in terms of color style, content elements, etc., significantly improving the personalized generation quality. In summary, in this embodiment, by combining image processing technology, computer vision technology, and user psychological state modeling, an intelligent image generation system driven by the user state as the core is constructed.

[0139] Preferably, in some embodiments of the present application, refer to Figure 5 , S4 specifically includes:

[0140] S41. Input the visual image personality vector into the image style parameter mapping network to map it into an image style control parameter set (that is, a set of style control parameters);

[0141] In this step S41, the constructed visual image personality vector is input into the image style parameter mapping network for decoding. The image style parameter mapping network adopts a deep neural network architecture, and its typical structure may include fully connected layers, residual blocks, multi-head attention mechanisms, etc., for learning the complex non-linear mapping relationship between the visual image personality vector and the image style control parameters.

[0142] The introduction of the image style parameter mapping network solves the semantic gap problem in the "mapping of psychological features to image control parameters". On the one hand, it has the ability to decode high-dimensional psychological feature vectors, and on the other hand, through training, it can learn the structural rules in the style space to achieve the generation of highly personalized and style-coordinated control parameters, thereby greatly improving the accuracy and interpretability of image personalized generation.

[0143] Among them, the image style parameter mapping network is a deep neural network for mapping the visual image personality vector into the style control parameters required for image generation.

[0144] The image style parameter mapping network is a key module based on deep neural networks. Its core function is to convert the input visual image personality vector into style control parameters required for image generation. The visual image personality vector integrates the user's long-term visual preferences (such as color sensitivity, style inclination, etc.) and the current interaction emotional state (such as calm, excited, anxious, etc.), and has rich semantic information and psychological clues. The image style parameter mapping network uses a multi-layer fully connected neural network structure, combined with non-linear activation functions (such as ReLU, Leaky ReLU, etc.), to achieve deep decoding of these complex features, and thus outputs a set of structured style parameters. These parameters include color bias, texture control, light and dark contrast, emotional tone distribution, etc., which can finely adjust the behavior of the image generation model in terms of style expression.

[0145] In practical applications, the image style parameter mapping network can be flexibly docked with various image generation frameworks. For example, in generative adversarial networks represented by StyleGAN, after receiving the visual image personality vector, the image style parameter mapping network outputs multiple "style layer control vectors" to adjust the feature map distribution of each layer inside the generator, so as to achieve delicate style injection. In diffusion models, the image style parameter mapping network can provide conditional guidance of emotions or preferences for the feature fusion stage in the diffusion process. This structure not only improves the personalization degree of image generation, but also enhances the model's adaptability to different style requirements.

[0146] For example, assume that a user shows a color preference for "cool blue tones" and an interactive state of "mild anxiety" in the visual image personality vector. The image style parameter mapping network will decode a set of style parameters including high blue weight, enhanced texture fineness, and slightly reduced overall brightness. Subsequently, the image generation model will generate an image with cool colors, soft rhythm, and an overall atmosphere that fits the user's current mental state according to these parameters. This method effectively realizes the direct conversion from "mental state" to "image expression", significantly improving the experience and quality of personalized generation.

[0147] For example, a typical image style parameter mapping network consists of multiple key components, and each part works together to jointly complete the efficient mapping of the visual image personality vector to the image style control parameters. First is the input layer, whose main function is to receive the preprocessed visual image personality vector. This vector usually has a dimension of 64 to 512 and has integrated semantic features such as the user's long-term style preferences and the current emotional state. The input form can be a one-dimensional vector or a high-dimensional tensor combined with position information or emotional modulation to adapt to different style control requirements and context scenarios.

[0148] Next is the non - linear transformation layer, which is the core component of the image style parameter mapping network. This part is usually composed of 4 to 8 fully connected layers, followed by a non - linear activation function such as ReLU, Leaky ReLU, or GELU after each layer, which is used to enhance the network's ability to fit and express complex features. To improve the training stability and convergence speed of the network, layer normalization or batch normalization is introduced in some structures to ensure the consistency of feature distribution during training and prevent the problems of gradient vanishing or explosion.

[0149] Finally, after multiple layers of non - linear transformation, the information will enter the style control decoding layer, which is responsible for mapping high - dimensional features into style control parameters that can be used for image generation. The output form is flexible and diverse. It can be a set of style code vectors, which are used to drive the multi - layer feature modulation of generators such as StyleGAN; it can also be a modulation parameter tensor, which is used to control the feature map form of certain convolutional layers of the generation model; or a set of attention weights, which are used to adjust the proportion of different style features in the generation process. This output determines the performance of the final generated image in terms of color tendency, texture details, light and dark contrast, shape composition, etc., ensuring that the image style highly matches the user's psychological characteristics.

[0150] S42. Input the image style control parameter set into the image generation and rendering module. The image generation and rendering module uses an image generation model to process the three - dimensional reconstruction model of the user's face obtained in advance according to the image processing process through the image style control parameter set, and outputs the generated personalized static image or dynamic image.

[0151] The image generation model is a deep generation model based on a generative adversarial network or a diffusion model.

[0152] Specifically, the image generation model in this embodiment is composed of multiple functional modules working together to complete the whole process from the input vector to image generation.

[0153] First is the input module, which is used to receive style control parameters. These parameters are derived from the mapping result of the user's visual image personality vector and determine the style tendency, texture structure, and visual features of the final image.

[0154] Next is the feature mapping module, which is usually composed of multiple layers of fully connected networks, embedding layers, or lightweight encoders, and is used to convert the input style control parameters into style feature vectors in the latent space. These latent representations will further modulate the features of each level of the generated image inside the generator to achieve style - driven content shaping.

[0155] The core of image generation is the generator, whose network structure is usually composed of a series of convolutional layers, residual blocks, and upsampling modules, and is used to generate image details layer by layer, converting low-resolution latent features into high-resolution images. For generative models that support style control (such as StyleGAN), a style modulation module (StyleModulation) is introduced inside the generator, which injects style control parameters into the image generation process through normalization operations at specific layers (such as AdaIN, StyleNorm), dynamically adjusting visual attributes such as color, texture, and structure, so as to generate image content that better suits the user's preferences.

[0156] In models based on generative adversarial networks (GANs), there is also a discriminator, which is used to judge the authenticity of the generated images. The generator and the discriminator are continuously iteratively optimized through adversarial training to make the generated images more realistic. In the generative structure based on diffusion models, the generation of images is achieved through multiple rounds of "denoising" processes. Compared with traditional GAN models, diffusion models have obvious advantages in generation stability and detail expression, and are especially suitable for complex image content and high-resolution output.

[0157] In addition, to ensure the quality and style consistency of the generated images, the entire generation system relies on a set of carefully designed loss functions for training. Common loss functions include adversarial loss (used to improve image realism), perceptual loss (used to capture semantic levels), reconstruction loss (ensuring content fidelity), and style loss (maintaining style consistency), which jointly constrain the model output to meet the user's psychological preferences and aesthetic expectations.

[0158] The image generation model in this embodiment can be the StyleGAN2 model (Style-based Generative Adversarial Network v2) as a specific implementation. StyleGAN2 is a high-performance image generation model based on generative adversarial networks (GANs), with excellent image synthesis quality and style controllability. Compared with traditional generative networks, StyleGAN2 introduces an image style parameter mapping network and a layer-by-layer modulation mechanism, enabling the image generation process to not only express rich textures and details, but also achieve refined style adjustment through externally input style control parameters.

[0159] In StyleGAN2, the visual image personality vector is first transformed into a style code in the latent space through an image style parameter mapping network. Then, this code is injected into multiple levels of the generator, and the distribution of feature maps in each layer is modulated through adaptive normalization (such as AdaIN) to achieve precise control over visual attributes such as image color, shape, structure, and texture. Through this mechanism, StyleGAN2 can synthesize static images, dynamic images, or virtual character appearances with diverse styles and significant personalization according to the visual personality characteristics of different users.

[0160] The StyleGAN2 model has high expressiveness and controllability. The generated image resolution can reach 1024×1024, with rich details and consistent styles, and is widely used in fields such as virtual character modeling, anime style conversion, face generation, and digital art creation. In this embodiment, the image generation system constructed based on StyleGAN2 can significantly improve the matching degree between the generated result and the user's psychological preference, meeting the diverse needs for personalized image content.

[0161] Optionally, in some embodiments of the present application, the image generation and rendering module uses an image generation model based on a generative adversarial network or a diffusion model. For the image style control parameter set, color remapping and brightness adjustment, texture feature refinement, edge structure optimization and anti-aliasing, light and shadow mapping simulation, and facial expression and pose-driven rendering are sequentially performed in the multi-layer perception channel to achieve personalized image synthesis.

[0162] Optionally, in some embodiments of the present application, S14 specifically includes:

[0163] S141: Use an image segmentation model to process the collected user facial image, accurately extract the hair area, and generate a high-resolution hair mask image;

[0164] The image segmentation model is a multi-scale semantic segmentation network;

[0165] In this embodiment, in step S141, an image segmentation model is used to process the collected user facial image, accurately separate the hair region in the image, and generate a corresponding high-resolution hair mask image. The image segmentation model can be a deep neural network with multi-scale semantic modeling capabilities, such as U-Net, DeepLab, or HRNet, etc. These models utilize technologies such as encoder-decoder structures, dilated convolutions, and skip connections to effectively enhance the segmentation ability for complex textures and edge details in the hair region. Achieving high-quality hairstyle mask extraction lays a clear and accurate image foundation for subsequent edge extraction and geometric analysis, significantly improving the ability to restore the structure of the user's original hairstyle. S142. Based on the hair mask image, use a contour detection algorithm to extract the outer edge line of the hairstyle;

[0166] Based on the generated hair mask image, apply a contour detection algorithm (such as Canny, Sobel, edge tracking algorithm, etc.) to extract the outer edge line of the hairstyle region. This contour line represents the overall morphological boundary of the user's current hairstyle, including key information such as hair strand distribution, hairline contour, and hair tip trend. Through precise edge extraction, this step can construct a geometric representation framework for the hairstyle, making subsequent modeling more structured and comparable. S143. Conduct geometric feature analysis on the extracted hairstyle edge curve, and extract contour parameters characterizing the hairstyle form. The contour parameters include:

[0167] Contour length: The total length of the hairstyle edge curve, used to measure the degree of expansion of the hairstyle and reflect the overall coverage of the hair;

[0168] Curvature: The curvature change value within the unit length of the edge curve, used to reflect the tortuous degree of the edge curve, describe the curvature change per unit length of the edge curve, and quantify the tortuous degree of the hair edge, such as the distinction between straight hair and curly hair;

[0169] Closure: The degree of connectivity between the head and tail of the edge curve, used to evaluate the possibility of forming a closed figure, judge the connection degree between the head and tail of the edge curve, and evaluate whether the hairstyle forms a closed structure. For example, judge whether the bangs surround the forehead area, etc.

[0170] Converting the complex visual information in the original image into quantifiable and comparable structural features helps the expression of the hairstyle form in the feature space.

[0171] S144. Normalize or standardize the contour parameters to eliminate scale differences, and further combine the standardized contour parameters to form a hairstyle contour feature vector for comprehensively describing the shape category and structural features of the user's hairstyle.

[0172] In step S144, the extracted contour parameters are normalized or standardized to eliminate the scale differences between the parameters and ensure the numerical stability and expression consistency of subsequent model learning. After processing, these standardized parameters are combined in multiple dimensions to generate a hairstyle contour feature vector. This vector, as part of the user appearance feature vector, is linked with the image style parameter mapping network and the image generation model to control the expression and consistency of the hairstyle in the generated image.

[0173] The present application discloses a personalized content generation method based on user image data processing, which relates to the fields of computer vision, image recognition and synthesis, and belongs to the technical category of the international patent classification G06T (image data processing or generation, especially image analysis or recognition performed by a computer).

[0174] This method collects and processes images of the user's face and the environment in which he is located, and uses image processing algorithms and multimodal deep learning models to model and identify the user's current interaction state, thereby achieving context-awareness-driven content generation optimization. It has multiple technical innovations and significant application effects.

[0175] First, the solution collects user facial images and environmental features, combines gesture recognition, expression recognition, hairstyle segmentation, key point detection and other multi-model collaborative analysis, and constructs an image interaction state vector, which can accurately reflect the user's psychological and physiological state in real time and improve the personalization and adaptability of generated content. Compared with the traditional method of roughly understanding the user's state and lagging response, this solution achieves fine recognition of user attention, emotion, and participation through high-dimensional information fusion and deep feature modeling.

[0176] Secondly, the solution innovatively introduces color stimulus image sequences and real-time facial micro-expression collection mechanisms to automatically establish user color response curves, overcoming the defect that traditional subjective questionnaire methods are difficult to truly reflect user preferences, and realizing emotion-driven personalized content generation. In addition, the use of deep learning technologies such as CNN ensures recognition accuracy and robustness, and has strong practicality and scalability.

[0177] In summary, the embodiments of the present application significantly enhance the emotional perception ability of human-computer interaction, breaking through the bottlenecks of the prior art in terms of single dimension of user status recognition and insufficient personalization capabilities.

[0178] In addition, an embodiment of the present application further proposes a computer device, which includes a processor and a memory, wherein the processor is used to execute instructions stored in the memory so that the computer device executes the content generation method based on the user image described in the above embodiment.

[0179] Finally, the embodiment of the present application also proposes a computer-readable storage medium, including computer program instructions, which, when executed by a processor, implement the content generation method based on user images described in the above embodiments.

[0180] Embodiment 2

[0181] This embodiment discloses a content generation method based on user images, aiming to generate highly personalized static avatars or virtual character appearances for users. This method not only captures user images, but also integrates their facial expressions, head postures, hairstyle contours, and emotional responses to color stimuli, and finally generates a virtual image that conforms to the user's style and emotional characteristics.

[0182] In the first step, the user's facial image is captured by a front camera, and at the same time, the environmental perception module is used to record the light intensity and color temperature of the current environment to obtain accurate image input conditions. In the image processing stage, the following analysis operations are performed in sequence:

[0183] Use a model based on convolutional neural network (CNN) (such as MediaPipe FaceMesh) to identify 68 facial key points including the corners of the eyes, the corners of the mouth, and the tip of the nose, providing a positioning basis for subsequent action recognition;

[0184] Adopt a pose recognition algorithm such as OpenPose to calculate the user's head angles (pitch, yaw, roll) and shoulder postures;

[0185] Separate the hair area through a multi-scale semantic segmentation network to generate a hair mask map;

[0186] Based on the hair mask map, use a contour detection algorithm to extract the hairstyle edge, further calculate its length, curvature, closure, etc. indicators, and perform normalization processing to form numerical hairstyle contour features;

[0187] Apply a CNN expression recognition model to identify facial expression action units such as AU1 (inner brow raise), AU6 (cheek raise), AU12 (lip corner stretch), etc. and their intensity values;

[0188] After the above feature data are aligned according to the time stamp, they are respectively encoded into feature embedding vectors and input into the fusion network to output a unified image interaction state vector.

[0189] In the second step, to deeply explore the user's emotional response to visual stimuli, a set of images with different hues (such as cold blue, warm orange, light green, deep red, etc.) are shown to the user, each frame lasting about 1 second and continuously playing 10 frames. During this process:

[0190] Record the user's facial reactions to each frame of color image in real time through the camera, and identify facial micro-expressions such as raised eyebrows and raised corners of the mouth and their intensities;

[0191] Taking each frame of image as the time reference point, plot the intensities of the extracted expression action units as multiple time-intensity response curves;

[0192] Fuse each response curve according to the preset weights (such as 0.5 for raised corners of the mouth, 0.3 for frowning, and 0.2 for zygomatic muscle elevation) to form a physiological response curve representing the user's color emotion preference.

[0193] In the third step, based on the above color physiological response curve, further extract key parameters:

[0194] Response start time: For example, a certain action (such as raised corners of the mouth) first exceeds the response threshold in the 3rd frame of the image;

[0195] Maximum response intensity: Record the peak intensity of key expression actions such as AU12;

[0196] High-level maintenance time: Statistically calculate the time interval during which a certain intensity (such as ≥0.68) lasts, such as maintaining for 1.2 seconds.

[0197] These parameters describing the user's emotional response pattern are concatenated and fused with the previously generated image interaction state vector to generate a low-dimensional visual image personality vector. This vector summarizes the user's emotional preference, style characteristics, and appearance performance tendency in the current state.

[0198] In the fourth step, input the visual image personality vector into an image style parameter mapping network (such as a fully connected neural network containing BatchNorm and ReLU activation units), and this network outputs multi-dimensional style control parameters, including but not limited to:

[0199] Color preference: Such as a cool color system preference value of 0.3 and a warm color system preference value of 0.7;

[0200] Expression brightness: Such as "high" representing a clear and lively expression style;

[0201] Hair style complexity: Such as a high curvature of the hair contour, indicating a preference for curly or fluffy hair styles.

[0202] These style parameters will be used as inputs into an image generation model (such as advanced image generators like StyleGAN2 or Stable Diffusion) to synthesize a virtual avatar image with strong realism and a style matching the user's characteristics. For example: If the user shows a strong positive reaction to warm orange tones, the generated avatar may have a warm color background, natural curly hair, and a slightly smiling expression with raised corners of the mouth, fully reflecting the user's personality style.

[0203] This embodiment realizes starting from the user's real image, combining facial micro-expressions, physiological responses, and environmental characteristics to comprehensively model the user's appearance and emotional preferences, and drive the image generation process. The finally output virtual image has highly personalized features, which not only truly reflects the user's physical appearance characteristics but also incorporates their potential style tendencies and emotional expressions.

[0204] In another specific embodiment, after obtaining the visual image personality vector, the visual image personality vector is input into the image style parameter mapping network. This mapping network is a type of deep neural network model, and its function is to perform feature extraction and non-linear mapping on the input personality vector, thereby outputting a set of control parameters representing the image style, which is called the image style control parameter set.

[0205] The visual image personality vector integrates the user's key facial features (such as hairstyle, expression, posture), environmental characteristics, and the user's physiological response characteristics to color stimuli. Therefore, it can reflect multi-dimensional information such as the user's personality, emotion, and style preferences. The mapping network models these high-dimensional semantic information, learns the mapping relationship in the image style preferences among different users, and thus generates a parameter set that can drive the change of the image style.

[0206] The training data of the image style parameter mapping network can include a large number of user image samples and their corresponding style labels. Using a supervised or self-supervised learning mechanism, the network can stably output controllable and personalized style parameters through the constraint of the loss function.

[0207] Next, the image style control parameter set is input into the image generation and rendering module. This module is based on a deep generation model, uses the parameter set as style guidance information, and performs image synthesis and stylization processing on the pre-obtained three-dimensional reconstruction model of the user's face.

[0208] Specifically, the image generation and rendering module includes an image generation model, and this model can be one of the following types:

[0209] Generative Adversarial Network (GAN): Through adversarial training, the generator can learn to generate realistic personalized images under the supervision of the discriminator;

[0210] Diffusion Model: By modeling the process of gradually adding noise and denoising to the image, high-quality and high-fidelity image synthesis is achieved.

[0211] During the image generation and rendering process, for the user's three-dimensional reconstruction model, the following image processing procedures are sequentially performed:

[0212] Color remapping and brightness adjustment: Adjust the overall color tone and lighting atmosphere according to style parameters;

[0213] Texture detection and refinement generation: Generate detailed textures for areas such as the face, hairstyle, and clothing;

[0214] Edge enhancement and anti-aliasing processing: Improve image clarity and naturalness;

[0215] Lighting simulation and shadow rendering: Enhance the sense of spatial three-dimensionality and realism;

[0216] Expression and pose animation processing: Dynamically restore the user's emotional expressions or head movements in dynamic images.

[0217] The final output is a static or dynamic image with highly personalized features. This image not only retains the structural information of the user himself but also reflects his unique preferences in visual style and emotion.

[0218] In the description of this application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of these features. In the description of this application, "a plurality" means two or more unless otherwise specifically defined.

[0219] In this application, unless otherwise clearly defined and limited, terms such as "installed", "connected", "connected to", "fixed" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium; it can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0220] In this application, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature can be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature can be that the first feature is directly above or diagonally above the second feature, or merely indicates that the first feature has a higher horizontal height than the second feature. The first feature being "under", "below" and "beneath" the second feature can be that the first feature is directly below or diagonally below the second feature, or merely indicates that the first feature has a lower horizontal height than the second feature.

[0221] In the description of this specification, the descriptions of terms such as "one embodiment", "some embodiments", "embodiment", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0222] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for generating content based on user images, characterized in that, It includes the following steps: S1. Collect the user's facial image through an image acquisition device, and collect the environmental characteristics where the user is located through an environmental perception module. Extract the facial key points, hairstyle features, expression intensity parameters, and posture features of the user in the facial image, and generate an image interaction state vector in combination with the environmental characteristics; S2. Present a sequence of color stimulus images with different hues to the user, and synchronously collect the user's facial image data during the presentation process. Extract the user's facial micro-expressions in the facial image data, establish a facial expression response curve, and further generate a color physiological response curve; S3. Extract specified feature parameters based on the color physiological response curve, and fuse them with the image interaction state vector to generate a visual image personality vector for controlling the image style; S4. Input the visual image personality vector into an image style parameter mapping network to map and obtain an image style control parameter set. The image generation and rendering module processes the pre-acquired three-dimensional reconstruction model of the user's face according to the image processing flow, and outputs the generated personalized static image or dynamic image; The image processing flow sequentially includes: color remapping and brightness adjustment, texture detection and refinement generation, edge enhancement and anti-aliasing processing, lighting simulation and shadow rendering, expression and posture animation processing.

2. The content generation method based on user images according to claim 1, wherein, The specific content of S1 includes: S11. Use an image acquisition device to collect the user's facial image; S12. Analyze the collected user's facial image through a posture recognition algorithm to extract the user's posture features; The posture features include: head posture angle and body posture information; S13. Use a facial key point detection model based on a convolutional neural network to identify the positions of facial key points in the user's facial image; The facial key points include: facial contour, eyes, nose, mouth; S14. Use an image segmentation model to identify the hair area in the user's facial image and obtain a hairstyle contour feature vector; S15. Use an expression recognition model to identify the expression action units included in the user's facial image, and record the intensity values of each expression action unit as expression intensity parameters; S16. Synchronously align the user's posture features, facial key points, hairstyle contour feature vector, expression action unit intensity parameters with the environmental characteristics at the same timestamp, respectively input them into a feature embedding encoder for normalization processing, and generate a low-dimensional joint feature vector through a feature fusion network as the image interaction state vector representing the current user interaction state; The environmental characteristics include: light intensity, color temperature.

3. The content generation method based on user images according to claim 2, wherein The specific content of S14 includes: S141. Use an image segmentation model to process the collected user's facial image, accurately extract the hair area, and generate a high-resolution hair mask image; The image segmentation model is a multi-scale semantic segmentation network; S142. Based on the hair mask image, use a contour detection algorithm to extract the outer edge line of the hairstyle; S143. Conduct geometric feature analysis on the extracted hairstyle edge curve, and extract contour parameters representing the hairstyle form. The contour parameters include: Contour length: The total length of the hairstyle edge curve; Curvature: The value of the change in curvature per unit length of the edge curve, which is used to reflect the tortuosity of the edge curve; Closure: The degree of connectivity between the head and tail of the edge curve, which is used to evaluate the possibility of forming a closed figure; S144. Normalize or standardize the contour parameters to eliminate scale differences, and further combine the standardized contour parameters to form a hairstyle contour feature vector for comprehensively describing the hairstyle shape category and structural characteristics of the user.

4. The method for generating content based on a user image according to claim 3, wherein The expression action units include: inner eyebrow raising action unit, outer eyebrow raising action unit, frowning action unit, upper eyelid lifting action unit, zygomaticus major muscle lifting action unit, eyelid closing action unit, nose wrinkling action unit, upper lip lifting action unit, mouth corner raising action unit, mouth corner pulling down action unit; The expression recognition model is a convolutional neural network model.

5. The content generation method based on user images according to claim 4, wherein The S2 specifically includes: S21. Continuously present a sequence of color stimulus images with different hues to the user. During the presentation of the color stimulus images, use an image acquisition device to collect the facial image data of the user in real time; S22. Input the collected facial image data into the expression recognition model to extract the facial micro-expression features of the user. The facial micro-expression features include multiple recognized expression action units and their corresponding intensity values; S23. According to the intensity values of the expression action units in each frame of the color stimulus image sequence and the corresponding time points, establish a set of facial expression response curves of the user; During the establishment of the facial expression response curves of the user, use the intensity value of each expression action unit as the vertical axis and the time sequence of the color stimulus images as the horizontal axis to separately establish a facial expression response curve for each expression action unit; S24. Based on the facial expression response curve corresponding to each expression action unit, construct a color physiological response curve reflecting the user's color physiological response under different hue color stimulus conditions.

6. The method for generating content based on a user image according to claim 5, wherein Among them, The color physiological response curve is obtained by weighted fusion of the response curves of all expression action units according to a preset weight.

7. The content generation method based on user images according to claim 6, wherein, The S3 specifically includes: S31. Extract the specified feature parameters in the color physiological response curve; The specified feature parameters include: response start time, maximum response intensity, high-level maintenance time; Among them, the response start time is: with the time axis on the curve as the horizontal axis, the time point corresponding to when the intensity value first exceeds a preset threshold; The maximum response intensity is: the maximum intensity value in the color physiological response curve; The high-level maintenance time is the length of the time period during which the intensity value continuously remains above the high-intensity threshold; The high-intensity threshold is 80% of the maximum response intensity; S32. Fuse the specified feature parameters with the image interaction state vector to form a combined feature vector, and use the combined feature vector as the visual image personality vector.

8. The content generation method based on user images according to claim 7, characterized in that The S4 specifically includes: S41. Input the visual image personality vector into the image style parameter mapping network to map it into an image style control parameter set; S42. Input the image style control parameter set into the image generation and rendering module. The image generation and rendering module processes the three-dimensional reconstruction model of the user's face obtained in advance according to the image processing flow by using the image generation model through the image style control parameter set, and outputs the generated personalized static image or dynamic image; Among them, the image style parameter mapping network is a deep neural network for mapping the visual image personality vector into the style control parameters required for image generation, and the image generation model is a deep generation model based on a generative adversarial network or a diffusion model.

9. An electronic device, characterized in that, It includes a memory and a processor, where: The memory is used to store computer programs; The processor is used to execute the computer program to implement the method for generating content based on user images according to any one of claims 1 to 8.

10. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to implement the method for generating content based on user images according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Virtual character image construction device and method

    CN112164135A

  • Image cartoonalization method and device

    CN113470147A

  • Method and device for acquiring emotion data

    CN109730701A

  • Quantifiable face model reconstruction method and system

    CN117315154A

  • Three-dimensional digital human generation and interaction method and system

    CN117496072A

Cited By

  • Virtual image generation method based on machine learning

    CN120726194A

  • Intelligent photo system based on crowd classification

    CN121074523A