Text-guided three-dimensional face reconstruction and emotion analysis method, device and equipment

By combining a text-guided 3D face reconstruction method with attribute matrices generated from natural language and image features, the problems of privacy leakage and reconstruction failure in 3D face reconstruction are solved. This method achieves high-precision facial semantic consistency and emotion recognition, and is applicable to the diagnosis of mental illness and telemedicine.

CN120747417BActive Publication Date: 2025-11-21ZHEJIANG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511156807.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-21
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing 3D face reconstruction technology poses privacy risks in the diagnosis of mental illnesses. Reconstruction may fail or features may be lost when image quality is limited. Multimodal reconstruction schemes are difficult to align modally, making it difficult to achieve high-precision facial semantic consistency.

Method used

A text-guided 3D face reconstruction method is adopted. By jointly encoding the attribute matrix generated by natural language with image features, and combining a multimodal model and a semantic consistency alignment module, the geometric reconstruction accuracy and emotion recognition accuracy are improved, and a 3D face model with semantic consistency is generated.

Benefits of technology

It significantly improves the accuracy of geometric reconstruction and emotion recognition in complex scenarios, reduces the risk of privacy leakage, and is suitable for remote psychological intervention and intelligent medical tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747417B_ABST
    Figure CN120747417B_ABST
Patent Text Reader

Abstract

The application discloses a kind of three-dimensional face reconstruction and emotion analysis method, device and equipment based on text guide, comprising the following steps: obtaining the facial image information of human body and pre-processing and marking, construct training set;Three-dimensional face reconstruction model is trained based on text guide using training set, the face three-dimensional grid is obtained by inputing the face image to be analyzed into the three-dimensional shape model trained, and the face three-dimensional grid is input into emotion analysis model to judge face emotion category.The attribute matrix generated by natural language is jointly coded with image features and input into the three-dimensional reconstruction framework, which can significantly improve the geometric reconstruction accuracy in complex scenarios such as blurred face, local occlusion or dramatic / fine changes in facial expression, thereby effectively alleviating the problem of reconstruction failure or feature loss under non-ideal image conditions in traditional methods.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a three-dimensional face reconstruction and emotion analysis method, device and equipment based on text guidance. BACKGROUND

[0002] With the deep integration of artificial intelligence and medical image analysis, three-dimensional face reconstruction technology has gradually shown broad application potential in the medical field, especially in the auxiliary diagnosis of mental and psychological diseases. Depression, as a high-incidence mental disease, its non-invasive and objective recognition and diagnosis has always been the focus of academic and clinical attention. Abnormal changes in facial expressions and facial muscle activity are often considered as one of the important manifestations of depression.

[0003] Chinese patent document with publication number CN119786021A discloses a multi-modal multi-factor depression recognition system fusing emotional information. However, in sensitive scenarios such as mental disease assessment, directly collecting and transmitting original images, videos or voice data poses potential privacy leakage risks.

[0004] In actual medical environments, there are often complex collection conditions such as image blur, extreme illumination, and partial occlusion, and patients may have abnormal expression states or subtle facial dynamics (such as micro-expression). Existing methods mainly rely on low-dimensional parameterized models (such as 3DMM), which are limited by insufficient model expression ability, difficulty in capturing facial detail features, and difficulty in effectively expressing complex facial deformations. The reconstructed results are prone to distortion of key structural details, affecting the geometric accuracy and expression feature fidelity. This is a significant gap from the demand for high-precision facial features in medical diagnosis.

[0005] Existing partial multi-modal reconstruction schemes attempt to introduce audio or other auxiliary modalities to improve modeling accuracy. For example, Chinese patent document with publication number CN118800274A discloses an AI digital human automatic expression generation system based on voice driving, including a voice expression generation module, a facial expression database, an expression feature extraction module, and a facial three-dimensional reconstruction module. However, there are often problems such as non-uniform embedding space and weak feature matching in the modal alignment process. For example, the semantic information of the voice is hidden in the audio signal, the voice expression is more ambiguous and implicit, and it is difficult to accurately describe the facial details. The semantic association between visual information is weak, making it difficult to achieve collaborative optimization between geometric shapes, expression parameters and texture features, thereby affecting the semantic consistency and expression integrity of the overall reconstruction effect. It is difficult to achieve efficient fusion of multi-modal semantic information, limiting its application and promotion in fine tasks such as facial recognition and emotion understanding.

[0006] Therefore, a three-dimensional face reconstruction system for auxiliary diagnosis of mental diseases is needed, which can effectively fuse multi-modal information (text and vision), improve the facial semantic expression ability, and enhance the modeling accuracy of complex emotional states. SUMMARY

[0007] The application provides a three-dimensional face reconstruction and emotion analysis method, device and equipment based on text guidance, which improves the accuracy of three-dimensional face reconstruction under limited image information conditions and realizes the auxiliary diagnosis of mental diseases under the premise of protecting patient privacy.

[0008] The technical solutions of the application are as follows:

[0009] A three-dimensional face reconstruction and emotion analysis method based on text guidance comprises the following steps:

[0010] (1) Obtain facial image information of a human body and perform preprocessing and labeling to construct a training set;

[0011] (2) Train a three-dimensional face reconstruction model using the training set, comprising:

[0012] (2-i) Disassemble the facial image in the training set into a text description T containing multiple facial attribute information, and convert the text description T into a structured attribute matrix P attr ;

[0013] (2-ii) Input the facial image in the training set into a three-dimensional shape model to perform three-dimensional face reconstruction, and introduce the attribute matrix into the three-dimensional face reconstruction process to obtain a face three-dimensional mesh with semantic consistency ;

[0014] (2-iii) Input the facial image in the training set into an image style generator to synthesize a texture image, and introduce the attribute matrix into the texture image synthesis process to obtain a texture image with semantic consistency I UV ;

[0015] (2-iv) Combine the face three-dimensional mesh and the texture image I UV , and generate a rendered image after rendering;

[0016] (2-v) Construct a loss function between the rendered image and the original facial image, and between the rendered image and the text description, respectively, and perform training;

[0017] (3) Input a face image to be analyzed into the trained three-dimensional shape model to obtain a face three-dimensional mesh , and combine the face three-dimensional mesh The input emotion analysis model judges a facial emotion category.

[0018] The three-dimensional structure reconstruction of a face not only needs to accurately restore the overall contour of the face, but also needs to capture the dynamic changes of expression muscles at a fine-grained level to support high-precision tasks such as emotion recognition and auxiliary diagnosis of mental illness. The present application significantly enhances the semantic alignment capability between facial structure and emotional attributes in three-dimensional face modeling by constructing a collaborative modeling mechanism that integrates text description and image visual features. In the case of limited image quality or significant pose changes, stable and high-fidelity three-dimensional reconstruction results can still be achieved. It is particularly suitable for application scenarios such as auxiliary diagnosis of mental illness and remote psychological intervention.

[0019] In step (1), the labeled content includes basic attribute classes, emotional expression classes, and geometric structure classes; the basic attribute classes include gender, age, etc.; the emotional expression classes include expression categories (happy, sad, angry, surprised, frightened, disgusted, neutral, confused, etc.), expression intensity scores (continuous values between 0 and 1, used to express the intensity of emotion); the geometric structure classes include face contour types (square / ellipse / inverted triangle / long face, etc.), basic shapes of facial features (thickness, width, height, angle, etc.), hair features (mustache density, mustache color, mustache contour, etc., including sideburns, eight-character mustache, goat mustache, etc.), etc.

[0020] In step (1), the preprocessing includes data cleaning and enhancement operations, including but not limited to: image uniform size scaling and standardization processing; face region cropping and boundary detection; illumination condition adjustment and contrast enhancement; Gaussian filtering, edge smoothing, and other noise reduction strategies.

[0021] Step (2-i) includes:

[0022] (2-ia) Construct a set of face semantic-related queries Q = {q1, q2, …, q M}, where each query q i corresponds to a specific attribute dimension of the face;

[0023] (2-ib) For the input face image I, call the pre-trained image-text question answering model to jointly process the image and the query pair, and generate the corresponding natural language response set A = {a1, a2, …, a M}, where a i = BLIP(I, q i );

[0024] (2-ic) Use the language template T tem set A to organize the text description T containing multiple facial attribute information;

[0025] (2-id) obtaining a global semantic embedding vector z of the text description T by a multi-modal model T , and mapping the global embedding vector z T into a structured attribute matrix P attr .

[0026] In step (2-id), the multi-modal model can be a CLIP (Contrastive Language-Image Pre-training) model.

[0027] In step (2-id), the global embedding vector z T is mapped into a structured attribute matrix P attr :

[0028] P attr = F MLP (z T );

[0029] , K denotes the number of attribute categories, C denotes the number of possible discrete states of each attribute category; the first row denotes the state distribution of the first k attribute.

[0030] Step (2-ii) comprises:

[0031] (2-iia) inputting the face image into a multi-channel feature encoding network to obtain individual feature parameters , dynamic expression parameters and head pose parameters ;

[0032] (2-iib) inputting the attribute matrix P attr into nonlinear conversion modules and respectively, to obtain a semantic modulation vector m s related to geometric shape modeling m e :

[0033] m s = σ(Γ s ( Pattr )), m e = σ(Γ e ( P attr ));

[0034] wherein denotes a Sigmoid activation function;

[0035] (2-iii) performing element-wise multiplication operation between the individual feature parameters and the semantic modulation vector to obtain the modulated individual parameters ; and performing fusion between the expression parameters and the semantic modulation vector to generate the modulated dynamic parameters :

[0036] , ;

[0037] (2-iv) performing three-dimensional shape model decoding on the , , to obtain the face three-dimensional mesh with semantic consistency .

[0038] Step (2-iii) comprises:

[0039] (2-iiia) inputting the face image into the encoding network to extract an initial latent vector z;

[0040] (2-iiib) performing fusion between the attribute matrix P attr and the initial latent vector z to form a joint feature expression u = f(z, P attr ), wherein denotes a nonlinear fusion mapping function;

[0041] (2-iiic) taking the joint feature expression u as the input of the StyleGAN2 model to output the texture image with semantic consistency I UV .

[0042] Step (2-v) comprises:

[0043] (2-va) generating K different block-level random mask views for the rendered image, inputting the K mask views into the image-text encoding model to obtain visual representations , and finally performing aggregation to obtain the image feature vector :

[0044] ;

[0045] (2-vb) inputting the text description T into the text-image encoding model to obtain a text feature vector b ;

[0046] (2-vc) constructing a bidirectional cross-contrast loss function between the image and the text l align :

[0047] ;

[0048] τ is a temperature factor; B is the number of samples; , is a similarity score from image to text; , is a similarity score from text to image;

[0049] (2-vd) constructing an image reconstruction loss function l rec :

[0050] ;

[0051] λ 1、 λ 2、 λ 3、 λ 4 is a hyperparameter weight; l pix 、 l land 、 l reg 、 l emo respectively correspond to pixel error loss, landmark position error loss, regularization loss, and emotion consistency error loss;

[0052] After training, a trained three-dimensional shape model is obtained.

[0053] The pixel error loss measures the difference between the rendered image and the original face image in the pixel space; the landmark position error is based on the Euclidean distance constraint between two-dimensional and three-dimensional face key points; the regularization loss imposes a priori restriction on the model parameters or representation space; the emotion consistency error extracts emotion features with a pre-trained expression recognition model, and minimizes the semantic difference with the target image.

[0054] In step (3), the emotion analysis model is a multi-layer perceptron (MLP).

[0055] To adapt to the discrimination demand of weak and atypical expressions in mental illness scenes, preferably, a pathological sample dataset is introduced in the training stage to fine-tune the emotion analysis model, improve the recognition accuracy of slight emotional changes, and be applicable to intelligent medical tasks such as remote psychological intervention and emotion disorder monitoring.

[0056] The application further provides a text-guided three-dimensional face reconstruction and emotion analysis device, comprising:

[0057] A data acquisition module acquires facial image information of a human body and performs preprocessing and labeling to construct a training set;

[0058] A model training module trains a three-dimensional shape model using the training set, comprising:

[0059] The facial images in the training set are deconstructed into text descriptions containing multiple facial attribute information, and the text descriptions are converted into structured attribute matrices;

[0060] The facial images in the training set are input into a three-dimensional shape model for three-dimensional face reconstruction, and the attribute matrices are introduced into the three-dimensional face reconstruction process to obtain a face three-dimensional mesh with semantic consistency ;

[0061] The facial images in the training set are input into an image style generator to synthesize texture images, and the attribute matrices are introduced into the texture image synthesis process to obtain texture images with semantic consistency I UV ;

[0062] The face three-dimensional mesh is combined with the texture images I UV to render a rendered image;

[0063] Loss functions between the rendered image and the original facial image, and between the rendered image and the text description are constructed respectively, and training is performed;

[0064] A three-dimensional face reconstruction and emotion analysis module stores a trained three-dimensional shape model and an emotion analysis model, the three-dimensional shape model generates a face three-dimensional mesh according to a face image to be analyzed, and the emotion analysis model judges a face emotion category according to the face three-dimensional mesh .

[0065] The application further provides a text-guided three-dimensional face reconstruction and emotion analysis device, comprising:

[0066] An image acquisition unit is configured to acquire facial image information of a user;

[0067] The face reconstruction unit stores a three-dimensional shape model trained by the above method, and the three-dimensional shape model obtains a face three-dimensional mesh according to face image information collected by the image collection unit ;

[0068] The emotion analysis unit stores a trained emotion analysis model, and judges the face emotion category according to the face three-dimensional mesh

[0069] The data storage unit is used for storing face three-dimensional reconstruction and emotion classification results

[0070] The man-machine interaction visualization unit is used for outputting face three-dimensional reconstruction and emotion classification results

[0071] The communication and control unit is used for information transmission and control between the processing unit and the man-machine interaction visualization unit.

[0072] The three-dimensional face reconstruction and emotion analysis device based on text guidance of the application is suitable for intelligent medical tasks such as remote psychological intervention and emotion disorder monitoring.

[0073] Compared with the prior art, the application has the following beneficial effects:

[0074] (1) The attribute matrix generated by natural language is jointly encoded with image features and input into a three-dimensional reconstruction framework, which can significantly improve the geometric reconstruction accuracy in complex scenes such as face blur, local occlusion or dramatic / subtle changes in facial expressions, thereby effectively alleviating the problem of reconstruction failure or feature loss of traditional methods under non-ideal image conditions.

[0075] (2) The application performs fine-grained semantic modeling on the face region under the multi-view mask mechanism, enhances the coupling strength between the image and the text through semantic-driven forward and backward matching, makes the corresponding relationship between facial expression changes and geometric morphology more clear, and improves the accuracy and interpretability of emotion recognition.

[0076] (3) The three-dimensional parametric reconstruction output used in the application can be used as a structured medical feature expression, replacing the traditional image or video form of patient face information transmission, significantly reducing the risk of original privacy image leakage. The generated three-dimensional reconstruction result not only retains the face structure and expression dynamics, but also can be used for clinical auxiliary analysis and remote medical diagnosis, taking into account data security and application practicality, realizing effective coordination between data privacy protection and functional processing in medical scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0077] Figure 1 It is a flowchart of the face reconstruction and emotion analysis method

[0078] Figure 2 ​A structural schematic diagram of a face reconstruction and emotion analysis device. DETAILED DESCRIPTION

[0079] The application will be described in further detail below with reference to the drawings and embodiments, it should be pointed out that the following embodiments are intended to facilitate the understanding of the application and do not limit the application in any way.

[0080] Compared with the three-dimensional face modeling scheme relying on a single visual mode in the prior art, the application provides a multi-modal reconstruction method fusing text semantic information. By jointly encoding attribute semantic vectors generated by natural language and image features and inputting them into a three-dimensional reconstruction framework, the geometric reconstruction accuracy in complex scenes such as blurred face, local occlusion or dramatic / subtle changes in facial expression can be significantly improved, thereby effectively alleviating the problem of reconstruction failure or feature loss of traditional methods under non-ideal image conditions.

[0081] Further, the semantic consistency alignment module proposed in the application performs fine-grained semantic modeling on the face region under the multi-view mask mechanism, enhances the coupling strength between the image and the language through semantic-driven forward and reverse matching, makes the corresponding relationship between the facial expression changes and the geometric morphology clearer, and thus improves the accuracy and interpretability of emotion recognition.

[0082] In addition, the three-dimensional parametric reconstruction output used in the application can be used as a structured medical feature expression, replacing the transmission of patient facial information in the form of traditional images or videos, and significantly reducing the risk of leakage of original privacy images. The generated three-dimensional reconstruction result not only retains the facial structure and expression dynamics, but also can be used for clinical auxiliary analysis and remote medical diagnosis, taking into account data security and application practicality, and realizing effective coordination between data privacy protection and functional processing in medical scenarios.

[0083] A three-dimensional face reconstruction and emotion analysis method based on text guidance, the steps of which are:

[0084] S100 system building: first, a composite computing system for face three-dimensional reconstruction and diagnostic analysis is built. The system includes: an image acquisition unit based on an imaging sensor, used to acquire facial visual image information, preferably using an RGB digital camera device; a model processing module, used to execute three-dimensional reconstruction algorithms and subsequent emotion analysis tasks; and a human-computer interaction visualization unit for result display, used to output the reconstruction model and analysis results.

[0085] S101 data collection: the dynamic image information of the user's face is acquired by the imaging module described in S100. In order to control the memory resource occupation and avoid the time sequence redundancy of image frames, the application adopts an interval frame extraction strategy, periodically extracts key frames for subsequent analysis, thereby reducing the processing cost while ensuring the diversity of facial features.

[0086] After the collection is completed, preliminary cleaning is performed, including but not limited to: removing images with severe blur / overexposure / over-shading, and removing images without complete facial structure (e.g., only half of a face).

[0087] The obtained facial image samples are formed into a structured data set through manual and automatic labeling processes, and are divided into samples in a set proportion, preferably 80% as a training data set and 20% as an evaluation test set, to ensure the balance and generalization ability of model training and verification.

[0088] The labeling content includes but is not limited to the following three types: 1) basic attribute type: gender, age range; 2) emotional expression type: expression category (happy, sad, angry, surprised, scared, disgusted, neutral, confused, etc.), expression intensity score (a continuous value between 0 and 1, used to express the intensity of emotion); 3) geometric structure type: face contour type (square / ellipse / inverted triangle / long face, etc.), basic shape of five organs (thickness, width, height, angle, etc.) such as eyebrows, eyes, nose, and mouth, hair features (mustache density, mustache color, mustache contour, etc., including sideburns, eight-character mustache, and goat mustache).

[0089] S102 Data preprocessing: To improve the stability and accuracy of the reconstruction model, a series of data cleaning and enhancement operations need to be performed on the image samples collected in S101, including but not limited to: image uniform size scaling and standardization processing; face region cropping and boundary detection; illumination condition adjustment and contrast enhancement; Gaussian filtering, edge smoothing, and other noise reduction strategies.

[0090] S103 Discrete attribute encoding method based on semantic query: The present application provides a semantic attribute extraction method for three-dimensional face reconstruction, which is characterized by decomposing the information of the input face image into a set of interpretable, discrete semantic feature vectors constrained by natural language, to enhance the semantic distinction degree in the subsequent geometric and texture generation process, and to provide structured semantic guidance for subsequent three-dimensional modeling and texture enhancement, thereby improving the robustness and expression integrity in complex collection environments.

[0091] In the present embodiment, the input image is denoted as For this image, a set of face semantic related queries is first constructed where each query corresponds to a specific attribute dimension of the face, including but not limited to age, gender, face contour, eye structure, nose type, lip shape, skin texture, current facial expression, and current environmental conditions, and other multi-level face description elements.

[0092] Next, the image and query pair are jointly processed by calling a pre-trained image-text question answering model (e.g., a BLIP model) to generate a corresponding set of natural language responses wherein After obtaining the above response content, a language template set in advance is used to organize the answer set into a coherent facial natural language description which structurally contains multiple explicit facial attribute information.

[0093] Subsequently, the text prompt is input into a multi-modal model (such as CLIP) to obtain its global semantic embedding vector wherein represents the dimension of the text encoding space. Considering that there is a high degree of coupling between attributes, directly using this embedding representation is not conducive to fine-grained facial attribute modeling. Therefore, the present application further designs an attribute structure decoding module for mapping the global embedding vector into a structured attribute matrix wherein represents the number of attribute categories, represents the number of possible discrete states for each attribute category.

[0094] Specifically, the attribute structure decoding module is composed of a set of multi-layer feedforward neural networks (MLP), and its mapping process can be represented as:

[0095] ;

[0096] wherein the th row represents the state distribution of the th

[0097] attribute. To enhance discriminability, the output adopts a one-hot encoding structure, and a supervision mechanism is introduced to optimize prediction accuracy.

[0098] ;

[0099] wherein represents the cross-entropy loss, which is used to measure the difference between the predicted state distribution and the true label.

[0100] S104 Facial three-dimensional reconstruction method based on text condition: the purpose is to introduce a semantic regulation mechanism on the basis of a traditional three-dimensional shape regression framework to enhance the model's ability to model details in complex environments or different expression states and improve the accuracy of geometric and expression features.

[0101] The present application adopts a three-dimensional shape model (e.g. FLAME) with wide applicability to parameterize the representation of facial geometry, in which the shape control variables include: individual feature parameters , dynamic expression parameters and head pose parameters . The above parameters are uniformly converted into a canonical three-dimensional mesh representation via a three-dimensional shape model, with topological consistency and reconstruction stability. The specific process is as follows:

[0102] First, the input image is predicted by a multi-channel feature encoding network to predict the input parameters (individual feature parameters , dynamic expression parameters and head pose parameters ) of the three-dimensional shape model.

[0103] The multi-channel feature encoding network is composed of three sub-modules, each responsible for extracting identity structure, expression state and pose change information. Each sub-module is based on a lightweight network architecture, such as MobileNetV3 as the backbone network, to ensure high inference efficiency while maintaining expression ability. After the backbone network receives the input image, it extracts the deep semantic features of the image layer by layer through multi-layer convolution, nonlinear activation, batch normalization, etc. Finally, a high-dimensional feature map is output. The high-dimensional feature map output by the backbone network is compressed into a fixed-length one-dimensional feature vector through global average pooling operation. The feature vector obtained above is sent into multiple parameter prediction sub-networks (i.e. multiple fully connected layers), each of which is specifically used to regress a parameter type (including: individual feature parameters, dynamic expression parameters and head pose parameters).

[0104] After obtaining the preliminary three-dimensional modeling parameters, the present application further introduces text attribute information generated from the image as a prior semantic guidance signal to improve the personalization and semantic consistency of the three-dimensional reconstruction result. The text description information is first converted into an attribute matrix by an embedding network, which is generated by step S103. Then it is sent into two independently designed nonlinear conversion modules and , respectively, to output semantic modulation vectors and related to geometric shape modeling and expression modeling, respectively. The process can be expressed as:

[0105] m s = σ(Γ s ( P attr )), me = σ(Γ e ( P attr ));

[0106] wherein represents a Sigmoid activation function, used to enhance the non-linear ability of expression and limit the output value to the stable interval [0, 1].

[0107] The nonlinear conversion module used in the present application is a semantic mapping network composed of a multilayer perceptron (MLP), and the typical structure includes two to three fully connected layers, nonlinear activation, normalization and Dropout, which is used to convert the text semantic vector into a numerical stable and semantically effective modulation vector, and then guide the three-dimensional face parameter generation process.

[0108] Subsequently, the shape parameters output by the multi-channel feature encoding network are element-wise multiplied with the semantic modulation vector to obtain the fused individual parameters ; similarly, the expression parameters are fused with the semantic modulation vector to generate the modulated dynamic parameters . That is:

[0109] , ;

[0110] The above fused parameters are decoded by the FLAME three-dimensional model to finally generate a face three-dimensional mesh model with semantic consistency .

[0111] S105 Text-driven texture map generation method: the purpose is to introduce semantic features into the modeling process of StyleGAN2 in the image generation field to enhance the detail restoration of texture images, thereby enhancing semantic consistency and expression fidelity at the texture modeling level.

[0112] Unlike the traditional StyleGAN2 model which only uses random latent variables subject to standard normal distribution as input, the present application introduces text conditions to regulate the generation process. Specifically, first, semantic text description information is extracted based on the input image, and the semantic information is converted into an attribute matrix by the S103 pre-defined attribute structure decoding module. At the same time, the image itself extracts an initial latent vector The potential representation is not dependent on random sampling, but is generated adaptively according to image content, with stronger expression pertinence. The encoding network refers to a convolutional neural network (such as ResNet) high-dimensional vector for mapping the input image to a potential vector, which is mapped to a fixed-dimensional (such as 512-dimensional) potential vector through a fully connected layer (FC).

[0113] Then, the above attribute matrix is fused with the image encoding vector to form a joint feature expression u = f(z, P attr ), wherein represents a nonlinear fusion mapping function. The joint feature is then mapped to the style space to generate a style vector . The style vector is injected into multiple generation modules according to the hierarchical structure in the StyleGAN2 generator, to adjust the style expression of each layer, so as to realize the progressive enhancement of texture resolution while maintaining semantic consistency.

[0114] Finally, the texture image output by the generator is , which has high visual fidelity and texture structure highly consistent with the original image semantic features.

[0115] The reconstructed three-dimensional face mesh is combined with the texture image I UV to generate a two-dimensional image through a differentiable renderer. That is, the generated texture image is accurately attached to the surface of the three-dimensional face mesh through UV coordinate mapping, forming a complete three-dimensional face model with realistic appearance. In other words, the texture image is accurately attached to the surface of the three-dimensional mesh through the UV coordinate system, thereby restoring a three-dimensional face appearance with high visual fidelity. Then, the three-dimensional mesh with attached texture is rendered using a differentiable renderer to generate the final two-dimensional image.

[0116] The UV coordinate system refers to a set of two-dimensional coordinates (U, V) defined on each vertex of a three-dimensional model, used to represent the position of the vertex on a two-dimensional texture image. U and V correspond to the horizontal and vertical positions of the texture image respectively, and the value range is usually 0 to 1. Through the UV coordinate system, the two-dimensional texture image can be accurately mapped to the surface of the three-dimensional model, realizing the texture rendering of the model.

[0117] S106 Semantic alignment: Feature extraction of the rendered image (i.e., the two-dimensional image generated by the differentiable renderer) and the generated text T using a text-image encoding model (e.g., CLIP), combined with a cosine similarity-based bidirectional alignment loss and a multi-view semantic enhancement strategy to enhance the cross-modal semantic consistency between the image and the text, and to address the weak semantic association and the difficulty in achieving collaborative optimization between geometric shapes, expression parameters, and texture features.

[0118] For the rendered two-dimensional image, first generate K different block-level random mask views, then input the K mask views into the image encoder of the text-image encoding model (e.g., CLIP) to obtain visual representations . Finally, aggregate to obtain cross-view Figure 1 consistency visual representations. Then the overall visual representation after aggregation is represented as:

[0119] .

[0120] Block-level random mask views, i.e., generate K two-dimensional images from different perspectives (e.g., regions, occlusion conditions), and apply a set of regional masks (masks) on each image, thereby forming a local image sample set with semantic region attention, to enhance the robustness and generalization ability of text-image semantic alignment.

[0121] To reduce the error caused by the modal difference between image encoding and text encoding, the present invention designs a bidirectional cross-contrast loss function to align the representations between visual and text modalities. Let τ be the temperature factor and B be the batch sample size, then the loss function is defined as follows:

[0122] ;

[0123] The first log term is the contrast between the image and all texts : the numerator is the similarity of and corresponding positive samples ; the denominator is the sum of the similarity of and all (including and other negative samples); the second log term is the contrast between the text and all images , which is symmetrical to the first log term.

[0124] wherein, is the overall visual representation after encoding by the image encoder of the text-image encoding model; b is the text representation after encoding by the text encoder of the text-image encoding model; , For cosine similarity, the similarity scores of image to text and text to image are calculated respectively, which are used as the optimization reference in the training process. Specifically, it is expressed as:

[0125] , .

[0126] In addition, in order to ensure the reconstruction consistency in visual perception between the rendered image and the original image, the present application also introduces an image reconstruction loss, including but not limited to the following categories: pixel error loss: measures the difference between the generated image and the original image in the pixel space; landmark position error: based on the Euclidean distance constraint between two-dimensional and three-dimensional facial key points; regularization loss: imposes prior constraints on model parameters or representation space; emotion consistency error: use a pre-trained expression recognition model to extract emotion features, and minimize the semantic difference with the target image. The above loss functions can be jointly constructed into an image-level (i.e. the loss between the rendered image and the original image) reconstruction loss term, denoted as:

[0127] ;

[0128] where each loss term corresponds to the pixel, landmark, regularization, and emotion consistency supervision term, is a hyperparameter weight.

[0129] S107 Expression classification: After completing the three-dimensional face reconstruction, based on the reconstructed face three-dimensional grid automatically judge the facial emotion category, especially suitable for mental disease auxiliary diagnosis and remote psychological intervention and other application scenarios.

[0130] The expression classifier is preferably a multi-layer perceptron (MLP), including at least three layers of non-linear fully connected network, supporting high-order non-linear mapping and classification decision of input features. The output is the probability distribution corresponding to each expression category (such as happy, sad, angry, and scared). The output form supports single-label classification, and can also be extended to regression prediction of continuous emotion dimension, such as valence (Valence) and arousal (Arousal) scores, to adapt to different emotion analysis needs.

[0131] Further, to adapt to the judgment needs of weak and atypical expressions in mental disease scenarios, the classifier introduces pathological sample dataset for fine-tuning in the training stage, to improve the recognition accuracy of slight emotional changes, suitable for remote psychological intervention, emotion disorder monitoring and other intelligent medical tasks.

[0132] On the disclosed Now dataset, the reconstruction accuracy of the application is tested by three error indicators of median error (Median), mean error (Mean) and standard deviation (Std), and low error performances of 1.38mm, 1.63mm and 1.42mm are achieved, which indicates that it is significantly superior to the existing method in geometric accuracy. In the emotion recognition task, the AffectNet dataset is used to evaluate the classification accuracy, and the model of the application achieves an expression classification accuracy (E-ACC) of 0.75. Further, in the continuous emotion dimension regression index, the valence prediction achieves a harmony correlation coefficient (Valence Concordance Correlation Coefficient, V-CCC) of 0.79 and a root mean square error (Valence Root Mean Square Error, V-RMSE) of 0.29; the arousal prediction achieves a harmony correlation coefficient (Arousal Concordance Correlation Coefficient, A-CCC) of 0.71 and a root mean square error (Arousal Root Mean Square Error, A-RMSE) of 0.30, which fully verifies the effectiveness and robustness of the application in the fine-grained emotion modeling and recognition task.

[0133] As shown in Figure 2 The application also provides a text-guided three-dimensional face reconstruction and emotion analysis device, which comprises:

[0134] An image acquisition unit is configured to acquire facial image information of a user;

[0135] A face reconstruction unit stores a three-dimensional shape model trained by the above method, and the three-dimensional shape model obtains a three-dimensional face mesh from the facial image information acquired by the image acquisition unit .

[0136] An emotion analysis unit stores an emotion analysis model trained by the above method, and judges a facial emotion category according to the three-dimensional face mesh .

[0137] A data storage unit is configured to store the three-dimensional face reconstruction and emotion classification results.

[0138] A human-computer interaction visualization unit is configured to output the three-dimensional face reconstruction and emotion classification results.

[0139] A communication and control unit is configured to transmit and control information between the processing unit and the human-computer interaction visualization unit.

[0140] The above embodiments describe the technical solutions and advantages of the present application in detail. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the present application. Any modification, supplement, and equivalent replacement within the principle range of the present application should be included in the protection scope of the present application.

Claims

1. A text-guided 3D face reconstruction and emotion analysis method, characterized in that, Includes the following steps: (1) Obtain facial image information of the human body and perform preprocessing and annotation to construct a training set; (2) The 3D face reconstruction model is trained using a training set, including: (2-i) Deconstruct the facial images in the training set into text descriptions T containing multiple facial attribute information, and transform the text descriptions T into structured attribute matrices. P attr ; (2-ii) Input the facial images from the training set into the 3D shape model to perform 3D face reconstruction, and simultaneously introduce the attribute matrix into the 3D face reconstruction process to obtain a semantically consistent 3D face mesh. M face ; (2-iii) Input the facial images from the training set into the image style generator to synthesize texture images, and simultaneously introduce the attribute matrix into the texture image synthesis process to obtain texture images with semantic consistency. I UV ; (2-iv) 3D mesh of face M face With texture image I UV Combined, the rendered image is generated. (2-v) Construct loss functions between the rendered image and the original facial image, and between the rendered image and the text description, and train them, including: (2-va) For the rendered image, generate K different block-level random mask views, and input the K mask views into the graph coding model to obtain visual representations. Finally, the feature vectors are aggregated to obtain the image feature vectors. : ; (2-vb) Input the text description T into the graph-text coding model to obtain the text feature vector. b ; (2-vc) Construct a bidirectional cross-contrast loss function between images and text. l align : ; τ is the temperature factor; B is the sample size; Image-to-text similarity score; A similarity score is given for text to image. (2-vd) Constructing the image reconstruction loss function l rec : ; λ 1. λ 2. λ 3. λ 4 represents the hyperparameter weights; l pix , l land , l reg , l emo These correspond to pixel error loss, marker position error loss, regularization loss, and sentiment consistency error loss, respectively. After training, a trained 3D shape model is obtained. (3) Input the face image to be analyzed into the trained 3D shape model to obtain the 3D face mesh. M face 3D mesh of face M face The emotion analysis model is used to determine the emotion category of a person's face.

2. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that, In step (1), the annotation content includes: basic attribute category, including gender and age; emotion expression category, including expression category and expression intensity score; geometric structure category, including facial contour type, basic shape of facial features, and hair features.

3. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that, Preprocessing includes data cleaning and data augmentation; Data augmentation includes: uniform image scaling and normalization; facial region cropping and boundary detection; lighting condition adjustment and contrast enhancement; and noise reduction.

4. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that, Step (2-i) includes: (2-ia) Construct a set of semantically related query sets for faces Q={q1, q2,…, q M }, where each query q i A specific attribute dimension corresponding to a human face; (2-ib) For the input facial image I, a pre-trained image-text question answering model is invoked to jointly process the image and query pair, generating the corresponding natural language response set A={a1,a2,…,a…} M }, where a i =BLIP(I,q i ); (2-ic) Utilizing a priori language templates T tem , will set The organization is a text description T containing multiple facial attribute information; (2-id) Obtain the global semantic embedding vector z of the text description T through a multimodal model. T , the global embedding vector z T Mapped to a structured attribute matrix P attr .

5. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that, Step (2-ii) includes: (2-iia) Input the facial image into the multi-channel feature encoding network Obtain individual feature parameters Dynamic expression parameters and head pose parameters ; (2-iib) will convert the attribute matrix P attr Input to nonlinear conversion module respectively and Semantic modulation vectors related to geometric shape modeling were obtained respectively. m s and semantic modulation vectors related to facial expression modeling m e : m s = σ(Γ s ( P attr )), m e = σ(Γ e ( P attr )); in This represents the Sigmoid activation function; (2-iic) represents individual characteristic parameters With semantic modulation vector Perform element-wise multiplication to obtain the modulated individual parameters. ; Set the facial expression parameters With semantic modulation vector The components are fused together to generate modulated dynamic parameters. : , ; (2-iid) will , , After decoding the 3D shape model, a semantically consistent 3D face mesh is obtained. .

6. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that, Step (2-iii) includes: (2-iiia) Input the facial image into the encoding network to extract the initial latent vector z; (2-iiib) Attribute matrix P attr The joint feature representation u = f(z, is fused with the initial latent vector z to form a joint feature representation u = f(z, P attr ),in Represents a nonlinear fusion mapping function; (2-iiic) uses the joint feature representation u as input to the StyleGAN2 model and outputs a texture image with semantic consistency. I UV .

7. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that, In step (3), the emotion analysis model is a multilayer perceptron.

8. A text-guided 3D face reconstruction and emotion analysis device, characterized in that, include: The data acquisition module acquires facial image information of the human body, performs preprocessing and annotation, and constructs a training set; The model training module uses a training set to train the 3D shape model, including: The facial images in the training set are deconstructed into text descriptions containing multiple facial attribute information, and the text descriptions are transformed into structured attribute matrices. Facial images from the training set are input into a 3D shape model for 3D face reconstruction. Simultaneously, the aforementioned attribute matrix is ​​incorporated into the 3D face reconstruction process to obtain a semantically consistent 3D facial mesh. ; Facial images from the training set are input into an image style generator to synthesize texture images. Simultaneously, the attribute matrix is ​​incorporated into the texture image synthesis process to obtain texture images with semantic consistency. I UV ; 3D mesh of the face With texture image I UV Combined, the rendering process generates a rendered image; Loss functions are constructed and trained for the relationships between the rendered image and the original facial image, and between the rendered image and the text description, including: For rendering an image, K different block-level random mask views are generated, and these K mask views are input into a graph encoding model to obtain a visual representation. Finally, the feature vectors are aggregated to obtain the image feature vectors. : ; The text description T is input into the image encoding model to obtain the text feature vector. b ; Construct a bidirectional cross-contrast loss function between images and text. l align : ; τ is the temperature factor; B is the sample size; Image-to-text similarity score; A similarity score is given for text to image. Constructing an image reconstruction loss function l rec : ; λ 1. λ 2. λ 3. λ 4 represents the hyperparameter weights; l pix , l land , l reg , l emo These correspond to pixel error loss, marker position error loss, regularization loss, and sentiment consistency error loss, respectively. After training, a trained 3D shape model is obtained. The 3D face reconstruction and emotion analysis module stores pre-trained 3D shape models and emotion analysis models. The 3D shape model generates a 3D face mesh based on the face image to be analyzed. The sentiment analysis model is based on a 3D facial mesh. Determine the emotion category of a person's face.

9. A text-guided 3D face reconstruction and emotion analysis device, characterized in that, include: The image acquisition unit is used to acquire facial image information of the user; The face reconstruction unit stores a three-dimensional shape model trained using the method described in any one of claims 1-7. The three-dimensional shape model obtains a three-dimensional face mesh based on facial image information acquired by the image acquisition unit. ; The emotion analysis unit stores an emotion analysis model based on a 3D facial mesh. Determine the emotion category of a person's face; Data storage unit, used to store the results of 3D facial reconstruction and emotion classification; The human-computer interaction visualization unit is used to output the results of 3D facial reconstruction and emotion classification; The communication and control unit is used for information transmission and control between the processing unit and the human-machine interaction visualization unit.

Citation Information

Patent Citations

  • AI digital human automatic expression generation system based on voice driving

    CN118800274A

  • Multi-mode multi-factor depression identification system fusing emotion information

    CN119786021A

  • High-fidelity three-dimensional face model generation method based on natural text description

    CN115984485A

  • Three-dimensional face reconstruction method based on CLIP model

    CN116563457A