Three-dimensional face reconstruction and emotion analysis method, device and equipment based on text guidance
Through a text-guided 3D face reconstruction method combined with natural language and image features, the accuracy and privacy issues of 3D face reconstruction under complex conditions are solved, and high-precision facial feature reconstruction and emotion recognition are achieved, which is suitable for mental illness diagnosis and telemedicine.
Patent Information
- Application Number
- CN202511156807.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing three-dimensional face reconstruction technology has the risk of privacy leakage in the diagnosis of mental illness, and the reconstruction results are easily distorted under complex acquisition conditions, making it difficult to achieve high-precision facial feature fidelity and semantic consistency.
A text-guided 3D face reconstruction method is adopted. By jointly encoding the attribute matrix generated by natural language and image features, combining it with a multimodal model for 3D reconstruction, introducing a semantic consistency alignment module and forward and reverse matching, the geometric reconstruction accuracy and emotion recognition accuracy are enhanced.
It significantly improves the geometric reconstruction accuracy and emotion recognition accuracy in complex scenarios, reduces the risk of privacy leakage, and is suitable for remote psychological intervention and intelligent medical tasks.
Smart Images

Figure CN120747417A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a text-guided three-dimensional face reconstruction and emotion analysis method, device and equipment. Background Art
[0002] With the deep integration of artificial intelligence and medical imaging analysis, 3D facial reconstruction technology is increasingly showing broad potential for application in the medical field, particularly in the auxiliary diagnosis of psychiatric and psychological disorders. Depression, a common mental illness, has long been a focus of academic and clinical attention for its non-invasive, objective identification and diagnosis. Abnormal changes in facial expressions and facial muscle activity are often considered a key sign of depression.
[0003] Chinese patent publication CN119786021A discloses a multimodal, multifactorial depression recognition system that integrates emotional information. However, in sensitive scenarios such as psychiatric assessment, directly collecting and transmitting raw image, video, or voice data poses a potential privacy risk.
[0004] In real-world medical environments, complex acquisition conditions such as image blur, extreme lighting, and partial occlusion often exist. Furthermore, patients may exhibit unusual facial expressions or subtle facial dynamics (e.g., micro-expressions). Existing methods primarily rely on low-dimensional parametric models (e.g., 3DMM). These methods are limited by their limited expressiveness and difficulty capturing detailed facial features. They struggle to effectively represent complex facial deformations, and the reconstruction results are prone to distortion of key structural details, affecting geometric accuracy and fidelity of facial features. This significantly lags behind the high-precision facial features required for medical diagnosis.
[0005] Some existing multimodal reconstruction schemes attempt to introduce audio or other auxiliary modalities to improve modeling accuracy. For example, Chinese patent document CN118800274A discloses a voice-driven AI digital human automatic expression generation system, which includes a voice expression generation module, a facial expression database, an expression feature extraction module, and a facial three-dimensional reconstruction module. However, during the modal alignment process, there are often problems such as inconsistent embedding space and weak feature matching. For example, the semantic information of speech is hidden in the audio signal, and speech expression is more vague and implicit, making it difficult to accurately describe facial details. The semantic association with visual information is weak, making it difficult to achieve coordinated optimization between geometric shapes, expression parameters, and texture features, thereby affecting the semantic consistency and expression integrity of the overall reconstruction effect. It is impossible to achieve efficient fusion of multimodal semantic information, limiting its application and promotion in delicate tasks such as facial recognition and emotion understanding.
[0006] Therefore, there is a need for a three-dimensional face reconstruction system for auxiliary diagnosis of mental illnesses that can effectively integrate multimodal information (text and vision), improve the ability of facial semantic expression, and enhance the modeling accuracy of complex emotional states. Summary of the Invention
[0007] The present invention provides a text-guided 3D face reconstruction and emotion analysis method, device and equipment to improve the accuracy of 3D face reconstruction under conditions of limited image information, and achieve the goal of assisting in the diagnosis of mental illness while protecting patient privacy.
[0008] The technical solutions of the present invention are as follows: A text-guided 3D face reconstruction and emotion analysis method comprises the following steps: (1) Obtain facial image information of the human body, perform preprocessing and annotation, and construct a training set; (2) Using the training set to train the 3D face reconstruction model, including: (2-i) Deconstruct the facial images in the training set into text descriptions T containing multiple facial attribute information, and convert the text descriptions T into a structured attribute matrix P attr ; (2-ii) Input the facial images in the training set into the 3D shape model to perform 3D face reconstruction, and at the same time introduce the attribute matrix into the 3D face reconstruction process to obtain a 3D face mesh with semantic consistency ; (2-iii) Input the facial images in the training set into the image style generator to synthesize the texture image, and at the same time introduce the attribute matrix into the texture image synthesis process to obtain a texture image with semantic consistency I UV ; (2-iv) 3D face mesh With texture image I UV Combining, rendering and generating a rendered image; (2-v) Construct loss functions between rendered images and original facial images, and between rendered images and text descriptions, and perform training. (3) Input the face image to be analyzed into the trained 3D shape model to obtain the 3D face mesh , the 3D face mesh Input the emotion analysis model to judge the facial emotion category.
[0009] Reconstructing the three-dimensional structure of a face requires not only accurately restoring the overall contours of the face, but also capturing the dynamic changes in facial muscles at a fine-grained level to support high-precision tasks such as emotion recognition and auxiliary diagnosis of mental illness. This invention significantly enhances the semantic alignment between facial structure and emotional attributes in three-dimensional face modeling by constructing a collaborative modeling mechanism that integrates text descriptions and image visual features. Even when image quality is limited or posture changes significantly, stable, high-fidelity three-dimensional reconstruction results can still be achieved. This technology is particularly suitable for application scenarios such as auxiliary diagnosis of mental illnesses and remote psychological intervention.
[0010] In step (1), the annotation content includes basic attribute class, emotion expression class and geometric structure class; basic attribute class includes gender, age, etc.; emotion expression class includes expression category (happy, sad, angry, surprised, fear, disgust, neutral, confused, etc.), expression intensity score (continuous value between 0 and 1, used to express the intensity of emotion); geometric structure class includes facial contour type (square / ellipse / inverted triangle / long face, etc.), basic form of facial features (thickness, width, height, angle, etc.), hair features (beard density, beard color, beard outline, etc., beard outline includes sideburns, mustache, goatee, etc.), etc.
[0011] In step (1), preprocessing includes data cleaning and enhancement operations, including but not limited to: image uniform size scaling and standardization; facial region cropping and boundary detection; lighting condition adjustment and contrast enhancement; Gaussian filtering, edge smoothing and other noise reduction strategies.
[0012] Step (2-i) includes: (2-ia) Construct a set of face semantic related query sets Q={q1, q2,…, q M}, where each query q i Corresponding to a specific attribute dimension of the face; (2-ib) For the input facial image I, call the pre-trained image-text question answering model, jointly process the image and query pair, and generate the corresponding natural language response set A={a1,a2,…,a M}, where a i =BLIP(I,q i ); (2-ic) Using a priori language templates T tem , will be collected Organized into a text description T containing multiple facial attribute information; (2-id) Obtain the global semantic embedding vector z of the text description T through the multimodal model T , embed the global vector z T Mapping to a structured attribute matrix Pattr .
[0013] In step (2-id), the multimodal model may be a CLIP (Contrastive Language-Image Pre-training) model.
[0014] In step (2-id), the global embedding vector z is transformed through a set of multi-layer feedforward neural networks (MLPs) T Mapping to a structured attribute matrix P attr : P attr = F MLP (z T ); , K Indicates the number of attribute types, C Indicates the number of possible discrete states for each type of attribute; OK Indicates the k The state distribution of an attribute.
[0015] Step (2-ii) includes: (2-iia) Input facial images into the multi-channel feature encoding network , obtain individual characteristic parameters , dynamic expression parameters and head pose parameters ; (2-iib) The attribute matrix P attr Input nonlinear transformation module and , respectively obtain the semantic modulation vectors related to geometric shape modeling m s and semantic modulation vectors related to expression modeling m e : m s = σ(Γ s ( P attr )), m e = σ(Γ e ( P attr )); in Represents the Sigmoid activation function; (2-iic) Individual characteristic parameters and semantic modulation vector Perform element-by-element multiplication to obtain the modulated individual parameters ; Set the expression parameters and semantic modulation vector Fusion generates modulated dynamic parameters : , ; (2-iic) will 、 、 After decoding the 3D shape model, a 3D face mesh with semantic consistency is obtained. .
[0016] Step (2-iii) includes: (2-iiia) Input the facial image into the encoding network to extract the initial latent vector z; (2-iiib) The attribute matrix P attr Fuse with the initial latent vector z to form a joint feature expression u = f(z, P attr ),in represents the nonlinear fusion mapping function; (2-iiic) Use the joint feature expression u as the input of the StyleGAN2 model and output a texture image with semantic consistency I UV .
[0017] Step (2-v) includes: (2-va) For the rendered image, generate K different block-level random mask views and input the K mask views into the image-text encoding model to obtain visual representation , and finally aggregate to obtain the image feature vector : ; (2-vb) Input the text description T into the image-text encoding model to obtain the text feature vector b ; (2-vc) Construct a bidirectional cross-contrast loss function between images and text l align : ; τ is the temperature factor; B is the number of samples; , is the image-to-text similarity score; , is the similarity score from text to image; (2-vd) Constructing image reconstruction loss functionl rec : ; λ 1. λ 2. λ 3. λ 4 is the hyperparameter weight; l pix 、 l land 、 l reg 、 l emo They correspond to pixel error loss, landmark position error loss, regularization loss, and sentiment consistency error loss respectively; After the training is completed, the trained 3D shape model is obtained.
[0018] The pixel error loss measures the difference between the rendered image and the original facial image in pixel space; the landmark position error is based on the Euclidean distance constraint between two-dimensional and three-dimensional facial key points; the regularization loss imposes prior restrictions on model parameters or representation space; the emotion consistency error uses a pre-trained expression recognition model to extract emotional features and minimize the semantic difference with the target image.
[0019] In step (3), the sentiment analysis model is a multi-layer perceptron (MLP).
[0020] In order to adapt to the needs of distinguishing weak and atypical expressions in mental illness scenarios, it is preferred to introduce pathological sample datasets in the training stage to fine-tune the emotion analysis model to improve its recognition accuracy of slight emotional changes, which is suitable for intelligent medical tasks such as remote psychological intervention and emotional disorder monitoring.
[0021] The present invention also provides a text-guided 3D face reconstruction and emotion analysis device, comprising: The data acquisition module obtains facial image information of the human body and performs preprocessing and annotation to build a training set; The model training module uses the training set to train the 3D shape model, including: Deconstruct facial images in the training set into text descriptions containing multiple facial attribute information, and convert the text descriptions into a structured attribute matrix; The facial images in the training set are input into the 3D shape model for 3D face reconstruction. At the same time, the attribute matrix is introduced into the 3D face reconstruction process to obtain a 3D face mesh with semantic consistency. ; The facial images in the training set are input into the image style generator to synthesize the texture image, and the attribute matrix is introduced into the texture image synthesis process to obtain a texture image with semantic consistency.I UV ; 3D face mesh With texture image I UV Combining, rendering to generate a rendered image; Construct loss functions between rendered images and original facial images, and between rendered images and text descriptions, and perform training. The 3D face reconstruction and emotion analysis module stores the trained 3D shape model and emotion analysis model. The 3D shape model generates a 3D face mesh based on the face image to be analyzed. , the emotion analysis model is based on the 3D mesh of the face Determine the facial emotion category.
[0022] The present invention also provides a text-guided 3D face reconstruction and emotion analysis device, comprising: An image acquisition unit, used to acquire user facial image information; The face reconstruction unit stores the 3D shape model trained by the above method, and the 3D shape model obtains the 3D mesh of the face based on the facial image information collected by the image acquisition unit. ; Emotion analysis unit, which stores the trained emotion analysis model and calculates the emotion analysis model based on the three-dimensional grid of the face. Determine facial emotion categories; A data storage unit for storing the results of three-dimensional face reconstruction and emotion classification; Human-computer interaction visualization unit, used to output 3D facial reconstruction and emotion classification results; The communication and control unit is used for information transmission and control between the processing unit and the human-computer interaction visualization unit.
[0023] The text-guided 3D face reconstruction and emotion analysis device of the present invention is suitable for intelligent medical tasks such as remote psychological intervention and emotional disorder monitoring.
[0024] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention jointly encodes the attribute matrix generated by natural language and the image features and inputs them into the 3D reconstruction framework, which can significantly improve the geometric reconstruction accuracy in complex scenes such as face blur, partial occlusion, or drastic / subtle changes in facial expressions, thereby effectively alleviating the problem of reconstruction failure or feature loss in traditional methods under non-ideal image conditions.
[0025] (2) The present invention performs fine-grained semantic modeling of the facial region under a multi-view mask mechanism, and enhances the coupling strength between the image and the text through semantic-driven forward and backward matching, making the correspondence between facial expression changes and geometric forms clearer, thereby improving the accuracy and interpretability of emotion recognition.
[0026] (3) The three-dimensional parametric reconstruction output used in this invention can be used as a structured medical feature expression, replacing the traditional transmission of patient facial information in the form of images or videos, significantly reducing the risk of leakage of the original private image. The generated three-dimensional reconstruction results not only preserve the facial structure and expression dynamics, but can also be used for clinical auxiliary analysis and remote medical diagnosis, taking into account both data security and application practicality, and achieving effective synergy between data privacy protection and functional processing in medical scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flowchart of the face reconstruction and emotion analysis method; Figure 2 This is a schematic block diagram of the structure of the face reconstruction and emotion analysis equipment. DETAILED DESCRIPTION
[0028] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not have any limiting effect on the present invention.
[0029] Compared to existing 3D face modeling approaches that rely on a single visual modality, this paper provides a multimodal reconstruction method that integrates textual semantic information. By jointly encoding attribute semantic vectors generated from natural language with image features and inputting them into the 3D reconstruction framework, it significantly improves geometric reconstruction accuracy in complex scenarios such as face blur, partial occlusion, and dramatic or subtle changes in facial expression. This effectively alleviates the problems of reconstruction failure and feature loss that plague traditional methods in non-ideal image conditions.
[0030] Furthermore, the semantic consistency alignment module proposed in the present invention performs fine-grained semantic modeling of the facial area under a multi-view mask mechanism, enhances the coupling strength between image and language through semantic-driven forward and reverse matching, and makes the correspondence between facial expression changes and geometric shapes clearer, thereby improving the accuracy and interpretability of emotion recognition.
[0031] Furthermore, the 3D parametric reconstruction output employed in this invention can be used as a structured medical feature representation, replacing the traditional transmission of patient facial information in image or video form, significantly reducing the risk of leaking the original private image. The resulting 3D reconstruction not only preserves facial structure and expression dynamics but can also be used for clinical analysis and remote medical diagnosis, balancing data security with practical application, achieving effective synergy between data privacy protection and functional processing in medical scenarios.
[0032] A text-guided 3D face reconstruction and emotion analysis method, the steps of which are as follows: S100 System Construction: First, a composite computing system for 3D facial reconstruction and diagnostic analysis was constructed. This system includes: an image acquisition unit based on an imaging sensor, preferably an RGB digital camera, for capturing visual facial images; a model processing module for executing the 3D reconstruction algorithm and subsequent emotion analysis tasks; and a human-computer interaction visualization unit for displaying the results, outputting the reconstruction model and analysis results.
[0033] S101 Data Collection: This invention acquires dynamic image information of the user's face through the imaging module described in S100. To control memory resource usage and avoid temporal redundancy of image frames, the present invention adopts an interval frame extraction strategy, regularly extracting key frames for subsequent analysis, thereby ensuring facial feature diversity while reducing processing costs.
[0034] After the acquisition is completed, preliminary cleaning is performed, including but not limited to: removing images that are severely blurred / overexposed / occluded, and removing images that do not contain complete facial structures (such as only half a face).
[0035] The acquired facial image samples are formed into a structured data set through manual and automatic annotation processes, and the samples are divided according to a set ratio, preferably 80% as a training data set and 20% as an evaluation test set, to ensure the balance and generalization ability of model training and verification.
[0036] The annotation content includes but is not limited to the following three categories: 1) Basic attributes: gender and age group; 2) Emotional expression: expression category (happy, sad, angry, surprised, fearful, disgusted, neutral, confused, etc.), expression intensity score (a continuous value between 0 and 1, used to express the intensity of the emotion); 3) Geometric structure: facial contour type (square / oval / inverted triangle / long face, etc.), basic shapes of facial features such as eyebrows, eyes, nose, and mouth (thickness, width, height, angle, etc.), hair characteristics (beard density, beard color, beard outline, etc., beard outlines include sideburns, mustaches, goatees, etc.).
[0037] S102 Data Preprocessing: To improve the stability and accuracy of the reconstructed model, a series of data cleaning and enhancement operations are performed on the image samples collected by S101, including but not limited to: image uniform resizing and normalization; facial region cropping and boundary detection; lighting condition adjustment and contrast enhancement; Gaussian filtering, edge smoothing and other noise reduction strategies.
[0038] S103 Discrete attribute encoding method based on semantic query: The present invention provides a semantic attribute extraction method for three-dimensional face reconstruction. The core of the method is to deconstruct the information of the input facial image into a set of interpretable discrete semantic feature vectors constrained by natural language, so as to enhance the semantic distinction in the subsequent geometry and texture generation process, provide structured semantic guidance for subsequent three-dimensional modeling and texture enhancement, and thus improve the robustness and expression integrity in complex acquisition environments.
[0039] In this embodiment, the input image is denoted as For this image, we first construct a set of query sets related to facial semantics. , where each query Corresponding to a specific attribute dimension of the human face, it includes but is not limited to age, gender, facial contour, eye structure, nose shape, lip shape, skin texture, current facial expression, and current environmental conditions, among other multi-level facial description elements.
[0040] Next, by calling a pre-trained image-text question answering model (such as the BLIP model), the image and query pairs are jointly processed to generate the corresponding natural language response set. ,in After obtaining the above response content, use the language template set in advance , the answer set Organize into a coherent natural language description of the face , the description structuredly contains multiple explicit facial attribute information.
[0041] Then, the text prompt Input to a multimodal model (such as CLIP) to obtain its global semantic embedding vector ,in Represents the dimension of the text encoding space. Considering the high coupling between the semantics of attributes, directly using this embedding representation is not conducive to fine-grained face attribute modeling. Therefore, the present invention further designs an attribute structure decoding module to convert the global embedding vector Mapping to a structured attribute matrix ,in Indicates the number of attribute types, Represents the number of possible discrete states for each attribute.
[0042] Specifically, the attribute structure decoding module is composed of a set of multi-layer feedforward neural networks (MLPs), and its mapping process can be expressed as: ; Among them OK Indicates the To enhance the discriminability, the output adopts a one-hot encoding structure and introduces a supervision mechanism to optimize the prediction accuracy.
[0043] During the training phase, the present invention uses attribute label set As a supervisory signal, a loss function is used for optimization. The loss function is defined as follows: ; in Represents the cross entropy loss, which is used to measure the difference between the predicted state distribution and the true label.
[0044] S104 Text-based 3D face reconstruction method: The purpose is to introduce a semantic control mechanism based on the traditional 3D shape regression framework to enhance the model's detailed modeling in complex environments or different facial expressions, and improve the accuracy of geometric and expression features.
[0045] The present invention uses a widely applicable three-dimensional shape model (such as FLAME) to parameterize the facial geometry, where the shape control variables include: individual feature parameters , dynamic expression parameters and head pose parameters The above parameters are uniformly converted into a standardized 3D mesh representation through the 3D shape model, which has topological consistency and reconstruction stability. The specific process is as follows: First, the input image passes through a multi-channel feature encoding network Predict the input parameters of the 3D shape model (individual feature parameters , dynamic expression parameters and head pose parameters ).
[0046] Multi-channel feature encoding network It consists of three submodules, each responsible for extracting identity structure, expression state, and posture change information. Each submodule is built on a lightweight network architecture, such as the MobileNetV3 backbone network, to ensure high inference efficiency while maintaining expressiveness. After receiving the input image, the backbone network extracts the deep semantic features of the image layer by layer through multi-layer convolution, nonlinear activation, batch normalization, and other operations. The final output is a high-dimensional feature map. The high-dimensional feature map output by the backbone network is compressed into a fixed-length one-dimensional feature vector through a global average pooling operation. The resulting feature vectors are fed into multiple parameter prediction subnetworks (i.e., multiple fully connected layers), each of which is dedicated to regressing a parameter type (including individual feature parameters, dynamic expression parameters, and head posture parameters).
[0047] After obtaining the preliminary 3D modeling parameters, the present invention further introduces text attribute information generated from the image as a priori semantic guidance signal to enhance the personalization and semantic consistency of the 3D reconstruction results. The text description information is first converted into an attribute matrix through the embedding network. The attribute matrix is generated by step S103 and then fed into two independently designed nonlinear conversion modules. and , which output semantic modulation vectors related to geometric shape modeling and expression modeling respectively and The process can be expressed as: m s = σ(Γ s ( P attr )), m e = σ(Γ e ( P attr )); in Represents the Sigmoid activation function, which is used to enhance the nonlinear ability of expression and limit the output value to a stable range [0,1].
[0048] The nonlinear conversion module used in this invention is a semantic mapping network composed of a multi-layer perceptron (MLP). Its typical structure includes two to three fully connected layers, nonlinear activation, normalization and dropout. It is used to convert text semantic vectors into numerically stable and semantically valid modulation vectors, thereby guiding the three-dimensional face parameter generation process.
[0049] Subsequently, the shape parameters of the multi-channel feature encoding network output and semantic modulation vector Perform element-by-element multiplication to obtain the fused individual parameters ; Similarly, expression parameters and semantic modulation vector Fusion generates modulated dynamic parameters .Right now: , ; The above fusion parameters are then decoded by the FLAME 3D model to finally generate a 3D face mesh model with semantic consistency. .
[0050] S105 Text-driven texture map generation method: The purpose is to introduce semantic features into the modeling process of StyleGAN2 in the field of image generation to enhance the detail restoration of texture images, thereby enhancing semantic consistency and expression fidelity at the texture modeling level.
[0051] Unlike the traditional StyleGAN2 model which only uses random latent variables that obey the standard normal distribution as input, the present invention regulates the generation process by introducing text conditions. Specifically, semantic text description information is first extracted based on the input image, and the semantic information is converted into an attribute matrix by the attribute structure decoding module predefined in S103. At the same time, the image itself is passed through the encoding network to extract an initial latent vector This latent representation does not rely on random sampling, but is adaptively generated based on image content, making it more expressive and targeted. The encoding network refers to a convolutional neural network (such as ResNet) that maps the input image to a high-dimensional latent vector. This is then mapped to a latent vector of fixed dimension (such as 512 dimensions) through a fully connected layer (FC).
[0052] Next, the above attribute matrix With the image encoding vector Fusion is performed to form a joint feature expression u = f(z, P attr ),in Represents a nonlinear fusion mapping function. Joint features It is then mapped to the style space to generate a style vector This style vector is injected into multiple generation modules in the StyleGAN2 generator according to a hierarchical structure to adjust the style expression of each layer, thereby achieving a progressive enhancement of texture resolution while maintaining semantic consistency.
[0053] Finally, the texture image output by the generator is , which has high visual fidelity and a texture structure that is highly consistent with the semantic features of the original image.
[0054] The reconstructed 3D face mesh With texture imageI UV Combined with the texture map, a 2D image is generated using a differentiable renderer. The generated texture map is precisely attached to the surface of the 3D facial mesh using UV coordinate mapping, forming a complete and realistic 3D facial model. In other words, the texture map is accurately attached to the 3D mesh surface using the UV coordinate system, restoring the 3D facial appearance with high visual fidelity. The textured 3D mesh is then rendered using a differentiable renderer to generate the final 2D image.
[0055] The UV coordinate system is a set of two-dimensional coordinates (U, V) defined at each vertex of a 3D model, representing the vertex's position on a 2D texture image. U and V correspond to the horizontal and vertical positions of the texture image, respectively, and typically range from 0 to 1. Using UV coordinates, the system can accurately map a 2D texture image to the surface of a 3D model, enabling texture rendering of the model.
[0056] S106 Semantic Alignment: Utilizes image-text encoding models (e.g., CLIP) to extract features from rendered images (i.e., two-dimensional images generated using a differentiable renderer) and generated text T. Combined with a cosine similarity-based bidirectional alignment loss and a multi-view semantic enhancement strategy, this approach enhances cross-modal semantic consistency between images and text, addressing weak semantic associations and the difficulty in achieving collaborative optimization between geometric shapes, expression parameters, and texture features.
[0057] For the rendered 2D image, we first generate K different block-level random mask views, and then input the K mask views into the image encoder of the image coding model (such as CLIP) to obtain the visual representation. Finally, aggregation is performed to obtain the cross-view Figure 1 The overall visual representation after aggregation is expressed as: .
[0058] Block-level random mask views generate K two-dimensional images from different perspectives (such as regions and occlusion conditions) and apply a set of regional masks on each image to form a set of local image samples with semantic region attention, which is used to enhance the robustness and generalization ability of image-text semantic alignment.
[0059] To reduce the error caused by the modality difference between image and text encoding, this paper designs a bidirectional cross-contrast loss function to align the representations between visual and text modalities. Let τ be the temperature factor and the number of batch samples be B. The loss function is defined as follows: ; The first log item is the image With all text Comparison: The molecule is and The corresponding positive sample The similarity of and all (include and other negative samples); the second log item is the sum of the similarities of the text With all images The comparison is symmetrical with the first log term.
[0060] in, It is the overall visual representation after being encoded by the image encoder of the image-text coding model; b It is the text representation encoded by the text encoder of the image-text encoding model; 、 is the cosine similarity, and the similarity scores of image to text and text to image are calculated respectively, which are used as optimization reference during the training process. Specifically expressed as: , .
[0061] Furthermore, to ensure visual consistency between the rendered image and the original image, the present invention introduces image reconstruction losses, including but not limited to the following: pixel error loss, which measures the difference between the generated image and the original image in pixel space; landmark position error, which is based on the Euclidean distance constraint between 2D and 3D facial key points; regularization loss, which imposes a priori constraints on model parameters or representation space; and emotion consistency error, which uses a pre-trained expression recognition model to extract emotional features and minimize the semantic difference with the target image. These loss functions can be combined to form an image-level reconstruction loss term (i.e., the loss between the rendered image and the original image), denoted as: ; The loss terms correspond to pixels, landmarks, regularization and sentiment consistency supervision terms respectively. is the hyperparameter weight.
[0062] S107 Expression Classification: After completing 3D face reconstruction, based on the reconstructed 3D face mesh Automatically judge facial emotion categories, which is particularly suitable for application scenarios such as auxiliary diagnosis of mental illnesses and remote psychological intervention.
[0063] The expression classifier is preferably a multilayer perceptron (MLP), consisting of at least three layers of nonlinear, fully connected networks. This supports high-order nonlinear mapping of input features and classification decisions. The output is a probability distribution corresponding to each expression category (such as happiness, sadness, anger, fear, etc.). The output format supports both single-label classification and can be expanded to regression prediction of continuous emotion dimensions, such as valence and arousal scores, to meet different emotion analysis needs.
[0064] Furthermore, in order to adapt to the needs of distinguishing weak and atypical expressions in mental illness scenarios, the classifier introduces a pathological sample dataset for fine-tuning during the training phase to improve the recognition accuracy of slight emotional changes, making it suitable for intelligent medical tasks such as remote psychological intervention and emotional disorder monitoring.
[0065] On the public Now dataset, the reconstruction accuracy of the present invention was tested using three error indicators: median error (Median), mean error (Mean), and standard deviation (Std), achieving low error performances of 1.38mm, 1.63mm, and 1.42mm, respectively, indicating that it is significantly superior to existing methods in terms of geometric accuracy. In the emotion recognition task, the AffectNet dataset was used to evaluate the classification accuracy. The proposed model achieved an expression classification accuracy (E-ACC) of 0.75. Furthermore, in the continuous emotion dimension regression indicators, the valence prediction had a valence concordance correlation coefficient (V-CCC) of 0.79 and a root mean square error (V-RMSE) of 0.29; the arousal prediction had an arousal concordance correlation coefficient (A-CCC) of 0.71 and a root mean square error (A-RMSE) of 0.30, which fully verified its effectiveness and robustness in fine-grained emotion modeling and recognition tasks.
[0066] like Figure 2 As shown, the present invention also provides a text-guided 3D face reconstruction and emotion analysis device, comprising: An image acquisition unit, used to acquire user facial image information; The face reconstruction unit stores the 3D shape model trained by the above method, and the 3D shape model obtains the 3D mesh of the face based on the facial image information collected by the image acquisition unit. ; Emotion analysis unit, which stores the trained emotion analysis model and calculates the emotion analysis model based on the three-dimensional grid of the face. Determine facial emotion categories; A data storage unit for storing the results of three-dimensional face reconstruction and emotion classification; Human-computer interaction visualization unit, used to output 3D facial reconstruction and emotion classification results; The communication and control unit is used for information transmission and control between the processing unit and the human-computer interaction visualization unit.
[0067] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A text-guided 3D face reconstruction and emotion analysis method, characterized in that: The following steps are involved: (1) Obtain facial image information of the human body, perform preprocessing and annotation, and construct a training set; (2) Using the training set to train the 3D face reconstruction model, including: (2-i) Deconstruct the facial images in the training set into text descriptions T containing multiple facial attribute information, and convert the text descriptions T into a structured attribute matrix P attr ; (2-ii) Input the facial images in the training set into the 3D shape model to perform 3D face reconstruction, and at the same time introduce the attribute matrix into the 3D face reconstruction process to obtain a 3D face mesh with semantic consistency ; (2-iii) Input the facial images in the training set into the image style generator to synthesize the texture image, and at the same time introduce the attribute matrix into the texture image synthesis process to obtain a texture image with semantic consistency I UV ; (2-iv) 3D face mesh With texture image I UV Combining, rendering and generating a rendered image; (2-v) Construct loss functions between rendered images and original facial images, and between rendered images and text descriptions, and perform training. (3) Input the face image to be analyzed into the trained 3D shape model to obtain the 3D face mesh , the 3D face mesh Input the emotion analysis model to judge the facial emotion category.
2. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that: In step (1), the annotation content includes: basic attribute class, including gender and age; emotional expression class, including expression category and expression intensity score; geometric structure class, including facial contour type, basic form of facial features, and hair characteristics.
3. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that: Preprocessing includes data cleaning and data enhancement; Data enhancement includes: image uniform size scaling and standardization; facial area cropping and boundary detection; lighting condition adjustment and contrast enhancement; and noise reduction.
4. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that: Step (2-i) includes: (2-ia) Construct a set of face semantic related query sets Q={q1, q2,…, q M }, where each query q i Corresponding to a specific attribute dimension of the face; (2-ib) For the input facial image I, call the pre-trained image-text question answering model, jointly process the image and query pair, and generate the corresponding natural language response set A={a1,a2,…,a M }, where a i =BLIP(I,q i ); (2-ic) Using a priori language templates T tem , will be collected Organized into a text description T containing multiple facial attribute information; (2-id) Obtain the global semantic embedding vector z of the text description T through the multimodal model T , embed the global vector z T Mapping to a structured attribute matrix P attr .
5. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that: Step (2-ii) includes: (2-iia) Input facial images into the multi-channel feature encoding network , obtain individual characteristic parameters , dynamic expression parameters and head pose parameters ; (2-iib) The attribute matrix P attr Input nonlinear transformation module and , respectively obtain the semantic modulation vectors related to geometric shape modeling m s and semantic modulation vectors related to expression modeling m e : m s = σ(Γ s ( P attr )), m e = σ(Γ e ( P attr )) in Represents the Sigmoid activation function; (2-iic) Individual characteristic parameters and semantic modulation vector Perform element-by-element multiplication to obtain the modulated individual parameters ; Set the expression parameters and semantic modulation vector Fusion generates modulated dynamic parameters : , ; (2-iic) will 、 、 After decoding the 3D shape model, a 3D face mesh with semantic consistency is obtained. .
6. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that: Step (2-iii) includes: (2-iiia) Input the facial image into the encoding network to extract the initial latent vector z; (2-iiib) The attribute matrix P attr Fuse with the initial latent vector z to form a joint feature expression u = f(z, P attr ),in represents the nonlinear fusion mapping function; (2-iiic) Use the joint feature expression u as the input of the StyleGAN2 model and output a texture image with semantic consistency I UV .
7. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that: Step (2-v) includes: (2-va) For the rendered image, generate K different block-level random mask views and input the K mask views into the image-text encoding model to obtain visual representation , and finally aggregate to obtain the image feature vector : ; (2-vb) Input the text description T into the image-text encoding model to obtain the text feature vector b ; (2-vc) Construct a bidirectional cross-contrast loss function between images and text l align : ; τ is the temperature factor; B is the number of samples; , is the image-to-text similarity score; , is the similarity score from text to image; (2-vd) Constructing image reconstruction loss function l rec : ; λ 1. λ 2. λ 3. λ 4 is the hyperparameter weight; l pix 、 l land 、 l reg 、 l emo They correspond to pixel error loss, landmark position error loss, regularization loss, and sentiment consistency error loss respectively; After the training is completed, the trained 3D shape model is obtained.
8. The text-guided 3D face reconstruction and emotion analysis method according to claim 1, characterized in that: In step (3), the sentiment analysis model is a multi-layer perceptron.
9. A text-guided 3D face reconstruction and emotion analysis device, characterized in that: include: The data acquisition module obtains facial image information of the human body and performs preprocessing and annotation to build a training set; The model training module uses the training set to train the 3D shape model, including: Deconstruct facial images in the training set into text descriptions containing multiple facial attribute information, and convert the text descriptions into a structured attribute matrix; The facial images in the training set are input into the 3D shape model for 3D face reconstruction. At the same time, the attribute matrix is introduced into the 3D face reconstruction process to obtain a 3D face mesh with semantic consistency. ; The facial images in the training set are input into the image style generator to synthesize the texture image, and the attribute matrix is introduced into the texture image synthesis process to obtain a texture image with semantic consistency. I UV ; 3D face mesh With texture image I UV Combining, rendering to generate a rendered image; Construct loss functions between rendered images and original facial images, and between rendered images and text descriptions, and perform training. The 3D face reconstruction and emotion analysis module stores the trained 3D shape model and emotion analysis model. The 3D shape model generates a 3D face mesh based on the face image to be analyzed. , the emotion analysis model is based on the 3D mesh of the face Determine the facial emotion category.
10. A text-guided 3D face reconstruction and emotion analysis device, characterized in that: include: An image acquisition unit, used to acquire user facial image information; A face reconstruction unit stores a three-dimensional shape model trained using the method described in any one of claims 1 to 8, wherein the three-dimensional shape model obtains a three-dimensional face mesh based on the facial image information collected by the image acquisition unit. ; Emotion analysis unit, which stores the emotion analysis model and calculates the emotion analysis model based on the three-dimensional grid of the face. Determine facial emotion categories; A data storage unit for storing the results of three-dimensional face reconstruction and emotion classification; Human-computer interaction visualization unit, used to output 3D facial reconstruction and emotion classification results; The communication and control unit is used for information transmission and control between the processing unit and the human-computer interaction visualization unit.
Citation Information
Patent Citations
AI digital human automatic expression generation system based on voice driving
CN118800274A
Multi-mode multi-factor depression identification system fusing emotion information
CN119786021A
High-fidelity three-dimensional face model generation method based on natural text description
CN115984485A
Three-dimensional face reconstruction method based on CLIP model
CN116563457A
Texture-controllable three-dimensional fine face reconstruction method and device based on sketch input
CN118553001A