Face feature decoupling representation method and device with causal effect transmission and medium
By integrating nonlinear structural causal model and graph attention network on the face data set, the problem of insufficient causal effect transmission in the existing technology is solved, effective decoupling and transmission of causal features is achieved, and the explanatory and generalization capabilities of the model are improved.
Patent Information
- Application Number
- CN202510112695.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The existing face decoupling method fails to effectively consider the transmission of causal effects when interfering with potential representations, resulting in the inability to correctly decouple the causal relationship between the parent node and the child node in the latent space.
A nonlinear structural causal model (SCM) combined with graph attention network (GAT) is used to integrate the causal structure model in a variational autoencoder, pass the causal effect through GAT, and design the loss function to train the causal representation model.
It realizes efficiently capturing causal features on the face dataset and learning causal structures, transmitting causal effects, improving the interpretability and transferability of feature representations, avoiding the situation where the generation results are contrary to common sense.
Smart Images

Figure CN120047983A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of causal discovery, image generation, and disentangled representation learning. More specifically, it relates to a method, apparatus, and medium for disentangled representation of facial features with causal effect transmission. Background Art
[0002] With the rapid development of deep learning and artificial intelligence technologies, traditional centralized learning methods have gradually revealed their deficiencies in dealing with large-scale data, heterogeneous tasks, and complex systems. Especially in the context of massive data and diverse tasks, how to efficiently perform task allocation, feature representation, and model training has become a major challenge in current research. In this context, disentangled representation learning, as an emerging technical framework, has gradually become an effective tool for solving this problem.
[0003] The core idea of disentangled representation learning is to decouple data representation and task processing, so that the learning model can better handle diverse tasks without relying on the global task objective. Compared with traditional unified models, disentangled representation learning can independently learn features at different levels by performing hierarchical representation of the input data, thus improving the generalization ability of the model in a multi-task environment. This method is particularly suitable for scenarios with extremely high requirements for real-time performance, flexibility, and efficiency in modern applications, such as autonomous driving, personalized recommendation, intelligent manufacturing, etc.
[0004] Currently, the popularization of advanced technologies such as 5G, Internet of Things (IoT), edge computing, and reinforcement learning provides strong support for the application of disentangled representation learning. The low-latency and high-bandwidth characteristics of 5G technology make edge computing an important platform for achieving efficient task processing. On this basis, disentangled representation learning can flexibly handle data and task allocation problems between edge computing and cloud computing. In this way, computing tasks can be intelligently scheduled and processed at local edge nodes or in the cloud according to their real-time performance and resource requirements, thereby improving the efficiency of the overall system. The rise of disentangled representation learning provides a new solution for dealing with multi-tasks and large-scale data. Especially in the current highly dynamic and complex application environment, disentangled representation learning not only improves the efficiency and performance of the model, but also lays a foundation for achieving more intelligent and efficient task processing.
[0005] Decoupled representation learning aims to learn a model that can identify and decouple the underlying factors hidden in observable data. This process of separating the varying underlying factors into semantically meaningful variables helps in learning an interpretable representation of the data, which mimics the meaningful understanding process of humans when observing an object or relationship. When humans observe an object, we try to understand various attributes of the object (such as shape, size, and color, etc.) and combine certain prior knowledge. Existing end-to-end black-box deep learning models adopt a shortcut strategy of directly learning object representations to conform to the data distribution and discrimination criteria, failing to extract hidden attributes with similar human generalization abilities. Thus, decoupled representation learning is proposed. However, all existing decoupling methods on human faces can intervene in these latent representations and reflect them onto the image after obtaining the latent representations of image causal features. But the methods they use for intervention are to change a certain dimension of the latent representation, without considering the transmission of causal effects. Therefore, even in the latent space, the causal relationship between the parent node and the child node cannot be ignored.
[0006] In view of the above research background, aiming at the problems of latent representation intervention and causal feature decoupling in the current big data era, a causal decoupling method integrating graph attention network is studied and designed, optimizing the semi-supervised decoupling model, combining relevant knowledge and technologies such as variational autoencoders, graph attention networks, and generative adversarial networks, and using a small amount of label information to learn a causal feature decoupling method on a stable, efficient, and reliable face dataset. Summary of the Invention
[0007] To solve the above technical problems, the present invention provides a face feature decoupled representation method, device, and medium with causal effect transmission. Considering the complex attribute features on the face dataset, in order to ensure that the structural causal model (SCM) can learn a causal structure that conforms to the prior, a non-linear SCM is used to process the data, and a graph attention network (GAT) is used to achieve the transmission of causal effects.
[0008] In a first aspect, the present invention provides a face feature decoupled representation method with causal effect transmission, the method comprising:
[0009] Establish a variational autoencoder, and in response to the input face data, learn a latent feature distribution from the latent space based on the variational autoencoder;
[0010] Integrate the structural causal model into the variational autoencoder to model the causal effect, so as to extract causal features from the latent feature distribution;
[0011] Establish a graph attention network, and when decoupling, input the causal features into the graph attention network to transmit the causal effect;
[0012] Design a loss function, using a variational autoencoder, a discriminator, a causal structure model, and a graph attention network as causal representation models, and use the loss function to train the causal representation models.
[0013] Further, the variational autoencoder includes an encoder and a decoder; wherein, the encoder is used to map the input data to a latent space representation and control the encoding distribution through mean and variance parameters, and the decoder reconstructs the original data based on the sampled latent space vector.
[0014] Further, establish a variational autoencoder through the following method:
[0015] Determine the loss function of the variational autoencoder, expressed as:
[0016]
[0017] In the formula, is the loss of the variational autoencoder; q(z|x), p(x|z), and p(z) are the latent variable, the reconstruction distribution, and the prior distribution respectively; x is the observed data, z is the latent endogenous variable; KL is the KL divergence function; E q(z|x) [logp(x|z)] is the pixel-based cross-entropy loss;
[0018] Use the adversarial similarity loss to replace the pixel-based cross-entropy loss E q(z|x) [logp(x|z)], and the adversarial similarity loss is expressed as:
[0019]
[0020] In the formula, Dis b (x) is the feature vector representation mined by the discriminator in each layer, satisfying the Gaussian distribution is the image generated by the decoder, and design the loss of the discriminator by the decoder as where z is encoded by the encoder Enc(x), N is the Gaussian distribution, Dis b (x) is the discriminant loss of the b-th layer, is the discriminant loss of the output result, I is the identity matrix, Dis(x) is the total loss of all discriminant layers, and Dec(z) is the decoding result of the latent variable;
[0021] Use and to update the gradient in the decoder:
[0022]
[0023] where λ is the weight parameter, Θ Dec is the gradient of the decoder, and is the derivative update of the gradient of the decoder;
[0024] Given the observed data x and the corresponding latent endogenous variable z, the latent exogenous variable ε of z, and the supervised label l, with the maximized log-likelihood log pΘ (x) as the objective, introduce the variational distribution q φ (ε,z|x,l) to approximate the true posterior distribution p Θ (ε,z|x,l), where φ and Θ are the parameters of the variational distribution and the model distribution respectively;
[0025] Introduce the KL divergence into the log-likelihood, rearrange and simplify it, and use Bayes' rule to rewrite log pΘ (x) as the conditional probability log pΘ (x|z) and the prior probability p Θ (z), thus obtaining the following evidence lower bound ELBO:
[0026]
[0027] where D KL (q φ (ε,z|x,l)||p Θ (ε,z|l)) is the KL divergence;
[0028] Simplify the variational distribution through the following formula:
[0029] q φ (ε,z|x,l) = q φ (ε|x,l)f(z=(I - A T ) -1 ε)
[0030] where f is a non-linear or linear function, and A is the weighted binary adjacency matrix of the causal graph;
[0031] Based on the simplified variational distribution, obtain a new evidence lower bound, expressed as:
[0032]
[0033] where is the mean of the variational of the overall observed data, is the mean of the variational posterior of the latent variable.
[0034] Furthermore, integrate the structural causal model into the variational autoencoder to model the causal effect, so as to extract causal features from the latent feature distribution, including:
[0035] Learn an adjacency matrix in a supervised manner given the causal relationship, where the weights of the adjacency matrix reflect the causal strength;
[0036] Based on the adjacency matrix, determine z through the following formula:
[0037] z = f 2 ((I - A T ) -1 f 1 (ε))
[0038] where A is the weighted binary adjacency matrix of the causal graph, and f 1 and f 2 are non - linear and linear functions;
[0039] Rewrite z = f 2 ((I - A T ) -1 f 1 (ε)) as
[0040] In , add a new term KL(q(z i |l)||p(l)) to get:
[0041]
[0042] where KL(q(z i |l)||p(l)) is the KL divergence between the prior and the i - th dimensional latent variable; z i is the i - th dimension of z;
[0043] Supervise the generation of causal features through the first i terms of z.
[0044] Furthermore, establish a graph attention network to transfer causal effects when decoupling the causal features input into the graph attention network, including:
[0045] By intervening on z i and using the supervision label l, obtain z without causal aggregation:
[0046] z = do(Supervise([z] 1 ,[z] 2 ,...,[z] i ; l) i = c)
[0047] where c is the number of latent endogenous variables [z] i with a distance from z within the interval [-i, i], where [z] i i Represents the i-th z;
[0048] Based on the graph attention network GAT, the graph attention mechanism is used to reconstruct the formula by dynamically assigning weights to neighbor nodes, and each z is processed separately i , concatenate the processed features, and transpose them into the reconstructed vector z':
[0049] z' = transpose([GAT(z 1 ), GAT(z 2 ),..., GAT(z i )])
[0050] In the formula, transpose is the matrix transpose, and GAT(z i ) applies the attention mechanism to the i-th dimensional latent variable;
[0051] The reconstructed vector z' learns the weights of the causal edges in A, aggregates the influence of the parent nodes, and by training GAT, focuses on the transmission of causal effects;
[0052] Let z = {z i |1 < i ≤ N, z i ∈ R F}, where N is the number of z, F is the dimension of z, and the output z' = {z' i |1 < i ≤ N, z' i ∈ R F} is generated by GAT, combining the aggregated information from adjacent nodes;
[0053] Train the weight matrix W ∈ R F×F , R is the real number field, and the weight matrix represents the relationship between the input and output features;
[0054] Use the attention coefficient e ij to represent the importance degree between the j-th dimension z j of z and the i-th dimension z i , e ij = a(Wz i , Wz j ), is the weight matrix of z i , is the weight matrix of z j , and α ij is the weight of the edge from z i to z j , expressed as:
[0055]
[0056] In the formula, softmaxi (e ij ) is the activation function for normalization, exp is the exponential function with the base of the natural constant, k is the k-th attention head, and N i is the total number of parent nodes, and e jk is the attention correlation score between nodes j and k;
[0057] The attention mechanism on a single node is used to formulate α ij :
[0058]
[0059] In the formula, LeakyRelu is a non-linear activation function, is the weight of the k-th dimension z of z k ;
[0060] Taking as the final output feature of each node, the multi-head attention mechanism is adopted to enhance stability. The multi-head attention mechanism uses K independent attention heads and converts z to z′ through the following formula:
[0061]
[0062] In the formula, ‖ represents the concatenation operation, represents the normalized attention coefficient calculated by the k-th attention head, and W k is the weight matrix corresponding to the input linear transformation.
[0063] Furthermore, a loss function is designed, including:
[0064] According to the KL divergence, reconstruction loss decoder loss and supervision loss to establish a loss function, expressed as:
[0065]
[0066] In the formula, is the total loss of the model; β is the weight parameter;
[0067] KL is the KL divergence, used to fit the standard Gaussian distribution, expressed as:
[0068]
[0069] is used to shorten the distance between the i-th dimension of z and the true label, expressed as:
[0070]
[0071] where z i represents the i-th dimension of z, l represents the label, and i is the number of attributes of interest.
[0072] Further, when designing the loss function, using the variational autoencoder, discriminator, causal structure model, and graph attention network as the causal representation model, after training the causal representation model using the loss function, the method further includes:
[0073] Analyzing the identifiability of the causal representation model.
[0074] Further, analyzing the identifiability of the causal representation model includes:
[0075] Let ∼ be a binary relation on the parameter Θ, and define:
[0076]
[0077] where Θ(E, D, A, T, λ) are the model parameters, E is the encoder, D is the decoder, A is the weighted binary adjacency matrix of the causal graph, and T is the time step; are the encoded output results of the corresponding parameters respectively, B 1 is the parameter matrix, B 2 is the bias vector, b 2 , b 1 is the scalar, x is a single observed data, is the linear mapping, is the inverse mapping transformation, is the entire data set;
[0078] Determine that the number Θ is identifiable when the following conditions are met:
[0079] The set has a measure of zero, where φ ξ represents the characteristic function of the density p Θ (ε, z|l) = p ε (ε)p Θ (z|l) in p ξ ;
[0080] The function D as the decoder is differentiable, and the Jacobian matrix of the decoder function remains full rank;
[0081] The sufficient statistic T i,s (z i ) is non-zero at all points, T i,s (z i ) represents the s-th statistic of the i-th dimensional variable z i ;
[0082] The additional observations satisfy li ≠ 0.
[0083] In a second aspect, the present invention provides a face feature decoupled representation device with causal effect transmission, and the device includes:
[0084] A feature distribution acquisition module, configured to establish a variational autoencoder, and in response to the input face data, learn a latent feature distribution from the latent space based on the variational autoencoder;
[0085] A causal feature extraction module, configured to integrate a structural causal model into the variational autoencoder, model the causal effect, and extract causal features from the latent feature distribution;
[0086] A causal feature decoupling module, configured to establish a graph attention network, and transmit the causal effect when decoupling the causal features input into the graph attention network;
[0087] A model training module, configured to design a loss function, use the variational autoencoder, discriminator, causal structure model, and graph attention network as a causal representation model, and train the causal representation model using the loss function.
[0088] In a third aspect, the present invention provides a readable storage medium storing one or more programs, and the one or more programs can be executed by one or more processors to implement the method as described above.
[0089] The present invention has at least the following beneficial effects:
[0090] The present invention explores for the first time the decoupled representation learning with causal effect transmission, solves the problem that the change of intervention features cannot trigger the corresponding causal features, and thus avoids the situation where the generated results are contrary to common sense. In order to efficiently capture causal features and learn causal structures, a non-linear or linear SCM is constructed, which encodes latent exogenous variables into latent endogenous variables, and a discriminator with hierarchical feature loss is designed to replace the pixel-level loss in the VAE, thereby enhancing the decoupled representation. In order to transmit causal effects, a GAT intervention mechanism is designed to aggregate the causal information of adjacent nodes by continuously learning the weights of the edges in the causal matrix, thus naturally decoupling the causal features. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 Shows a schematic diagram of a system architecture for implementing a face feature decoupled representation method with causal effect transmission according to an embodiment of the present invention.
[0092] Figure 2 Shows the flow of a face feature decoupled representation method with causal effect transmission according to an embodiment of the present inventionFigure 1 。
[0093] Figure 3 Shows the process of a face feature decoupled representation method with causal effect transmission according to an embodiment of the present invention Figure 2 。
[0094] Figure 4 Shows the structural diagram of a face feature decoupled representation device with causal effect transmission according to an embodiment of the present invention. Detailed implementation manners
[0095] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners. The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and specific examples, but this is not a limitation to the present invention. For the various steps described herein, if there is no necessity for a sequential relationship between them, the order in which they are described as examples herein should not be regarded as a limitation. Those skilled in the art should know that they can be adjusted in order as long as the logic between them is not destroyed and the entire process cannot be realized.
[0096] Most traditional face feature representation learning methods rely on pixel-level reconstruction or representation learning, and it is difficult to effectively capture causal relationships and their intervention effects on features. In many application scenarios, understanding and transmitting causal effects between features is crucial for improving the interpretability and generalization ability of the model. To solve this problem, an embodiment of the present invention provides a face feature decoupled representation method with causal effect transmission. This method combines the advantages of variational autoencoders (VAEs) and generative adversarial networks (GANs), performs causal structure learning and causal feature intervention in the latent space, thereby achieving more robust and meaningful face feature decoupling. The innovations are as follows: First, it uses the intermediate layer features of the discriminator as the reconstruction loss, instead of the pixel-level cross-entropy loss in traditional VAEs, ensuring the effective learning of the causal structure; second, an additive noise model is used to process the latent space with partial labels, thereby transforming latent exogenous variables into latent endogenous variables to achieve the decoupling of causal features in the latent space; finally, by introducing a graph attention network (GAT) mechanism, causal effects are effectively transmitted, and a probabilistic generation mechanism is injected into the model to further optimize the decoupling effect. This method can effectively mine and transmit the causal relationships between face features, improve the interpretability and transferability of feature representations, and has strong application potential, especially suitable for tasks such as face recognition and expression analysis.
[0097] Such as Figure 1As shown in the figure, it is a schematic diagram of the system architecture for implementing the decoupled representation method of face features with causal effect transmission. This system architecture consists of four parts: a variational autoencoder (VAE), a discriminator, a causal structure model (SCM), and a graph attention network (GAT). Based on the complexity of the prior causal structure of the input image data, the SCM is divided into a non-linear SCM and a linear SCM.
[0098] The VAE mainly consists of two parts: an encoder and a decoder. Among them, the encoder is responsible for mapping the input data to the latent space representation and controlling the encoding distribution through the mean and variance parameters. The decoder reconstructs the original data based on the sampled latent space vector. The output space of the encoder is a Gaussian distribution, and when new data is input, it will be encoded and represented in the latent space according to the rules of variational inference.
[0099] Considering the characteristic that causal relationships are invariant in the high-dimensional feature space, we propose a new feature extraction method. Specifically, we introduce the Discriminator as the core component of feature extraction and input both the original input image and the output result of the GAT into the discriminator. Different from the VAE using a pixel-level loss function, we use the hierarchical feature loss of the discriminator to guide the model learning. This method can effectively capture the causal features in the data and avoid the noise interference that may be brought about by directly extracting features in the pixel space.
[0100] The SCM is divided into two types: linear and non-linear, which are used to process simple and complex data sets respectively. Since the features on the human face are mixed and difficult to distinguish, the non-linear SCM is used to process the face data set. The latent feature distribution after passing the face data through the Encoder is input into the non-linear SCM to obtain the causal feature distribution, and at the same time, the weights of the causal graph edges are updated.
[0101] Although the causal features obtained after passing through the SCM reflect the causal relationships between variables, these features do not fully consider the influence mechanism from the parent nodes. Therefore, we propose a causal effect transmission method for the GAT. Inputting the extracted causal features into the GAT can effectively simulate the information flow process in the causal graph. The attention mechanism of the GAT can adaptively learn the influence weights of different parent nodes on the current node, not only considering the direct causal relationship but also being able to capture the complex causal effects generated by the joint action of multiple parent nodes, enabling the model to more comprehensively understand and express the causal dependencies between variables.
[0102] Considering that the performance of the GAT on the face data set will vary with the change in the number of attention heads, we change the number of attention heads while keeping the SCM parameters unchanged.
[0103] In the above-mentioned face feature decoupled representation method with causal effect transmission, the process of decoupled representation is as follows:
[0104] (1) Considering that the single image in the original face dataset is too large, operations such as cropping, rotating, and normalizing the dataset are required. The processed data is input into the Encoder to obtain the latent distribution of features. To improve the training efficiency and generalization ability of the model, we first use a face detection algorithm to locate and crop the face region, and uniformly adjust the image size to 224×224 pixels. Subsequently, through data augmentation by randomly rotating ±10 degrees, the robustness of the model to pose changes is improved. The pixel values of the image are mapped to the interval [-1, 1] after normalization processing to accelerate network convergence. After these preprocessing steps, the data is fed into the Encoder composed of a multi-layer convolutional neural network. Under the prior p(l), the Encoder compresses the high-dimensional image features into a low-dimensional latent space ε = E(x, l) + ζ through dimensionality reduction operations, where ε is the latent exogenous variable, l is the label, and ζ is the noise variable. This makes the distribution of similar face images in the latent space more compact, laying the foundation for subsequent face recognition tasks.
[0105] (2) The latent exogenous variable ε needs to be processed by a non-linear SCM before it can be transformed into an endogenous variable related to face attributes. The overall encoding process is q φ (z, ε|x, l) ≡ q(z|ε)q ζ (ε - E(x, l)), and the decoding process is p Θ (x|z, ε, l) = p Θ (x|z) ≡ p ξ (x - D(z)), where the joint prior of z and ε is p Θ (ε, z|l) = p ε (ε)p Θ (z|l), where p ε (ε) satisfies the standard normal distribution N(0, I) (I represents the identity matrix), indicating that z and ε need to be restricted by the label and KL divergence. The prior p Θ (z|l) is a factored Gaussian distribution conditional on the label l where F 1 and F 2 represent arbitrary functions. To align z and l in the initial dimension of the latent space, it is set that F 1 (l) = l and F 2 (l) ≡ 1. The mean μ(z) and variance σ(z) of z are determined by the sufficient statistic T(z) = (μ(z), σ(z)) = (T 1,1 (z 1 ),..., T n,2 (z n)). In this case, the statistics of T(z) satisfy E[T(z)]=μ(z) and Var[T(z)]=σ 2 (z), which provides an important basis for the subsequent analysis of probability distribution.
[0106] (3) Given that intervention on a node will lead to changes in all levels of child nodes with causal relationships, the advantage of our method is that it is more natural than using the disentangled features in the latent space obtained by traditional models to generate images. The latent endogenous variable z needs to pass through GAT to aggregate the causal information of adjacent nodes so that the causal effect can be transmitted between parent and child nodes. i and the jth dimension z j There is a causal effect i →z j When the method is Right i Intervene and let z j Change accordingly.
[0107] Specifically, if Figure 2 As shown, the process of the facial feature decoupling representation method with causal effect transmission is shown in FIG. Figure 1 The facial feature decoupling representation method with causal effect transmission can be implemented by the following steps S10 to S40.
[0108] S10: Establish a variational autoencoder, and in response to input face data, obtain a potential feature distribution from a latent space based on the variational autoencoder.
[0109] In order to deeply explore the causal structure in face data, it is necessary to learn feature distribution from the latent space.
[0110] In some implementations, a variational autoencoder VAE is designed by:
[0111] Loss of VAE It includes a reconstruction loss and KL divergence, where q(z|x), p(x|z), and p(z) are the latent variables, the reconstruction distribution, and the prior distribution, respectively:
[0112]
[0113] In order to achieve a more accurate feature distribution in the latent space, an adversarial similarity loss is used Replace the pixel-based cross entropy loss E q(z|x) [logp(x|z)]. Considering that images of the same category present highly similar feature distributions, is from the discriminator Dis b The middle layer of (·) is derived as Among them, Dis b (x) is the feature vector representation mined by the discriminator in each layer. These obtained feature representations satisfy the Gaussian distribution where is the image generated by the decoder. Through the decoder in the VAE, the loss of the discriminator can be designed as where z is encoded by the encoder Enc(x). Using and to update the gradient in the decoder where λ is the weight parameter. Given the observed data x, the corresponding latent endogenous variable z, the latent exogenous variable ε of z, and the supervised label l, the goal is to maximize the log-likelihood: log pΘ (x). To effectively utilize the label and variables, a variational distribution q φ (ε, z|x, l) is introduced to approximate the true posterior distribution p Θ (ε, z|x, l), where φ and Θ are the parameters of the variational distribution and the model distribution respectively. To maximize the log-likelihood, in this embodiment, the KL divergence is introduced into the log-likelihood, and then it is rearranged and simplified. Using Bayes' rule, we rewrite log pΘ (x) as the conditional probability log pΘ (x|z) and the prior probability p Θ (z), thus obtaining the following evidence lower bound (ELBO):
[0114]
[0115] D KL (q φ (ε, z|x, l)||p Θ (ε, z|l)) is the KL divergence. The variable z is derived from ε through the SCM and subsequent non-linear or linear transformations. This process simplifies the variational posterior distribution:
[0116] q φ (ε, z|x, l) = q φ (ε|x, l)f(z = (I - A T ) -1 ε)
[0117] f(·) is a non-linear or linear function, and A is the weighted matrix. The new evidence lower bound can be obtained:
[0118]
[0119] The first term is the reconstruction loss, the second term ensures that ε conforms to the standard normal distribution, and the third term constrains the first i dimensions of z to correspond one-to-one with the label. It can be maximized by to minimize the KL divergence, and thus completing the design of the VAE, making the model pay more attention to the exploration of latent features and having better image generation effects.
[0120] S20: Integrate the structural causal model into the variational autoencoder to model the causal effect for extracting causal features from the latent feature distribution.
[0121] By integrating the structural causal model (SCM) into the VAE, the causal effect is modeled to extract causal features from the distribution of the latent space. Given the causal relationship, the method learns an adjacency matrix in a supervised manner, and the weights of the matrix reflect the causal strength. Then the structural causal model is used as the prior p z .
[0122] z = f 2 ((I - A T ) -1 f 1 (ε))
[0123] where A is the weighted binary adjacency matrix of the causal graph (directed acyclic graph, i.e., DAG), I is the identity matrix, and f 1 and f 2 are non-linear or linear functions. The exogenous noise ε follows the Gaussian distribution N(0, I). The parameter set β(f 1 , f 2 , A) needs to be optimized in the parameter space B. To adapt to the SCM in the method, rewrite z = f 2 ((I - A T ) -1 f 1 (ε)) as z = f β ((I - A T ) -1 ε). f(·) is invertible, and continue to rewrite it as ε in the SCM is encoded by x. In , a new term KL(q(z i |l)||p(l)) is added:
[0124]
[0125] so as to supervise the generation of causal features through the first i terms of z. The model uses z = f β ((I - A T ) -1ε) Implement the mining of causal features, which learns z by generating a distribution consistent with the supervised features. Initially, we fill A with zeros to approximate the causal graph, and then train and update f 2 ((I - A T ) -1 f 1 (ε)) and the parameters of, thus obtaining the causal features, that is, the latent endogenous variable z. This method realizes the learning and inference of causal relationships by combining the DAG structure, the encoder network, and the KL divergence regularization term, while ensuring the interpretability and generalization ability of the model.
[0126] S30: Establish a graph attention network, and input the causal features into the graph attention network to transfer causal effects during decoupling.
[0127] In some embodiments, the causal effects are transferred during decoupling by the following method.
[0128] By intervening on the first i dimensions of z i and using the supervised label l, z without causal aggregation is obtained:
[0129] z = do(Supervise([z] 1 , [z] 2 ,..., [z] i ; l) i = c)
[0130] c is the number of latent endogenous variables [z] i with a distance from z i in the interval [-i, i], where [z] i represents the i-th z. To further characterize causal transmission, the graph attention mechanism is used to reconstruct the formula z by dynamically assigning weights to neighbor nodes. Different from the traditional GAT model, our model processes each z i separately, and then splices the processed features and transposes them to z':
[0131] z' = transpose([GAT(z 1 ), GAT(z 2 ),..., GAT(z i )])
[0132] The reconstructed z' learns the weights of the causal edges in A through GAT and aggregates the influence of the parent nodes. By training GAT, the model can pay more attention to the transmission of causal effects. Let z = {z i |1 < i ≤ N, z i ∈R F}, where N is the number of z and F is its dimension. The output z′ = {z′ i |1 < i ≤ N, z′ i ∈R F} is generated by GAT, which combines the aggregated information from neighboring nodes. To derive the output from the input, a weight matrix W ∈ R F×F is trained, which represents the relationship between the input and output features. The attention coefficient e ij represents the degree of importance between z j and z i , and e ij = a(Wz i , Wz j ). α ij is the weight of the edge from z i to z j .
[0133]
[0134] GAT is a single-layer feed-forward neural network, parameterized by a weight vector α, and uses the non-linear activation function LeakyRelu. The attention mechanism on a single node is used to formulate α ij :
[0135]
[0136] In addition, the method α ij calculates the linear aggregation of the corresponding features, which will be used as the final output feature of each node We adopt the multi-head attention mechanism to enhance the stability of the model. This mechanism uses K independent attention heads to transform z into z′ through the formula . ‖ represents the concatenation operation, represents the normalized attention coefficient calculated by the k-th attention head, and W k is the weight matrix for the linear transformation of the corresponding input. The multi-head attention mechanism effectively captures different levels of relationships and feature dependencies between nodes through the parallel calculation of multiple independent attention heads, improving the expression ability and stability of the model in complex graph structures.
[0137] S40: Design a loss function. Using the variational autoencoder, discriminator, causal structure model, and graph attention network as causal representation models, train the causal representation models using the loss function.
[0138] Design a loss function which is derived from the discriminator and VAE. In , the loss of the discriminator is used to distinguish between the real image x and the generated image The total loss of the method consists of four parts: KL divergence, reconstruction loss Decoder loss and supervision loss
[0139]
[0140] Among them, KL divergence is used to fit the standard Gaussian distribution:
[0141]
[0142] for optimizing the decoder for shortening the distance between the first i dimensions of z and the true label, where β is the weight parameter:
[0143]
[0144] z i represents the first i dimensions of z, l represents the label, and i is the number of attributes of interest. is the reconstruction loss. By comprehensively considering KL divergence, reconstruction loss, decoder loss, and supervision loss, this loss function not only ensures the effective learning of the model for the data distribution but also realizes the decoupling and supervision constraints of potential causal features, thereby improving the quality of the generated images and the accuracy of causal effect transmission.
[0145] In some embodiments, as Figure 3 shown, for the process of the face feature decoupling representation method with causal effect transmission Figure 2 , after step S40, the method further includes:
[0146] S50: Using a variational autoencoder, discriminator, causal structure model, and graph attention network as the causal representation model; analyzing the identifiability of the causal representation model.
[0147] This embodiment conducts an identifiability analysis and extends the identifiability results of causal representation learning based on the structural causal prior. Let ~ be a binary relation on Θ, and define:
[0148]
[0149] Define If B 1 is invertible and is a diagonal matrix whose diagonal elements are associated with z i then the model parameters are said to be identifiable. Assume that the observed data is sampled from the generative model, and the model parameters are Θ(E, D, A, T, λ). Under the following conditions: (1) The set has measure zero, where φ ξ represents p Θ (ε,z|l) = p ε (ε)p Θ the density p in (z|l) ξ characteristic function of; (2) the function D as the decoder is differentiable and its Jacobian matrix remains full rank; (3) the sufficient statistic T i,s (z i ) is non-zero at almost all points, where T i,s (z i ) represents the s-th statistic of the variable z i , where 1 ≤ i ≤ n, 1 ≤ s ≤ 2; (4) the additional observation satisfies l i ≠ 0. Then, the parameter Θ is identifiable. Through the strict constraints of the above conditions, we have proven the identifiability of the model parameters, thus ensuring that the causal generation model can accurately learn the data distribution and causal structure at the theoretical level, providing a robust theoretical support for practical applications.
[0150] In some embodiments, the method further includes step S60: performing a decoupling experiment on the face dataset on the graphics card A100 using PyCharm software, minimizing the loss mentioned in step S40 to implement the decoupling method with causal effect transmission, and verifying the performance of the model in terms of reconstruction quality, causal feature extraction, and generation effect.
[0151] The embodiment of the present invention also provides a face feature decoupling representation device with causal effect transmission, as Figure 4 shown, the device includes:
[0152] A feature distribution acquisition module 401, configured to establish a variational autoencoder, and in response to the input face data, learn a latent feature distribution from the latent space based on the variational autoencoder;
[0153] A causal feature extraction module 402, configured to integrate a structural causal model into the variational autoencoder, model the causal effect, and extract causal features from the latent feature distribution;
[0154] A causal feature decoupling module 403, configured to establish a graph attention network, and transmit the causal effect when decoupling the causal features input into the graph attention network;
[0155] A model training module 404, configured to design a loss function, use the variational autoencoder, discriminator, causal structure model, and graph attention network as the causal representation model, and train the causal representation model using the loss function.
[0156] In some embodiments, the variational autoencoder includes an encoder and a decoder; wherein, the encoder is configured to map input data to a latent space representation and control the encoding distribution through mean and variance parameters, and the decoder reconstructs the original data based on the sampled latent space vector.
[0157] In some embodiments, the feature distribution acquisition module is further configured to establish a variational autoencoder by the following method:
[0158] Determine the loss function of the variational autoencoder, expressed as:
[0159]
[0160] wherein, is the loss of the variational autoencoder; q(z|x), p(x|z), and p(z) are the latent variable, the reconstruction distribution, and the prior distribution respectively; x is the observed data, z is the latent endogenous variable; KL is the KL divergence function; E q(z|x) [logp(x|z)] is the pixel-based cross-entropy loss;
[0161] Use the adversarial similarity loss to replace the pixel-based cross-entropy loss E q(z|x) [logp(x|z)], and the adversarial similarity loss is expressed as:
[0162]
[0163] wherein, Dis b (x) is the feature vector representation mined by the discriminator in each layer, satisfying the Gaussian distribution is the image generated by the decoder, and the loss of the discriminator is designed as where z is encoded by the encoder Enc(x), N is the Gaussian distribution, Dis b (x) is the discriminant loss of the b-th layer, is the discriminant loss of the output result, I is the identity matrix, Dis(x) is the total loss of all discriminant layers, and Dec(z) is the decoding result of the latent variable;
[0164] Use and to update the gradient in the decoder:
[0165]
[0166] wherein, λ is the weight parameter, Θ Dec is, is the derivative update of the gradient of the decoder;
[0167] Given the observed data \(x\), the corresponding latent endogenous variable \(z\), the latent exogenous variable \(\epsilon\) of \(z\), and the supervised label \(l\), with the maximized log-likelihood \(\log\) pΘ (x) as the objective, introduce the variational distribution \(q\) φ (\(\epsilon,z|x,l\)) to approximate the true posterior distribution \(p\) Θ (\(\epsilon,z|x,l\)), where \(\varphi\) and \(\Theta\) are the parameters of the variational distribution and the model distribution respectively;
[0168] Introduce the KL divergence into the log-likelihood, rearrange and simplify it, and use Bayes' rule to rewrite \(\log\) pΘ (x) as the conditional probability \(\log\) pΘ (x|z) and the prior probability \(p\) Θ (z), thus obtaining the following evidence lower bound ELBO:
[0169]
[0170] In the formula, \(D\) KL (q φ (\(\epsilon,z|x,l\))||p Θ (\(\epsilon,z|l\))) is the KL divergence;
[0171] Simplify the variational distribution through the following formula:
[0172] q φ (\(\epsilon,z|x,l\)) = \(q\) φ (\(\epsilon|x,l\))\(f(z=(I - A\) T ) -1 \(\epsilon)\)
[0173] In the formula, \(f\) is a non-linear or linear function, and \(A\) is the weighted binary adjacency matrix of the causal graph;
[0174] Based on the simplified variational distribution, obtain a new evidence lower bound, expressed as:
[0175]
[0176] In the formula, is, is.
[0177] In some embodiments, the causal feature extraction module is further configured to:
[0178] Learn an adjacency matrix in a supervised manner given the causal relationship, and the weights of the adjacency matrix reflect the causal strength;
[0179] Based on the adjacency matrix, determine \(z\) through the following formula:
[0180] \(z = f\)2 ((I - A T ) -1 f 1 (ε))
[0181] where A is the weighted binary adjacency matrix of the causal graph, and f 1 and f 2 are non - linear and linear functions respectively;
[0182] Rewrite z = f 2 ((I - A T ) -1 f 1 (ε)) as
[0183] In , add a new term KL(q(z i |l)||p(l)), and we get:
[0184]
[0185] where KL(q(z i |l)||p(l)) is; z i is the i - th dimension of z;
[0186] Supervise the generation of causal features through the first i terms of z.
[0187] In some embodiments, the causal feature decoupling module is further configured to:
[0188] By intervening in z i and using the supervision label l, obtain z without causal aggregation:
[0189] z = do(Supervise([z] 1 ,[z] 2 ,...,[z] i ; l) i = c)
[0190] where c is the number of potential endogenous variables [z] i with a distance from z i in the interval [-i, i], where [z] i represents the i - th z;
[0191] Based on the Graph Attention Network GAT, use the graph attention mechanism to reconstruct the formula by dynamically assigning weights to neighbor nodes, process each z i individually, splice the processed features, and transpose them into a reconstructed vector z':
[0192] z′ = transpose([GAT(z 1 ), GAT(z 2 ),..., GAT(z i )])
[0193] where transpose is matrix transposition, and GAT(z i ) applies the attention mechanism to the i-th dimensional latent variable;
[0194] The reconstructed vector z′ learns the weights of the causal edges in A through GAT, aggregates the influence of the parent nodes, and trains GAT to focus on the transmission of causal effects;
[0195] Let z = {z i |1 < i ≤ N, z i ∈R F}, where N is the number of z, F is the dimension of z, and the output z′ = {z′ i |1 < i ≤ N, z′ i ∈R F} is generated by GAT and combines the aggregated information from adjacent nodes;
[0196] Train the weight matrix W ∈ R F×F , where R is the real number field, and the weight matrix represents the relationship between the input and output features;
[0197] Use the attention coefficient e ij to represent the importance degree between the j-th dimension z j and the i-th dimension z i of z, e ij = a(Wz i , Wz j ), is the weight matrix of z i , is the weight matrix of z j , and α ij is the weight of the edge from z i to z j , expressed as:
[0198]
[0199] where softmax i (e ij ) is the activation function for normalization, exp is the exponential function with the base of the natural constant, k is the k-th attention head, N i is the total number of parent nodes, and e jk is the attention correlation score between nodes j and k;
[0200] Formulate α using the attention mechanism on a single node ij :
[0201]
[0202] In the formula, LeakyRelu is a non - linear activation function, is the weight of the k - th dimension z of z k ;
[0203] Take as the final output feature of each node, and use the multi - head attention mechanism to enhance stability. The multi - head attention mechanism uses K independent attention heads and transforms z into z′ through the following formula:
[0204]
[0205] In the formula, ‖ represents the concatenation operation, represents the normalized attention coefficient calculated by the k - th attention head, and W k is the weight matrix of the corresponding input linear transformation.
[0206] In some embodiments, the model training module is further configured to:
[0207] Establish a loss function according to the KL divergence, reconstruction loss decoder loss and supervision loss as:
[0208]
[0209] In the formula, is the total loss of the model; β is a weight parameter;
[0210] KL is the KL divergence, used to fit the standard Gaussian distribution, expressed as:
[0211]
[0212] is used to shorten the distance between the i - th dimension of z and the true label, expressed as:
[0213]
[0214] In the formula, z i represents the i - th dimension of z, l represents the label, and i is the number of attributes of interest.
[0215] In some embodiments, the device further includes an analysis module, which is configured to: use a variational autoencoder, a discriminator, a causal structure model, and a graph attention network as a causal representation model; and analyze the identifiability of the causal representation model.
[0216] In some embodiments, the analysis module is configured to:
[0217] Let ∼ be a binary relation on the parameter Θ, and define:
[0218]
[0219] where Θ(E, D, A, T, λ) are the model parameters, E is the encoder, D is the decoder, A is the weighted binary adjacency matrix of the causal graph, and T is the time step; are the encoded output results of the corresponding parameters, B 1 is the parameter matrix, B 2 is the bias vector, b 2 , b 1 is a scalar, x is a single observed data, is a linear mapping, is the inverse mapping transformation, is the entire data set;
[0220] Determine that the number Θ is identifiable when the following conditions are met:
[0221] The set has a measure of zero, where φ ξ represents the characteristic function of the density p Θ (ε, z|l) = p ε (ε)p Θ (z|l) in p ξ ;
[0222] The function D as the decoder is differentiable, and the Jacobian matrix of the decoder function remains full rank;
[0223] The sufficient statistic T i,s (z i ) is non-zero at all points, T i,s (z i ) represents the s-th statistic of the i-th dimensional variable z i of z;
[0224] The additional observation satisfies l i ≠ 0.
[0225] It should be noted that the structures of the various face feature decoupling representation devices with causal effect transmission described in this embodiment belong to the same inventive concept as the previously described face feature decoupling representation method with causal effect transmission, and achieve the same beneficial effects through the same principle, which will not be elaborated here.
[0226] An embodiment of the present invention also provides a readable storage medium storing one or more programs, which can be executed by one or more processors to implement the method described in any of the above embodiments.
[0227] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present invention having equivalent elements, modifications, omissions, combinations (e.g., schemes where various embodiments intersect), adaptations or changes. The elements in the claims will be broadly interpreted based on the language used in the claims and are not limited to the examples described in this specification or during the implementation of this application, and the examples will be interpreted as non-exclusive. Thus, this specification and examples are intended to be considered only as examples, and the true scope and spirit are indicated by the following claims and the full scope of their equivalents.
[0228] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of their aspects) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. Additionally, in the above detailed description, various features can be grouped together to simplify the present invention. This should not be construed as an intention that the features of an invention not claimed are necessary for any claim. On the contrary, the subject matter of the present invention can be less than all the features of a particular embodiment of the invention. Thus, the following claims are incorporated herein as examples or embodiments into the detailed description, where each claim stands alone as a separate embodiment, and it is contemplated that these embodiments can be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the appended claims and the full scope of the equivalents to which those claims are entitled.
Claims
1. A facial feature decoupling representation method with causal effect transmission, characterized in that: The method comprises: Establishing a variational autoencoder, and in response to input face data, learning a potential feature distribution from a latent space based on the variational autoencoder; Integrating a structural causal model into the variational autoencoder to model causal effects to extract causal features from the latent feature distribution; Establishing a graph attention network, and inputting the causal features into the graph attention network to transmit causal effects when decoupling; A loss function is designed, and a variational autoencoder, a discriminator, a causal structure model, and a graph attention network are used as a causal representation model, and the loss function is used to train the causal representation model.
2. The facial feature decoupling representation method with causal effect transfer according to claim 1, characterized in that: The variational autoencoder includes an encoder and a decoder; wherein the encoder is used to map input data to a latent space representation and control the encoding distribution through mean and variance parameters, and the decoder reconstructs the original data based on the sampled latent space vector.
3. The facial feature decoupling representation method with causal effect transfer according to claim 2 is characterized in that: The variational autoencoder is constructed as follows: Determine the loss function of the variational autoencoder, expressed as: L VAE =E q(z|x) [logp(x|z)]-KL(q(z|x)||p(z)) Where, L VAE is the loss of the variational autoencoder; q(z|x), p(x|z) and p(z) are the latent variable, the reconstructed distribution and the prior distribution respectively; x is the observed data, z is the latent endogenous variable; KL is the KL divergence function; E q(z|x) [logp(x|z)] is the pixel-based cross entropy loss; Using adversarial similarity loss L ren Replace the pixel-based cross entropy loss E q(z|x) [logp(x|z)], adversarial similarity loss L ren It is expressed as: L ren =-E q(z|x) [logp(Dis b (x)|z)] Where, Dis b (x) is represented by the feature vector mined by the discriminator in each layer, satisfying the Gaussian distribution is the image generated by the decoder, through which the loss of the discriminator is designed as Where z is encoded by the encoder Enc(x), N is a Gaussian distribution, Dis b (x) is the discriminative loss of the b-th layer, is the discriminant loss of the output result, I is the unit matrix, Dis(x) is the total loss of all discriminant layers, and Dec(z) is the decoding result of the latent variable; use and To update the gradient in the decoder: In the formula, λ is the weight parameter, Θ Dec is the gradient of the decoder, Update the gradient derivative of the decoder; Given observation data x and the corresponding latent endogenous variables z, z's latent exogenous variables ε and supervised labels l, the log-likelihood log pΘ (x) as the target, introducing the variational distribution q φ (ε,z|x,l) to approximate the true posterior distribution p Θ (ε,z|x,l), where φ and Θ are the parameters of the variational distribution and the model distribution, respectively; Introduce KL divergence into log-likelihood, reorganize and simplify it, and use Bayes' theorem to convert log pΘ (x) is rewritten as the conditional probability log pΘ (x|z) and prior probability p Θ (z), thus we get the following evidence lower bound ELBO: Where D KL (q φ (ε,z|x,l)||p Θ (ε,z|l)) is the KL divergence; The variational distribution is simplified by the following formula: q φ (ε,z|x,l)=q φ (ε|x,l)f(z=(I-A T ) -1 ε) Where f is a nonlinear or linear function, A is the weighted binary adjacency matrix of the causal graph; Based on the simplified variational distribution, a new lower bound of evidence is obtained, expressed as: In the formula, yes, yes.
4. The facial feature decoupling representation method with causal effect transfer according to claim 3 is characterized in that: Integrating a structural causal model into the variational autoencoder to model causal effects to extract causal features from the latent feature distribution includes: Given the causal relationships, an adjacency matrix is learned in a supervised manner, wherein the weights of the adjacency matrix reflect the causal strength; Based on the adjacency matrix, z is determined by the following formula: z=f2((I-A T ) -1 f1(ε)) Where A is the weighted binary adjacency matrix of the causal graph, f1 and f2 are nonlinear and linear functions; z=f2((IA T ) -1 f1(ε)) can be rewritten as exist In the above example, we add a new term KL(q(z i |l)||p(l)), we get: In the formula, KL(q(z i |l)||p(l)) is the KL divergence between the prior and the latent variable of the i-th dimension; z i is the i-th dimension of z; The generation of causal features is supervised by the top i items of z.
5. The facial feature decoupling representation method with causal effect transfer according to claim 4 is characterized in that: Establishing a graph attention network, inputting the causal features into the graph attention network to transmit causal effects when decoupling, including: Through z i Intervene and use the supervised label l to obtain z without causal aggregation: from=to(Supervise([from]1,[from]2,...,[from] i ;l) i =c) In the formula, c is in the interval [-i,i] and z i There is a distance between the potential endogenous variables [z] i The number of i represents the i-th z; Based on the graph attention network GAT, the graph attention mechanism is used to reconstruct the formula by dynamically assigning weights to neighbor nodes and processing each z separately. i , concatenate the processed features and transpose them into the reconstructed vector z′: z′=transpose([GAT(z1),GAT(z2),...,GAT(z i )]) In the formula, transpose is the matrix transposition, GAT(z i ) is to apply the attention mechanism to the i-th dimension latent variable; The reconstructed vector z′ learns the weights of the causal edges in A through GAT, aggregates the influence of the parent nodes, and trains GAT to focus on the transmission of causal effects; Let z = {z i ∣1<i≤N,z i ∈R F }, where N is the number of z, F is the dimension of z, and the output z′={z i ′|1<i≤N,z i ′∈R F }Generated by GAT, combining aggregate information from neighboring nodes; Training weight matrix W∈R F×F , R is the real number domain, and the weight matrix represents the relationship between input and output features; Using the attention coefficient e ij represents the j-th dimension z of z j and the i-th dimension z i The importance between ij =a(Wz i ,Wz j ), For z i The weight matrix of For z j The weight matrix, α ij It is from z i to z j The weight of the edge is expressed as: In the formula, softmax i (e ij ) is the activation function used for normalization, exp is the exponential function with the base being a natural constant, k is the kth attention head, N i is the total number of parent nodes, e jk is the attention correlation score between nodes j and k; Using the attention mechanism on a single node to formulate α ij : Where LeakyRelu is a nonlinear activation function. is the kth dimension z of z k The weight matrix of Will As the final output feature of each node, a multi-head attention mechanism is used to enhance stability. The multi-head attention mechanism uses K independent attention heads to transform z into z′ through the following formula: In the formula, ‖ represents the concatenation operation, represents the normalized attention coefficient calculated by the kth attention head, W k is the weight matrix corresponding to the linear transformation of the input.
6. The facial feature decoupling representation method with causal effect transfer according to claim 4, characterized in that: Design loss functions, including: According to KL divergence, reconstruction loss Decoder loss and monitoring loss Establish the loss function, expressed as: In the formula, is the total loss of the model; β is the weight parameter; KL is the KL divergence, which is used to fit the standard Gaussian distribution and is expressed as: It is used to shorten the distance between the i-th dimension of z and the true label, expressed as: In the formula, z i denotes the i-th dimension of z, l denotes the label, and i is the number of attributes of interest.
7. The facial feature decoupling representation method with causal effect transfer according to claim 1, characterized in that: After designing a loss function, using a variational autoencoder, a discriminator, a causal structure model, and a graph attention network as a causal representation model, and training the causal representation model using the loss function, the method further includes: The identifiability of the causal representation model is analyzed.
8. The facial feature decoupling representation method with causal effect transfer according to claim 7, characterized in that: The identifiability of the causal representation model is analyzed, including: Let ~ be a binary relation on parameter Θ and define: Where Θ(E, D, A, T, λ) is the model parameter, E is the encoder, D is the decoder, A is the weighted binary adjacency matrix of the causal graph, and T is the time step; are the encoding output results of the corresponding parameters, B1 is the parameter matrix, B2 is the bias vector, b2 and b1 are scalars, and x is a single observation data. is a linear mapping, is the inverse mapping transformation, X is the entire data set; A number Θ is identifiable if the following conditions are met: The set {x∈X|φ ξ (x)=0} is zero, where φ ξ Indicates p Θ (ε,z|l)=p ε (ε)p Θ (z|l) medium density p ξ The characteristic function of The function D as the decoder is differentiable, and the Jacobian matrix of the decoder function maintains full rank; Sufficient Statistics T i,s (z i ) is not zero at all points, T i,s (z i ) represents the i-th dimension variable z i The sth statistic of ; Additional observations satisfy l i ≠0.
9. A facial feature decoupling representation device with causal effect transmission, characterized in that: The device comprises: A feature distribution acquisition module is configured to establish a variational autoencoder, and in response to input face data, learn a potential feature distribution from a latent space based on the variational autoencoder; a causal feature extraction module configured to integrate a structural causal model into the variational autoencoder to model causal effects to extract causal features from the latent feature distribution; A causal feature decoupling module is configured to establish a graph attention network, and transmit causal effects when the causal features are input into the graph attention network for decoupling; The model training module is configured to design a loss function, use a variational autoencoder, a discriminator, a causal structure model, and a graph attention network as a causal representation model, and use the loss function to train the causal representation model. 10 . A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, perform the method according to claim 1 .
Citation Information
Patent Citations
Face key point identity and expression decoupling method and device
CN115050087A
Cloud API service quality prediction method based on three-dimensional tensor high-order feature interaction
CN115809721A
Causal decoupling representation learning method based on variational auto-encoder
CN118711225A
Medical equipment maintenance data monitoring method based on big data
CN119252453A
Signal coding using potential feature prediction
CN119317957A
Cited By
Method, device and equipment for extracting polymerizable causal information of medical image and medium
CN120047792A
Industrial time sequence event analysis method and device based on causal regularization and medium
CN120654104A
Causal decoupling method and device based on multi-scale noise and adversarial supervision
CN121415085A
Interaction quantity causal contribution prediction method and device, and storage medium
CN121685017A