Method, device and equipment for extracting polymerizable causal information of medical image and medium

By introducing measurement learning and graph attention mechanisms into the causal representation learning framework of medical images, SCM based on GAT is constructed, which solves the problem of complex causal relationship recognition in medical images, and achieves more accurate disease feature recognition and individualized diagnostic support.

CN120047792AActive Publication Date: 2025-05-27YANSHAN UNIV

Patent Information

Application Number
CN202510112408.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the complex nonlinear causal relationships in medical images, which makes it difficult to remove confounding factors introduced by data noise and deviation, affecting the accuracy of disease characteristic recognition.

Method used

Introducing metric learning and graph attention mechanisms in the traditional causal representation learning framework, constructing a structural causal model (SCM) based on graph attention network (GAT) to improve the quality of causal representation and infer causal relationships in images.

Benefits of technology

By capturing complex nonlinear causal relationships in medical images, it effectively removes confounding factors introduced by data noise and bias, improves the independence of extracted characterization, thereby more accurately identifying disease characteristics and providing accurate support for individualized diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047792A_ABST
    Figure CN120047792A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device and equipment for extracting polymerizable causal information of a medical image and a medium. Relates to the technical field of representation learning, causal inference and depth generation models. The method comprises the following steps: constructing a causal representation learning framework; the causal representation learning framework comprises an encoder, a structural causal model, a decoder and a discriminator, the encoder is used for encoding an input medical image into a low-dimensional exogenous variable, the structural causal model takes the low-dimensional exogenous variable as input and generates a causal representation based on the low-dimensional exogenous variable, the decoder is used for intervening and reconstructing the causal representation, and the discriminator is used for discriminating the causal representation. The discriminator is used for adversarial training; and establishing a model training loss function, training the causal representation learning framework based on the model training loss function, and extracting the polymerizable causal information in the medical image by using the trained causal representation learning framework. The causal graph identified by the method can clearly describe the causal relationship among different features, and more transparent decision support is provided for clinicians.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of representation learning, causal inference, and deep generative models, and particularly relates to a method, apparatus, device, and medium for extracting aggregable causal information from medical images. Background Art

[0002] With the rapid development of medical imaging technologies (such as X-rays, CT, MRI, and ultrasound), medical images play an increasingly important role in disease diagnosis, treatment planning, and efficacy evaluation. Traditional deep learning models require a large amount of labeled data for training, but the cost of labeling medical image data is high, and there are problems of scarce and inconsistent labeling. Moreover, deep learning models are often "black-box models" and it is difficult to provide interpretive information of practical value to medical professionals. Causal representation learning provides a new paradigm for solving the above problems. It explicitly models the causal relationships between variables, extracts causally meaningful representations, thereby enhancing the generalization ability and interpretability of the model.

[0003] By learning causal representations from medical images, it is possible to model the causal relationships in the data, effectively remove confounding factors introduced by data noise and biases, and thus more accurately identify disease characteristics. This causally driven analysis method helps to reduce misdiagnosis and missed diagnosis phenomena caused by subjective factors or equipment differences, and improve the reliability and consistency of diagnosis. At the same time, the medical field has relatively high requirements for model interpretability. Causal graphs can clearly reveal the causal relationships between variables, helping clinicians understand the basis behind the model output, thereby increasing trust in artificial intelligence decisions.

[0004] Current causal representation learning algorithms are mainly divided into two categories: VAE-based algorithms and GAN-based algorithms. In 2018, Kocaoglu et al. proposed the CausalGAN model, which is a causal implicit generative model for learning a given causal graph. They used the causal graph to represent the dependency structure between binary labels, and in the generation process, the binary labels can be used to control relevant attributes in face images, such as attributes like smiling and squinting. The limitation of CausalGAN is that it directly takes real causal information as a conditional attribute and inputs it into the generator, and the model does not adjust the dimensional alignment between the latent variables and the latent generation factors. Therefore, CausalGAN does not perform representation learning in the decoupling stage. In 2021, the CausalVAE model proposed by Yang et al. introduced a structural causal model (SCM) and binary label information of image attributes in the decoupling stage, input the encoded representation into the SCM for further modulation, output a representation vector with causal information, and performed non-linear processing on the latent vector and scene reconstruction in the generation process, achieving causally controllable decoupled representation learning.

[0005] Metric learning is an important branch of machine learning, which aims to learn effective distance or similarity metrics between data. Deep metric learning can perform non-linear mappings on input features and has been widely applied in the field of computer vision. Graph Attention Networks (GATs) are a novel type of convolutional neural network that includes masked self-attention layers and can process graph-structured data. The graph attention layers used in GAT can perform computations efficiently (without expensive matrix operations and can process all nodes in the graph in parallel), allowing different importance to be assigned to different nodes implicitly within the neighborhood while handling neighborhoods of different sizes.

[0006] In view of the above research background, how to capture the complex non-linear causal relationships in medical image data, thereby effectively removing confounding factors introduced by data noise and bias, increasing the independence of the extracted representations, and thus more accurately identifying disease features to provide precise support for individualized diagnosis and treatment is a technical problem that urgently needs to be solved at present. Summary of the Invention

[0007] This application provides a method, device, equipment and medium for extracting aggregable causal information from medical images. Considering the complex non-linear distribution in medical images, metric learning and graph attention mechanisms are introduced into the traditional causal representation learning framework to construct a GAT-based SCM, thereby improving the quality of causal representations and inferring causal relationships in images.

[0008] In a first aspect, this application provides a method for extracting aggregable causal information from medical images, including:

[0009] Construct a causal representation learning framework; wherein, the causal representation learning framework includes an encoder, a structural causal model, a decoder and a discriminator. The encoder is used to encode the input medical image into a low-dimensional exogenous variable. The structural causal model takes the low-dimensional exogenous variable as input and generates a causal representation based on the low-dimensional exogenous variable. The decoder is used to intervene and reconstruct the causal representation, and the discriminator is used for adversarial training;

[0010] Establish a model training loss function and train the causal representation learning framework based on the model training loss function to implement the extraction of aggregable causal information in medical images with the trained causal representation learning framework.

[0011] In a possible design, the structural causal model is expressed as:

[0012]

[0013] In the formula, z represents the causal representation, T represents matrix transpose, f and h represent non-linear functions, and g represents a parameterized graph attention network. Represents low-dimensional exogenous variables, Represents the real number field, d represents the number of causal attributes, and A represents the prior causal graph, Represents the causal structure matrix;

[0014] Optimize the graph attention network based on the gradient-based continuous update strategy to obtain the causal structure matrix

[0015] Input the causal structure matrix Into the structural causal model to generate causal representations.

[0016] In a possible design, optimize the graph attention network based on the gradient-based continuous update strategy to obtain the causal structure matrix Includes:

[0017] Using i and j as any two nodes in the causal structure matrix Learn the causal edge between i and j by multiplying the connection features with the shared attention mechanism a(·), generating the attention score e ij ; The attention score is expressed as:

[0018] e ij = a(W∈ i , W∈ j )

[0019] In the formula, Represents a shared weight matrix;

[0020] For node i, only calculate nodes j ∈ Pa i , where Pa i Represents the parent node, and use the softmax function to normalize them. The calculation process is as follows:

[0021]

[0022] In the formula, Represents the normalized attention coefficient generated by the h-th attention head, and H represents the number of attention heads;

[0023] Construct through the following formula

[0024]

[0025] In the formula, Represents the causal strength of the i-th row and j-th column in the causal structure matrix, τ > 0, is the temperature parameter that controls the result approaching 0 or 1; And Represents samples independently drawn from the Gumbel(0,1) distribution, and σ(·) represents the logistic sigmoid function;

[0026] Iteratively train the graph attention network to obtain the causal structure matrix

[0027] In a possible design, input the causal structure matrix into the structural causal model to generate causal representations, including:

[0028] Use an encoder based on metric learning to generate the ideal causal representation of the triplet samples; the triplet network fits the feature distribution in the latent space by measuring the distances between the anchor and the positive / negative samples, so as to obtain a causal representation that is not affected by distribution shift. The triplet loss is expressed as:

[0029]

[0030] In the formula, (z a , z p , z n ) represents a triplet, including an anchor sample z a , a positive sample z p and a negative sample z n . max represents the maximum value function, D(z a , z n ) represents the distance between the anchor sample and the negative sample, D(z a , z p ) represents the distance between the anchor sample and the positive sample, D(z i , z j ) = ||f(z i ) - f(z j )|| 2 represents the Euclidean distance between two vectors, and m represents the margin.

[0031] In a possible design, intervene and reconstruct the causal representation through the following formula:

[0032]

[0033] In the formula, f -1 represents the inverse mapping of f after intervention, represents a set of intervention operations, represents the causal representation after intervention, represents the causal structure matrix after intervention.

[0034] In a possible design, establish the model training loss function through the following method;

[0035] Learn the approximate variational distribution q of the true distribution p of the causal representation z to fit the actual posterior distribution;

[0036] Obtain a training data set and the joint distribution Maximize the evidence lower bound, which is expressed as:

[0037]

[0038] where ELBO represents the evidence lower bound; represents taking the expectation over all samples under the data distribution to ensure that the optimization of the model's objective function can reflect the overall performance of the data set; represents the expected value calculated under the approximate posterior distribution q φ (z|x,u) of the latent variable z; p θ (x|z) represents the distribution of generating observed data from the latent variable, parameterized by θ; q φ (∈|x,u) represents the probability distribution of ∈ given the observed data x and the conditional variable u; p ∈ (∈) represents the prior distribution of the noise variable ∈; q φ (z|x,u) represents the approximate posterior distribution of the latent variable z; p θ (z|u) represents the prior distribution of the latent variable z;

[0039] and represents the KL divergence term, which measures the difference between the two distributions separated by || respectively; x represents high-dimensional medical image data; u represents the corresponding label of the image;

[0040] Model learning of the causal structure is achieved by setting a prior on the representation generated by the structural causal model and calculating the KL divergence, using the following loss function:

[0041]

[0042] where

[0043]

[0044] where α and β are regularization hyperparameters, represents the expected log-likelihood of the reconstruction ability, represents constraining the distribution of the latent variable through prior knowledge; represents optimizing the distance of features in the latent space; represents the overall loss function of the VAE framework; D KL (q φ (∈|x,u)||p ∈ (∈)) represents q φ (∈|x,u) and q φThe KL divergence of (∈|x,u), where q φ (∈|x,u) is a variational approximation that models the posterior distribution of the latent variable ∈, p ∈ (∈); D KL (q φ (z|x,u)||p θ (z|u)) represents the KL divergence between the distributions q φ (z|x,u) and p θ (z|u), where q φ (z|x,u) is the approximate posterior distribution learned by the encoder network in the VAE, representing the probability distribution of the latent variable z given the observed data x and the condition u; x a represents the anchor sample; x p represents a sample that is semantically close to the anchor sample, i.e., a positive sample; x n represents a sample that is semantically different from the anchor sample, i.e., a negative sample; m represents the margin; D(x a ,x p ) represents the distance between the anchor sample and the positive sample in the feature space, generally calculated using the Euclidean distance; D(x a ,x n ) represents the distance between the anchor sample and the negative sample in the feature space;

[0045] The discriminator optimizes a binary classifier using the following loss function:

[0046]

[0047] In the formula, represents the discriminator loss; Dis(x) represents the output of the discriminator, and the value range is 0 ≤ Dis(x) ≤ 1, represents the data generated by the decoder;

[0048] Express the model training loss function as:

[0049]

[0050] In a second aspect, the present application provides an apparatus for extracting aggregable causal information of medical images, and the apparatus includes:

[0051] A framework construction module configured to construct a causal representation learning framework; wherein, the causal representation learning framework includes an encoder, a structural causal model, a decoder, and a discriminator. The encoder is used to encode the input medical image into a low-dimensional exogenous variable. The structural causal model takes the low-dimensional exogenous variable as an input and generates a causal representation based on the low-dimensional exogenous variable. The decoder is used to intervene and reconstruct the causal representation. The discriminator is used for adversarial training;

[0052] A framework training module, configured to establish a model training loss function, train the causal representation learning framework based on the model training loss function, and use the trained causal representation learning framework to extract aggregable causal information in medical images.

[0053] In a third aspect, an embodiment of the present application provides an electronic device, including: at least one processor and a memory; the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the method for extracting aggregable causal information from medical images as described in the first aspect and various possible designs of the first aspect above.

[0054] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method for extracting aggregable causal information from medical images as described in the first aspect and various possible designs of the first aspect above is implemented.

[0055] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for extracting aggregable causal information from medical images as described in the first aspect and various possible designs of the first aspect above is implemented.

[0056] The method, device, equipment and medium for extracting aggregable causal information from medical images provided by the present application can effectively quantify detection uncertainty, significantly improve the reliability and adaptability of the model in different scenarios, and have stronger robustness. The main advantages are summarized as follows:

[0057] 1. Based on the encoder-decoder framework, the present application introduces a decoupling method of metric learning at the encoder end to separate representations, increase the distance between negative samples and anchors, and reduce the distance between positive samples and anchors. A structural causal model is introduced between the encoder and the decoder, and a graph attention network (GAT) is used to identify the causal relationships between variables, and a causal layer is used to transform exogenous variables into endogenous variables, so as to capture complex non-linear causal relationships in medical image data, effectively removing confounding factors introduced by data noise and bias, increasing the independence of the extracted representations, and thus more accurately identifying disease features and providing precise support for individualized diagnosis and treatment.

[0058] 2. To learn causal representations from medical image data, remove confounding factors introduced by data noise and biases, more accurately identify disease features, and improve the interpretability of the model, a causal representation learning model based on a triple network and graph attention mechanism is proposed. Inject a graph attention network (GAT) into a structural causal model (SCM), obtain a causal structure matrix by aggregating causal information of context nodes and continuously updating the GAT using a gradient-based strategy. Use a triple loss to reduce the distance between similar samples in the latent space distribution to obtain more effective causal representations. This application can learn causal relationships in medical images with high accuracy; this application uses a triple network to generate more effective causal representations by reducing the distance between similar samples in the latent space distribution; the causal graph identified by this application can clearly describe the causal relationships between different features, providing more transparent decision support for clinicians. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0060] Figure 1 Schematic structural diagram of the causal representation learning framework provided by an embodiment of the present application;

[0061] Figure 2 Flowchart of a method for extracting aggregable causal information of a medical image provided by an embodiment of the present application;

[0062] Figure 3 Schematic structural diagram of the device for extracting aggregable causal information of a medical image provided by an embodiment of the present application.

[0063] Through the above accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0065] In the technical solution of this application, the processing of information such as financial data or user data, including collection, storage, use, processing, transmission, provision, and disclosure, complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0066] It should be noted that in the embodiments of this application, some existing solutions in the industry such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.

[0067] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the above technical problems. These several specific embodiments can be combined with each other. For the same or similar concepts or processes, they may not be elaborated in some embodiments. The following will describe the embodiments of this application with reference to the accompanying drawings.

[0068] The embodiment of this application provides a method for extracting aggregable causal information from medical images. The method includes: constructing a causal representation learning framework; establishing a model training loss function, and training the causal representation learning framework based on the model training loss function to extract aggregable causal information from medical images. The aggregable causal information extracted in this embodiment is a causal graph, which can provide more transparent decision support for clinicians.

[0069] Figure 1 FIG. [FIGURE NUMBER] is a schematic structural diagram of the causal representation learning framework provided by the embodiment of this application. The causal representation learning framework includes an encoder, a decoder, a discriminator, and a structural causal model SCM.

[0070] The encoder is used to encode an image with a non-linear feature distribution into a low-dimensional exogenous variable.

[0071] The SCM is used to define the causal relationship structure in the data. It consists of a graph attention network (GAT) and a set of structural equations containing non-linear transformations, and realizes the conversion from exogenous variables to endogenous variables.

[0072] The decoder is used to reconstruct the original image or the intervened image from the low-dimensional representation, and helps to verify the effectiveness of the potential causal representation.

[0073] The discriminator is used for adversarial training, enabling the discriminator to measure samples, thereby guiding the decoder to generate more realistic intervened samples.

[0074] As Figure 1 shown, the encoder accepts a triple (x a , x p , x n Please note that the figure number in "FIG. [FIGURE NUMBER] is a schematic structural diagram of the causal representation learning framework provided by the embodiment of this application." should be filled with the actual figure number. Also, the specific content of " a ", " p ", and " n " may need to be adjusted according to the actual context.) As inputs, it respectively includes anchor samples, positive samples with characteristics similar to those of the anchor samples, and negative samples with characteristics dissimilar to those of the anchor samples. The encoder encodes the inputs into independent exogenous variables, and their prior distributions are assumed to be standard Gaussian distributions. Then, the SCM converts them into causal representations (z a , z p , z n ) through causal layers, and calculates the triplet loss for updating the model. The discriminator calculates the discriminator loss through the generated images and real images of the inputs, and is used to update the discriminator for adversarial training.

[0075] To learn causal representations with high generalization ability from medical images, first, the graph attention network (GAT) is incorporated into the SCM to aggregate causal information from context nodes and encode the information embedded in the GAT parameters. Second, the triplet loss is used to reduce the distance between similar samples, adapt to the real distribution, and obtain effective causal representations. In addition, the triplet loss can promote the decoder to generate more realistic intervention samples for evaluating the generated causal representations, which is achieved by jointly training the variational autoencoder (VAE) and the generative adversarial network (GAN).

[0076] In a specific embodiment, Figure 2 This is a flowchart of a method for extracting aggregable causal information from medical images provided by an embodiment of the present application. As Figure 2 shown, the method for extracting aggregable causal information from medical images includes the following steps S10 to S40, which are introduced in detail as follows.

[0077] Step S10, construct a causal representation learning framework based on GAT and metric learning.

[0078] For the latent factors of data x, if the encoder E learns a decoupled representation (i.e., low-dimensional exogenous variables) about z, that is, for each i = 1,..., m, there exists a corresponding function [E(x)] i = z i .

[0079] On the contrary, most previous methods use independent priors to describe z. Assume a more general case where there are causal relationships between latent factors. A structural causal model SCM represents a tuple where z represents a set of endogenous variables ∈ represents a set of exogenous variables The structural equation is a function that determines z, where z i = f i (Pa i , ∈ i ), Pa iDenote the set of parent nodes of z in the causal structure matrix i . The SCM uses deterministic functions f and exogenous variables ∈ to construct a causal model, and uses an encoder E to generate independent exogenous variables from x.

[0080] Let X = {x n | 1 ≤ n ≤ N} be the data input set containing d latent factors, and U = {u n | 1 ≤ n ≤ N} be the corresponding set of attribute labels. By generating a triple dataset, we can obtain triples (x a , x p , x n ), representing the anchor, positive sample, and negative sample. For an image x, first use the encoder network E (e.g., ResNet) to generate the exogenous variable represents the real number field, and d represents the number of causal attributes,. This embodiment uses SCM and GAT to learn semantic association information from exogenous variables to obtain the matrix Therefore, set the SCM as:

[0081]

[0082] where g(·) represents the parameterized GAT, f i (·) and h(·) represent non-linear functions, A represents the prior causal structure, represents the causal structure matrix.

[0083] Optimize the multi-layer GAT through a gradient-based continuous update strategy to obtain the causal structure matrix Then input into the structural equation to generate the causal representation z.

[0084] The tasks of the causal representation learning framework include three parts:

[0085] 1) Learn the causal structure matrix

[0086] according to the supervision label u (u ∈ U); a , x p , x n ) to generate causal representations;

[0087] 3) Intervene in the causal representation z, use an independent causal mechanism to intervene and reconstruct the causal representation, and generate the intervened image through a decoder, that is, the aggregable causal information extracted from the medical image.

[0088] Step S20, learn the causal structure matrix

[0089] In this embodiment, the GAT based on aggregated causal information is used to learn the causal structure matrix. Considering the permutation invariance and message passing mechanism of the GAT, it can process nodes without loss of performance, effectively transmit causal effects, and obtain the dynamic causal strength on the edges. Let i and j be any two nodes in ij . The causal edge between i and j is learned by multiplying the connection features with the shared attention mechanism a(·), generating the attention score e ij ; e ij reflects the probability of the causal edge. By stacking multiple layers of networks, e

[0090] e ij = a(W∈ i , W∈ j )

[0091] where represents a shared weight matrix. Subsequently, the prior causal structure is injected into the causal attention mechanism by performing masked attention. For node i, only calculate node j ∈ Pa i , where Pa i represents the parent node. To balance the influence of different parent nodes, the softmax function is used to normalize them; the process is as follows:

[0092]

[0093] represents the normalized attention coefficient generated by the h-th attention head, and H represents the number of attention heads; the causal edge is identified by using Gumbel-Sigmoid, which can approximate the sample (threshold = 0.5) and ensure that the parameters are suitable for gradient optimization. The constructed can be expressed as follows:

[0094]

[0095] where represents the causal strength of the i-th row and j-th column in the causal structure matrix, τ > 0 is the temperature parameter that controls the result to be close to 0 or 1; '1' here represents the existence of a causal edge, and represent samples independently drawn from the Gumbel(0,1) distribution, and σ(·) represents the logistic sigmoid function; the GAT is iteratively trained using the above formula to obtain the causal structure matrix

[0096] Step S30, generate the causal representation z.

[0097] Generate an ideal causal representation of triplet samples using a VAE based on metric learning; the triplet network fits the feature distribution in the latent space by measuring the distances between the anchor and positive / negative samples, thereby obtaining a causal representation that is not affected by distribution shifts. For an anchor sample, we select samples with similar causal attributes as its positive samples, otherwise as negative samples; the triplet samples share the encoder weights, and the triplet loss is defined as follows:

[0098]

[0099] According to the above formula, the triplet (z a ,z p ,z n ) is generated by the SCM, which contains an anchor sample z a , a positive sample z p and a negative sample z n . The goal of the above formula is to ensure that the distance D(z a ,z n ) between the anchor sample and the negative sample exceeds the distance D(z a ,z p ) between the anchor sample and the positive sample, and the gap between the two is at least a margin m; where D(z i ,z j ) = ||f(z i ) - f(z j )|| 2 represents the Euclidean distance between two vectors; doing so aims to effectively utilize the triplet loss to enhance the performance of the VAE in learning causal representations.

[0100] Step S40, reconstruct the causal representation of the intervention

[0101] According to the causal ladder theory, "intervention" involves a change in the variable distribution, and the intervention only affects those nodes that belong to the same module as the intervened node.

[0102] By using "intervention" to operate and reconstruct variables, and using the SCM to simulate the propagation of causal effects, the change in the causal structure matrix can represent the causal relationship after the intervention;

[0103]

[0104] where, f -1 represents the inverse function of f, that is, the inverse mapping of f after the intervention. The above equation simplifies the complexity of the non-linear causal relationship into a linear form through f -1 . represents a set of intervention operations, Represents the causal representation after intervention, represents the causal structure matrix after intervention. The intervened causal node will change the effect node, otherwise it will not. This is because in the SCM, causal information only flows from the causal node to the effect node.

[0105] Step S50, define the model training loss function.

[0106] By learning the approximate variational distribution q of the true distribution p of the latent variable z to fit the actual posterior distribution; given the training dataset and the joint distribution attempt to maximize the evidence lower bound (ELBO), which is the following variational lower bound:

[0107]

[0108] where θ and φ are learnable parameters; considering the correlation between ∈ and z, the equation can be written in the following form:

[0109]

[0110] In the formula, ELBO represents the evidence lower bound; represents taking the expectation over all samples under the data distribution to ensure that the optimization of the model's objective function can reflect the overall performance of the dataset; represents the expected value calculated under the approximate posterior distribution q φ (z|x,u) of the latent variable z; p θ (x|z) represents the distribution of generating observed data from the latent variable, parameterized by θ; q φ (∈|x,u) represents the probability distribution of ∈ given the observed data x and the conditional variable u; p ∈ (∈) represents the prior distribution of the noise variable ∈; q φ (z|x,u) represents the approximate posterior distribution of the latent variable z; p θ (z|u) represents the prior distribution of the latent variable z;

[0111] and represent the KL divergence terms, which measure the differences between the two distributions separated by || respectively; x represents the high-dimensional medical image data; u represents the corresponding label of the image.

[0112] Model learning of the causal structure is achieved by setting a prior on the representation generated by the structural causal model and calculating the Kullback-Leibler (KL) divergence, using the following loss function:

[0113]

[0114] Among them

[0115]

[0116] Among them, α and β are regularization hyperparameters, represents the expected log-likelihood of the reconstruction ability, represents constraining the distribution of latent variables through prior knowledge; represents optimizing the distance of features in the latent space; represents the overall loss function of the VAE framework; D KL (q φ (∈|x, u)||p ∈ (∈)) represents the KL divergence between q φ (∈|x,u) and q φ (∈|x,u), where q φ (∈|x,u) is a variational approximation for modeling the posterior distribution of the latent variable ∈, and p ∈ (∈); D KL (q φ (z|x,u)||p θ (z|u)) represents the KL divergence between the distributions q φ (z|x,u) and p θ (z|u), where q φ (z|x,u) is the approximate posterior distribution learned by the encoder network in the VAE, representing the probability distribution of the latent variable z given the observed data x and the condition u; x a represents the anchor sample; x p represents the sample that is semantically close to the anchor sample, i.e., the positive sample; x n represents the sample that is semantically different from the anchor sample, i.e., the negative sample; m represents the margin; D(x a ,x p ) represents the distance between the anchor sample and the positive sample in the feature space, usually calculated using the Euclidean distance; D(x a ,x n ) represents the distance between the anchor sample and the negative sample in the feature space.

[0117] The discriminator aims to optimize a binary classifier using the following loss function to better distinguish real images and generated images, while encouraging the generator to fit the real distribution;

[0118]

[0119] Among them, Dis(x) represents the output of the discriminator, and its value range is 0 ≤ Dis(x) ≤ 1, represents the data generated by the decoder.

[0120] The model training loss function consists of the following equations:

[0121]

[0122] Wherein, is the model training loss function.

[0123] The embodiment of the present application also provides an apparatus for extracting aggregable causal information of medical images, as Figure 3 shown. The apparatus for extracting aggregable causal information of medical images includes:

[0124] A framework construction module 301, configured to construct a causal representation learning framework; wherein, the causal representation learning framework includes an encoder, a structural causal model, a decoder and a discriminator. The encoder is used to encode the input medical image into a low-dimensional exogenous variable. The structural causal model takes the low-dimensional exogenous variable as input and generates a causal representation based on the low-dimensional exogenous variable. The decoder is used to intervene and reconstruct the causal representation. The discriminator is used for adversarial training;

[0125] A framework training module 302, configured to establish a model training loss function, and train the causal representation learning framework based on the model training loss function to implement the extraction of aggregable causal information in medical images with the trained causal representation learning framework.

[0126] In some embodiments, the framework construction module is further configured to represent the structural causal model as:

[0127]

[0128] Wherein, z represents the causal representation, T represents the matrix transpose, f and h represent non-linear functions, g represents a parameterized graph attention network, represents the low-dimensional exogenous variable, represents the real number field, d represents the dimension of the causal representation, A represents the prior causal structure, represents the causal structure matrix;

[0129] Optimize the graph attention network based on the gradient-based continuous update strategy to obtain the causal structure matrix

[0130] Input the causal structure matrix into the structural causal model to generate a causal representation.

[0131] In some embodiments, the framework construction module is further configured to:

[0132] Use i and j as the causal structure matrix For any two nodes in [[]], the causal edge between i and j is learned by multiplying the connection feature by the shared attention mechanism a(·), generating an attention score e ij ; The attention score is expressed as:

[0133] e ij = a(W∈ i ,W∈ j )

[0134] In the formula, represents a shared weight matrix;

[0135] For node i, only calculate nodes j ∈ Pa i , where Pa i represents the parent node, and they are normalized using the softmax function. The calculation process is as follows:

[0136]

[0137] In the formula, represents the normalized attention coefficient generated by the h-th attention head, and H represents the number of attention heads;

[0138] Constructed through the following formula

[0139]

[0140] In the formula, represents the causal strength of the i-th row and j-th column in the causal structure matrix, τ > 0, which is a temperature parameter that controls the result to approach 0 or 1; and represent samples independently drawn from the Gumbel(0,1) distribution, and σ(·) represents the logistic sigmoid function;

[0141] Iteratively train the graph attention network to obtain the causal structure matrix

[0142] In some embodiments, the framework construction module is further configured to:

[0143] Use an encoder based on metric learning to generate an ideal causal representation of the triplet samples; The triplet network fits the feature distribution in the latent space by measuring the distance between the anchor and the positive / negative samples, thereby obtaining a causal representation that is not affected by distribution shift. The triplet loss is expressed as:

[0144]

[0145] In the formula, (z a ,z p ,z n) represents a triple, including an anchor sample z a , a positive sample z p and a negative sample z n , max represents the maximum value function, D(z a ,z n ) represents the distance between the anchor sample and the negative sample, D(z a ,z p ) represents the distance between the anchor sample and the positive sample, D(z i ,z j )=||f(z i )-f(z j )|| 2 represents the Euclidean distance between two vectors, and m represents the margin.

[0146] In some embodiments, the framework construction module is further configured to intervene and reconstruct the causal representation by the following formula:

[0147]

[0148] where, f -1 represents the inverse mapping of f after intervention, represents a set of intervention operations, represents the causal representation after intervention, represents the causal structure matrix after intervention.

[0149] In some embodiments, the framework training module is further configured to establish a model training loss function by the following method;

[0150] By learning an approximate variational distribution q of the true distribution p of the causal representation z to fit the actual posterior distribution;

[0151] Obtain the training data set and the joint distribution Maximize the evidence lower bound, and the evidence lower bound is expressed as:

[0152]

[0153] where, ELBO represents the evidence lower bound; represents taking the expectation over all samples under the data distribution to ensure that the optimization of the objective function of the model can reflect the overall performance of the data set; represents the expected value calculated under the approximate posterior distribution q φ (z|x,u) of the latent variable z; p θ (x|z) represents the distribution of generating observed data from the latent variable, parameterized as θ; q φ(∈|x,u) represents the probability distribution of ∈ given the observed data x and the conditional variable u; p ∈ (∈) represents the prior distribution of the noise variable ∈; q φ (z|x,u) represents the approximate posterior distribution of the latent variable z; p θ (z|u) represents the prior distribution of the latent variable z;

[0154] and represents the KL divergence term, which measures the difference between the two distributions separated by ||; x represents the high-dimensional medical image data; u represents the corresponding label of the image;

[0155] Model learning of the causal structure is achieved by setting priors on the representations generated by the structural causal model and calculating the KL divergence, using the following loss function:

[0156]

[0157] where

[0158]

[0159] where α and β are regularization hyperparameters, represents the expected log-likelihood of the reconstruction ability, represents constraining the distribution of the latent variable through prior knowledge; represents optimizing the distance of features in the latent space; represents the overall loss function of the VAE framework; D KL (q φ (∈|x, u)||p ∈ (∈)) represents the KL divergence between q φ (∈|x,u) and q φ (∈|x,u), where q φ (∈|x,u) is the variational approximation for modeling the posterior distribution of the latent variable ∈, p ∈ (∈); D KL (q φ (z|x,u)||p θ (z|u)) represents the KL divergence between the distributions q φ (z|x,u) and p θ (z|u), where q φ (z|x,u) is the approximate posterior distribution learned by the encoder network in the VAE, representing the probability distribution of the latent variable z given the observed data x and the condition u; x a represents the anchor sample; x p represents the sample that is semantically close to the anchor sample, i.e., the positive sample; x nrepresents samples that are semantically different from anchor samples, i.e., negative samples; m represents margin; D(x a ,x p ) represents the distance between the anchor sample and the positive sample in the feature space, which is generally calculated using the Euclidean distance; D(x a ,x n ) represents the distance between the anchor sample and the negative sample in the feature space;

[0160] The discriminator optimizes a binary classifier using the following loss function:

[0161]

[0162] In the formula, represents the discriminator loss; Dis(x) represents the output of the discriminator, and its value range is 0≤Dis(x)≤1. represents the data generated by the decoder;

[0163] The model training loss function It is expressed as:

[0164]

[0165] An embodiment of the present application provides an electronic device, which may include: a processor and a memory, wherein the processor and the memory may communicate with each other; illustratively, the processor and the memory communicate with each other via a communication bus.

[0166] The processor executes the computer execution instructions stored in the memory, so that the processor executes the scheme in the above embodiment. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit ASIC, a field programmable gate array FPGA or other programmable logic devices, discrete gates or transistor logic devices, and discrete hardware components.

[0167] The communication bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in illustration, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus. The transceiver is used to implement the communication between the database access device and other computers (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory.

[0168] The electronic device provided by the embodiments of the present application can be the terminal device in the above embodiments.

[0169] The embodiments of the present application also provide a computer-readable storage medium. Computer instructions are stored in the computer-readable storage medium. When the computer instructions run on a computer, the computer is enabled to execute the technical solutions of the method for extracting aggregable causal information of medical images in the above embodiments.

[0170] The embodiments of the present application also provide a computer program product. The computer program product includes a computer program which is stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When at least one processor executes the computer program, the technical solutions of the method for extracting aggregable causal information of medical images in the above embodiments can be implemented.

[0171] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.

[0172] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solutions of this embodiment.

[0173] In addition, in each embodiment of the present application, each functional module may be integrated in a processing unit, or each module may exist physically alone, or two or more modules may be integrated in one unit. The unit formed by the above modules may be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0174] The integrated module implemented in the form of a software functional module may be stored in a computer-readable storage medium. The above software functional module stored in a storage medium includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods in the various embodiments of the present application.

[0175] It should be understood that the above processor may be a central processing unit (CPU for short), or may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the execution of the hardware processor, or can be implemented by the combination of hardware and software modules in the processor.

[0176] The memory may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.

[0177] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.

[0178] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0179] An exemplary storage medium is coupled to a processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic control unit or a master control device.

[0180] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk or optical disk that can store program codes.

[0181] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for extracting aggregatable causal information from medical images, characterized in that: The method comprises: Constructing a causal representation learning framework; wherein the causal representation learning framework includes an encoder, a structural causal model, a decoder and a discriminator, wherein the encoder is used to encode the input medical image into a low-dimensional exogenous variable, the structural causal model takes the low-dimensional exogenous variable as input and generates a causal representation based on the low-dimensional exogenous variable, the decoder is used to intervene and reconstruct the causal representation, and the discriminator is used for adversarial training; A model training loss function is established, and the causal representation learning framework is trained based on the model training loss function, so as to realize the extraction of aggregatable causal information in medical images with the trained causal representation learning framework.

2. The method for extracting aggregatable causal information from medical images according to claim 1, characterized in that: The structural causal model is expressed as: In the formula, z represents the causal representation, T represents the matrix transpose, f and h represent nonlinear functions, and g represents the parameterized graph attention network. represents low-dimensional exogenous variables, represents the domain of real numbers, d represents the number of causal attributes, A represents the prior causal graph, Represents the causal structure matrix; Gradient-based continuous update strategy optimizes graph attention network to obtain causal structure matrix The causal structure matrix Input into the structural causal model to generate a causal representation.

3. The method for extracting aggregatable causal information from medical images according to claim 2, characterized in that: Gradient-based continuous update strategy optimizes graph attention network to obtain causal structure matrix include: With i and , as the causal structure matrix For any two nodes in , we learn the causal edge between i and j by multiplying the connection feature with the shared attention mechanism a(.) to generate an attention score e ij ; The attention score is expressed as: And ij =a(W∈ i ,W∈ j ) In the formula, represents a shared weight matrix; For node i, only node j∈Pa is calculated i , where Pa i Represents the parent nodes, and uses the softmax function to normalize them. The calculation process is as follows: In the formula, represents the normalized attention coefficient generated by the h-th attention head, and H represents the number of attention heads; Constructed by the following formula In the formula, It represents the causal strength of row i and column j in the causal structure matrix, τ>0, which is the temperature parameter that controls the result to be close to 0 or 1; and represents samples drawn independently from the Gumbel (0,1) distribution, and σ(·) represents the logistic sigmoid function; Iteratively train the graph attention network to obtain the causal structure matrix 4. The method for extracting aggregatable causal information from medical images according to claim 2, characterized in that: The causal structure matrix Input into the structural causal model to generate a causal representation, including: The encoder based on metric learning is used to generate the ideal causal representation of triple samples; the triple network fits the feature distribution in the latent space by measuring the distance between the anchor point and the positive / negative samples, thereby obtaining a causal representation that is not affected by distribution shift. The triple loss is expressed as: In the formula, (z a ,z p ,z n ) represents a triple, including an anchor sample z a , a positive sample z p and a negative sample z n , max represents the maximum value function, D(z a ,z n ) represents the distance between the anchor sample and the negative sample, D(z a ,z p ) represents the distance between the anchor sample and the positive sample, D(z i ,z j )=||f(z i )-f(z j )||2 represents the Euclidean distance between two vectors, and m represents the margin.

5. The method for extracting aggregatable causal information from medical images according to claim 2, characterized in that: Intervene and reconstruct the causal representation through the following formula: In the formula, f -1 represents the inverse mapping of f after intervention, represents a set of intervention operations, represents the causal representation after the intervention, Represents the causal structure matrix after the intervention.

6. The method for extracting aggregatable causal information from medical images according to claim 1, characterized in that: The model training loss function is established by the following method; By learning an approximate variational distribution q of the true distribution p of the causal representation z, in order to fit the actual posterior distribution; Get the training dataset and the joint distribution Maximize the evidence lower bound, which is expressed as: Where ELBO represents the lower bound of evidence; Indicates that the data is distributed Take the expectation for all samples to ensure that the optimization of the model's objective function can reflect the overall performance of the data set; represents the approximate posterior distribution q over the latent variable z φ The expected value calculated under (z|x,u); p θ (x|z) represents the distribution of observed data generated from the latent variable, parameterized by θ; q φ (∈|x,u) represents the probability distribution of ∈ given the observed data x and the conditional variable u; p ∈ (∈) represents the prior distribution of the noise variable ∈; q φ (z|x,u) represents the approximate posterior distribution of the latent variable z; p θ (z|u) represents the prior distribution of the latent variable z; and represents the KL divergence term, which measures the difference between two distributions separated by ||; x represents high-dimensional medical image data; u represents the corresponding label of the image; The model learns the causal structure by setting a prior on the representation generated by the structural causal model and computing the KL divergence, using the following loss function: in Where α and β are regularization hyperparameters, represents the expected log-likelihood of reconstruction ability, Represents constraining the distribution of latent variables through prior knowledge; Represents the distance between features in the latent space by optimizing; Represents the overall loss function of the VAE framework; Indicates q φ (∈|x,u) and q φ The KL divergence of (∈|x,u) where q φ (∈|x,u) is a variational approximation that models the posterior distribution of the latent variable ∈, p ∈ (∈) is the prior distribution of the exogenous variable ∈; D KL (q φ (z|x,u)||p θ (z|u)) represents the distribution q φ (z|x,u) and p θ The KL divergence between (z|u) where q φ (z|x,u) is the approximate posterior distribution learned by the encoder network in VAE, which represents the probability distribution of the latent variable z given the observed data x and condition u; x a represents the anchor point sample; x p represents samples that are semantically close to the anchor samples, i.e., positive samples; x n represents samples that are semantically different from anchor samples, i.e., negative samples; m represents margin; D(x a ,x p ) represents the distance between the anchor sample and the positive sample in the feature space, which is generally calculated using the Euclidean distance; D(x a ,x n ) represents the distance between the anchor sample and the negative sample in the feature space; The discriminator optimizes a binary classifier using the following loss function: In the formula, represents the discriminator loss; Dis(x) represents the output of the discriminator, and its value range is 0≤Dis(x)≤1. represents the data generated by the decoder; The model training loss function It is expressed as:

7. A device for extracting aggregatable causal information from medical images, characterized in that: The device comprises: A framework construction module is configured to construct a causal representation learning framework; wherein the causal representation learning framework includes an encoder, a structural causal model, a decoder and a discriminator, wherein the encoder is used to encode an input medical image into a low-dimensional exogenous variable, the structural causal model takes the low-dimensional exogenous variable as input and generates a causal representation based on the low-dimensional exogenous variable, the decoder is used to intervene and reconstruct the causal representation, and the discriminator is used for adversarial training; The framework training module is configured to establish a model training loss function, train the causal representation learning framework based on the model training loss function, and realize the extraction of aggregatable causal information in medical images with the trained causal representation learning framework.

8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method for extracting aggregatable causal information from a medical image according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method for extracting aggregatable causal information from medical images according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for extracting aggregatable causal information from a medical image according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Domain generalization image recognition method based on causal decoupling generation model

    CN114863213A

  • Traffic accident prediction method based on graph attention network

    CN118053095A

  • Causal decoupling representation learning method based on variational auto-encoder

    CN118711225A

  • Face feature decoupling representation method and device with causal effect transmission and medium

    CN120047983A

Cited By

  • Industrial time sequence event analysis method and device based on causal regularization and medium

    CN120654104A

  • Multi-task prediction method and device for medical focus segmentation and prognosis based on causal perception, medium, program product and terminal

    CN121190917A

  • Method, device, medium, program product and terminal for multi-task prediction of medical lesion segmentation and prognosis based on causal perception

    CN121190917B

  • Simulation model training method and device, sewage treatment method, equipment and medium

    CN121436046A

  • Simulation model training methods and devices, wastewater treatment methods, equipment and media

    CN121436046B