Automatic Image Caption Generation Method Based on Causal Reasoning

Through the non-aligned feature Transformer encoder and intervention Transformer decoder in the CIIC framework, combined with the intervention object detector and the causal intervention module, the visual and linguistic confusion factors are eliminated, the performance of the image description model is improved, and more realistic image titles are generated.

CN115239944BActive Publication Date: 2025-07-08CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210661517.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-13
Publication Date
2025-07-08
Estimated Expiration
2042-06-13

AI Technical Summary

Technical Problem

The existing Transformer-based image description model easily learns dataset bias and false correlations caused by visual and language confusion factors when processing image subtitles, and the existing causal reasoning methods fail to effectively eliminate visual feature confusion factors in the encoder.

Method used

Using an automatic image title generation method based on causal reasoning, by constructing a CIIC framework, using non-aligned feature Transformer encoder and intervention Transformer decoder, combining an intervention object detector and a causal intervention module, eliminate visual and linguistic confusion factors and achieve causal reasoning.

Benefits of technology

The performance of the image description model is significantly improved, effectively eliminating visual and linguistic confusion factors, and generating more realistic image titles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115239944B_ABST
    Figure CN115239944B_ABST
Patent Text Reader

Abstract

The present invention discloses a causal inference image caption generation method based on a causal graph, which is applicable to use in image captions. A causal inference method image caption CIIC framework based on a detailed causal graph is constructed, including a non-aligned feature Transformer encoder and an intervention-based Transformer decoder. The non-aligned feature Transformer decoder includes a FASTERR-CNN connected in sequence, an intervention-based object detector IOD, and a standard Transformer encoder; the intervention-based Transformer decoder is composed of inserting a causal intervention CI module after the feed-forward neural network layer module of the standard Transformer decoder; the intervention-based object detector IOD and the intervention-based Transformer decoder ITD jointly control the visual confounding factor and the text confounding factor to first encode the input image and then decode it. Through backdoor adjustment, confounding can be eliminated, effectively solving the problem of entangled visual features in the encoded image in traditional image description, and having strong robustness in image description.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for automatically generating image captions, and particularly to a causal inference image caption generation method based on a causal graph for use in image captions. Background Art

[0002] Existing image caption methods generally follow an encoder-decoder architecture, where features extracted from an image by a CNN are fed into an RNN (usually based on LSTM) to generate corresponding sentences. Since RNN-based models are limited by their sequential structure, convolutional language models have explored alternatives to traditional RNNs. Therefore, being essentially different from convolutional operations, the new Transformer-based caption models have achieved considerable results based on the paradigm of multi-head attention.

[0003] However, most Transformer-based image description models may still learn dataset biases brought about by hidden confounding factors. How to address the language confusion in image captions caused by dataset biases induced by vision and visual information still exists and has not been explored. In terms of visual performance, most models adopt pre-trained detectors, and these models ignore the problem of entangled visual features in images. In terms of model structure improvement, most current Transformer-based image descriptions seem to have two elusive confounding factors: visual confusion and language confusion, which usually lead to biases during training, spurious correlations during testing, and reduce the generalization ability of the model. Therefore, a new method is needed to address the spurious correlations and dataset biases brought about by the confounding factors that these image description models may learn.

[0004] Nowadays, there is a trend to introduce causal inference into different deep learning models. These efforts have made it possible to endow deep neural networks with the ability to learn causal effects. Causal effects have significantly improved the performance of many computer vision (CV) and natural language processing (NLP) models, including image classification, image semantic segmentation, visual feature representation, visual dialogue, image captioning, and dialogue generation. Existing research has analyzed the spurious correlations between visual features and captions through causality and proposed a framework for a deconstruction method for image captions (DIC) to address confounding factors, but there are still some limitations. In their causal graphs, the overall dataset is considered a confounding factor, which is a difficult-to-determine stratification and must be eliminated through complex front-door adjustment based on additional mediators. DIC focuses on eliminating the confounding factors in the decoder while ignoring the confounding factors of visual features in the encoder, resulting in a serious degradation in performance. Summary of the Invention

[0005] Technical problem: Aiming at the deficiencies of the above technologies, an automatic image caption generation method based on causal reasoning is proposed, which can simultaneously handle visual and language confounding factors in the sentence generation process, display a more detailed causal graph, and significantly improve the performance of the Transformer-based image description model.

[0006] Technical solution: To achieve the above technical purpose, the automatic image caption generation method based on causal reasoning and Transformer of the present invention is characterized in that: a causal inference method image caption CIIC framework based on a detailed causal graph is constructed, and the causal graph includes visual confounding factors and text confounding factors;

[0007] The causal inference method image caption CIIC framework includes a non-aligned feature Transformer encoder UFT and an intervention Transformer decoder ITD connected in sequence. The non-aligned feature Transformer decoder includes a FASTER R-CNN, an intervention object detector IOD, and a standard Transformer encoder connected in sequence; the intervention Transformer decoder is composed of inserting a causal intervention CI module after the feed-forward neural network layer module of the standard Transformer decoder; the intervention object detector IOD and the intervention Transformer decoder ITD jointly control the visual confounding factor and the text confounding factor to first encode the input image and then decode it;

[0008] Among them, the non-aligned feature Transformer encoder UFT first sends the de-confounded visual features extracted by the IOD and the bottom-up features extracted from the same image into two linear layers to map and generate Q, K, V vectors, which are integrated through self-attention and cross-attention, and then the AddNorm operation and the feed-forward propagation operation in the traditional Transformer are performed. The obtained output is transmitted to the next encoding block, with a total of L blocks, that is, the encoding is stacked L times; the input of the intervention Transformer decoder ITD is the currently generated sentence part, which undergoes cross-attention with the final output of the encoding end through the position embedding and the mask layer, performs the AddNorm operation and the feed-forward propagation operation, eliminates visual and language confusion in the decoding process through the causal intervention CI module, and then performs the AddNorm operation. Similarly, the decoding is repeated L times to obtain the final predicted output; the causal intervention CI module combines the fused visual and language features h2 with the expectations of the visual confounding factor D1 and the language confounding factor D2.

[0009] The Interventional Object Detector (IOD) separates region-based visual features by eliminating visual confounding factors: the features of the region of interest are separated by the Interventional Object Perceiver and then combined with the bottom-up features of the Faster Region-based Convolutional Neural Network (FASTER R-CNN) as the input to the Transformer encoder; the IOD integrates causal reasoning into the image features extracted by the Faster R-CNN to address the visual confounding in the traditional pre-trained models, thereby obtaining a region-based non-entangled representation; the results generated in the decoding stage are input into the Interventional Transformer Decoder (ITD), which introduces causal intervention into the Transformer decoder used in traditional image caption generation to alleviate the visual and language confounding in the decoding process.

[0010] By simultaneously establishing visual and language concepts through the encoder and decoder, the unobserved confounding factors between the IOD and the ITD are alleviated, visual and language confounding is eliminated, the spurious correlations occurring in visual feature representation and caption generation are effectively eliminated, and finally more realistic image captions are generated.

[0011] The specific steps are as follows:

[0012] The image for which the caption is to be generated is passed through the Faster R-CNN to extract image features, and the IOD is used to eliminate the visually confounding regional features in the image features.

[0013] Specifically, since the Faster R-CNN object detector uses the likelihood estimation method P(Y|X) as the training objective of the classifier, it leads to spurious correlations caused by the confounding factor Z.

[0014] P(Y|X) = ∑ z P(Y|X, Z = z)P(Z = z|X)

[0015] where X is the regional visual feature based on the input image, Z is the visual confounding factor of the image, and Y is the class label.

[0016] Therefore, causal reasoning intervention P(Y|do(X)) is used as the new classifier for object detection, where the do operator do(·) serves to cut the link Z→X. Since actual training requires sampling to estimate P(Y|do(X)), the training time is too long. Therefore, by applying the Normalized Weighted Geometric Mean (NWGM) approximation, the class probability output by the IOD is:

[0017]

[0018] where concat represents matrix concatenation. is the i-th class label, is the probability output that the x of the pre-trained classifier belongs to the class; x represents the regional features and y in the specific input image i c represents the feature corresponding to this region, and X and Y represent the random variables of x and y i c and x and y i c are represented as specific sample values;

[0019] approximate the confusion factor therein as a fixed confusion factor dictionary n represents the class size in the dataset, and z i represents the average RoI feature of the i-th one. Each RoI feature is pre-trained by FASTER R-CNN. The specific IOD feature extraction method is to first check the region of interest RoI on the feature map, use the faster region convolutional neural network FASTER R-CNN to extract the region of interest RoI on the feature map, and use the features of each region of interest RoI to predict the bounding box y B and the class probability output label y with surrounding visual confusion factor interference C , according to the class probability output label y C and the confusion dictionary Z, predict the final class label y by performing the do operator I to eliminate the interference of surrounding visual confusion factors;

[0020] Use the intervention object detector IOD to extract the deconfused object features from the candidate regions of all regions of interest RoI as the features of IOD. Since the bottom-up features extracted have the discriminative ability of different object attributes, the IOD features and the bottom-up features extracted from the same image are sent to two linear layers for mapping to generate Q, K, V vectors. Among them, Q represents the query vector, K represents the vector of the correlation between the query information and other information, and V represents the vector of the information to be queried. Integrate through self-attention and cross-attention to promote the visual representation of the CIIC model; since the bottom-up features and IOD features are not aligned, a multi-view Transformer encoder, that is, the unaligned feature Transformer encoder UFT, is introduced to adjust them. Input the bottom-up features and IOD features into the UFT encoder for alignment and fusion operations:

[0021] Suppose the bottom-up features and IOD features extracted from the image are respectively and where m≠n and d1≠d2. Use the two linear layers constructed in the Transformer network to map X F and X IConvert to a common d-dimensional space and represent them separately using and respectively. Select as the main visual feature and learn the cross-attention of through the main visual feature using the following formula:

[0022]

[0023] where MultiHead(·) represents the multi-head attention function of the standard Transformer, is the corresponding feature on . Similarly, establish a multi-head attention model for :

[0024]

[0025] Find the multi-head self-attention in the three , that is, Q, K, and V all come from . Therefore, note that all have the same shape, and then use the residual standard layer AddNorm to fuse and encapsulate . The fused feature information F is as follows:

[0026]

[0027] where LayerNorm represents layer normalization. Finally, send the fused feature information F into the FFN module, which is the feed-forward neural network in the Transformer, to generate the encoded result of the UFT;

[0028] To alleviate the spurious correlation between the participating visual features and the words with corresponding meanings, construct a standard Transformer decoder structure. Integrate the causal intervention module CI into each Transformer decoder layer based on the standard Transformer decoder structure. Take the region-based non-entangled representation obtained in the encoder and the text as the input of the decoder, and eliminate the visual and language confusion in the decoding process through the causal intervention module CI to generate the final image caption language description.

[0029] Furthermore, a causal relationship among the participating visual feature V, visual context D1, language context D2, participating word feature h1 of the partially generated sentence, fused feature h2, and predicted word W is constructed using the Structural Causal Model (SCM). Among them, the true causal effect is V→W. The visual context D1 and language context D2 affect the visual feature V and predicted word W respectively. The language context D2 affects the visual feature V through the participating word feature h1. The word feature h1 and visual feature V jointly affect the fused feature h2, which ultimately affects the predicted word W. Specifically, the causal effect V→W means that the participating visual feature leads to the generation of the corresponding word. The causal effect of D1 on V represents D1→V because when a caption model is trained, some frequently occurring visual contexts will seriously affect the participating visual feature. The causal effect D1→W means that the visual context directly affects the occurrence frequency of some relevant words in the generated description. D2→h1→V means that the participating word feature affected by the language context guides the participating visual feature through multi-head cross-attention. h1→h2, V→h2, and h2→W mean that the decoder fuses the visual feature and language feature and infers the next predicted word W using the fused feature h2. When using the observational probability P(W|V, h1) without causal intervention as the training objective, due to the confounding factors D1 and D2, the description generation model may learn some spurious correlations between the visual feature V and the predicted word W. To explain the causal intervention in image caption generation, P(W|V, h1) is expressed as:

[0030]

[0031] Among them, the confounding factors D1 and D2 usually introduce observational biases through P(d1|V) and P(d2|h1). Using the causal intervention P(W|do(V), do(h1)) instead of the traditional image caption training objective P(W|V, h1) can eliminate the causal effect of D1 on the visual feature V and the causal effect of D2 on the participating word feature h1, thus blocking the two backdoor paths V←D1→W and h1←D2→W and eliminating the spurious associations. Assuming that the confounding factors D1 and D2 can be stratified respectively, P(W|do(V), do(h1)) can be adjusted according to the backdoor as:

[0032]

[0033] Therefore, according to the adjustments in the formula, the image generation description model (P(W|do(V), do(h1)), which is the predicted output probability of the model) is forced to learn the true causal effect: V→W instead of the spurious associations caused by the visual confounding factor D1 and the language confounding factor D2; since both D1 and D2 are unobserved and outside the scope of the image generation description objective, approximate visual confounding factor dictionaries D1 and language confounding factor dictionaries D2 need to be constructed. The visual matrix is constructed by setting each entry in the image visual features as the average RoI feature of the objects in the categories of each image classification dataset where c is the number of classes in the training dataset, d v represents the dimension of each RoI feature. At the same time, d e dimensional word embeddings from a predefined vocabulary are used to construct the semantic space, N is the length of the vocabulary, and d e is the word feature dimension; then the description model is trained to learn two linear projections to transform the visual matrix V r and the word embeddings W e into D1 and D2 respectively through the formulas: D1 = V r P v , D2 = W e P w and then calculate using the NWGM approximation method:

[0034] P(W|do(V), do(h1)) ≈ Softmax{g(h2, E D1 [D1], E D2 [D2])},

[0035] where g(·) represents the fully connected layer, and by setting D1 and D2 to condition on the fused feature h2 to increase the representational ability of the interventionist Transformer decoder; do(h1)) means eliminating the participating word features affected by the language context, P(W|do(V)) represents the probability of predicting the generated word after eliminating the visual confounding features, and P(W|do(V), do(h1)) represents the predicted output probability after eliminating both the language context confounding and visual confounding features

[0036] Furthermore, the non - aligned feature Transformer encoder includes FASTER R - CNN, the interventionist object detector IOD, and a standard Transformer encoder including a multi - head attention layer, a residual standard layer, and a feed - forward neural network layer

[0037] The intervention-based Transformer decoder is formed by inserting a causal intervention CI module after the feed-forward neural network layer module of the standard Transformer decoder. The standard Transformer decoder includes a masked attention layer, a multi-head attention layer, a residual standard layer, and a feed-forward neural network layer;

[0038] Among them, the part composed of the multi-head attention layer, the residual standard layer, and the feed-forward neural network layer of the non-aligned feature Transformer encoder is stacked L times; the part composed of the masked attention layer, the multi-head attention layer, the residual standard layer, the feed-forward neural network layer, and the causal intervention CI module of the intervention-based Transformer decoder is stacked L times;

[0039] Both the Transformer decoder and the Transformer encoder include a multi-head attention layer, a residual standard layer, and a feed-forward neural network layer. The intervention-based Transformer decoder passes through the visual dictionary D1 and the language dictionary D2. The causal intervention module CI combines the fused feature h2 with the expectations of the visual confounding factor D1 and the language confounding factor D2 to predict the next word at each time step. At the beginning of the prediction, there is a start token as the text input, and then at each time step, the previously generated word is used as the text input; that is, through backdoor adjustment, cutting off the link of the confounding factor effectively eliminates the unobserved confounding factor to achieve causal intervention;

[0040] For the trained non-aligned feature Transformer encoder and the intervention-based Transformer decoder, first, the input image is used to extract bottom-up features through FASTER R-CNN, and the intervention-based object detector IOD extracts de-confounded object features from the RoI candidate regions. The UFT encoder takes the bottom-up features and the IOD features as inputs for alignment and fusion operations. The intervention-based Transformer decoder takes the integrated visual features as inputs and combines the input word information at each time step. The output of the last decoder layer is then projected into an N-dimensional space by a linear embedding layer, where N is the size of the vocabulary; finally, a softmax operation is used to predict the probability of words in the vocabulary to generate the final predicted word. That is, during training, each time step word comes from the true annotated sentence, and during the final prediction, the input is the output word of the previous time step.

[0041] Furthermore, pre-train the causal inference method image captioning CIIC framework:

[0042] First, pre-train using word-level cross-entropy. The training set contains images and corresponding descriptive sentences, and the loss function is:

[0043]

[0044] where θ are all the parameters of the causal inference method image captioning CIIC framework model, including weights and biases, w* 1:T is the target true sequence. The non-differentiable metric of the model is optimized through reinforcement learning (RL), and a variant of self-supervised sequence training (SCST) is adopted for the beam search sampling sequence, minimizing the negative expected score:

[0045]

[0046] where the reward r(·) is the CIDEr-D score;

[0047] The trained causal inference method image captioning CIIC framework is tested: Use beam search to generate sentences word by word in sequence. The trained model inputs the image to be recognized, and then the image passes through a series of processes and inputs the decoder. In the first decoding step, the top k candidates are considered. Generate k second words for these k first words. Considering the obtained scores, select the top k [first word, second word] combinations. For these k second words, select k third words and select the top k [first word, second word, third word] combinations. Repeat each decoding step. After k sequences, select the sequence with the best comprehensive score to obtain the sequence with the highest probability in the last beam.

[0048] Beneficial effects:

[0049] 1) This method adopts a new Transformer-based image captioning architecture CIIC from the perspective of causal relationship, seamlessly integrating causal intervention into object detection and description generation to jointly alleviate the confounding effect. On the one hand, the proposed IOD effectively untangles visual feature entanglement and promotes the deconfounding of image captioning. On the other hand, the proposed ITD adopts causal intervention to simultaneously handle visual and language confounding factors in the sentence generation process;

[0050] 2) This method decomposes the confounding factors into visual and text confounding factors and shows a more detailed causal graph.

[0051] 3) This method can significantly improve the performance of the Transformer-based image captioning model and achieve the current best image captioning performance in the single-model setting of the MS-COCO dataset. Description of the Drawings

[0052] Figure 1 is the image captioning framework diagram used in the method for automatically generating image captions based on causal inference of the present invention.

[0053] Figure 2 is the structural schematic diagram of the intervention-based object detector used in the method of the present invention.

[0054] Figure 3 This is the causal intervention schematic diagram in the image description of the present invention. Detailed implementation manners

[0055] The present invention will be further described below with reference to the accompanying drawings of the specification:

[0056] As Figure 1 shown, for the method for automatically generating an image caption based on causal reasoning of the present invention, first, confounding factors are divided, and the existing causal graphs are divided into two categories: visual confounding and text confounding. The causal inference method for image captioning (CIIC) framework structure based on causal graphs: The intervention object detector (IOD) and the intervention Transformer decoder (ITD) jointly face the two types of confoundings. The IOD integrates causal reasoning into FASTER R-CNN to cope with visual confounding, aiming to obtain a region-based disentangled representation. The ITD eliminates visual and language confoundings in the Transformer decoder stage. First, the region-of-interest features are separated by the intervention object perceptron (IOD), and then combined with the bottom-up feature of FASTER R-CNN as the input of the Transformer encoder. In CIIC, we propose a causal intervention module to cope with visual and language confoundings in word prediction. Our CIIC can effectively eliminate the spurious correlations occurring in visual feature representation and caption generation to obtain a more realistic image caption.

[0057] For the intervention object detector (IOD), FASTER R-CNN is used as the visual backbone to extract the region of interest (RoI) on the feature map. Each RoI feature is used to separately predict the class probability output y C and the bounding box y B . According to the class probability output y C and the confounding dictionary Z, we execute the do operator to predict the final class label y i .

[0058] The region-of-interest features are separated by the intervention object perceptron, and then combined with the bottom-up feature of FASTER R-CNN as the input of the Transformer encoder. In CIIC, a causal intervention module Casual Intervention is proposed to cope with visual and language confoundings in word prediction. Figure 1The symbol "L×" in the figure indicates that the encoding blocks (including the multi-head attention layer, the residual standard layer, and the feed-forward neural network layer) and the decoding blocks (including the masked attention layer, the multi-head attention layer, the residual standard layer, the feed-forward neural network layer, and the causal intervention CI module) in the dashed box are stacked L times. CIIC can effectively eliminate the spurious correlations that occur in visual feature representation and caption generation to obtain more realistic image captions.

[0059] Specifically:

[0060] The images for which captions are to be generated are respectively passed through FASTER R-CNN to extract image features, and the region features that eliminate visual confusion proposed by the interventionist object detector (IOD) are obtained.

[0061] The specific method used for the interventionist object detector is as follows:

[0062] Traditional object detectors, such as FASTER R-CNN, basically use the likelihood estimation method P(Y|X) as the training objective of the classifier, resulting in spurious correlations caused by the confounding factor Z.

[0063] P(Y|X) = ∑ z P(Y|X, Z = z)P(Z = z|X)

[0064] where X is the regional visual feature based on the input image, Z is the visual confounding factor of the image, and Y is the class label.

[0065] We propose to use causal inference intervention P(Y|do(X)) as the new classifier for object detection, where the do operator do(·) serves to cut the link Z → X. Since actual training requires time-consuming and laborious sampling to estimate P(Y|do(X)), which would make the training time prohibitive, we approximate it by applying the normalized weighted geometric mean (NWGM):

[0066] (The class probability output by the interventionist object detector)

[0067] where concat represents matrix concatenation, is the i-th class label, is the probability output that x of the pre-trained classifier belongs to the class; x represents the regional feature in the specific input image and y i c represents the feature corresponding to that region. X and Y represent the random variables of x and y i c and x and y i c are represented as specific sample values.

[0068] Approximate the confounding factors therein as a fixed confounding factor dictionary n represents the class size in the dataset, and z i represents the average RoI feature of the i-th one. Each RoI feature is pre-trained by Faster R-CNN. The specific structure of the IOD feature extractor is as Figure 2 , where first use Faster R-CNN (Faster Region Convolutional Neural Network) to extract the region of interest RoI on the feature map, such as "the upper body part of the child in blue clothes", and use each RoI feature to predict the class probability output label y C (with the interference of surrounding visual confounding factors) and the bounding box y B , according to the class probability output label y C and the confounding dictionary Z, predict the final class label y by performing the do operator I (eliminating the interference of surrounding visual confounding factors);

[0069] Step 3: Use the IOD extractor to extract the de-confounded object features from all RoI candidate regions, that is, the IOD features. Considering that the bottom-up features extracted have the discriminative ability of different object attributes, send the IOD features and the bottom-up features extracted from the same image into two linear layers to map and generate Q, K, V vectors. Among them, Q represents the query vector, K represents the vector of the correlation between the query information and other information, and V represents the vector of the information to be queried. Integrate through self-attention and cross-attention to promote the visual representation of the CIIC model; since the bottom-up features and the IOD features are not aligned, a multi-view Transformer encoder, that is, the unaligned feature Transformer encoder UFT, is introduced to adjust them. The UFT encoder takes the unaligned visual features (referring to the bottom-up features and the IOD features) as inputs and performs alignment and fusion operations simultaneously:

[0070] Let and respectively represent the bottom-up features and the IOD features extracted from the image, where m≠n and d1≠d2. Use two linear layers (constructed in the Transformer network) to transform X F and X I into a common d-dimensional space, represented by and respectively. Select as the main visual feature, and use the following formula to learn the cross-attention of :

[0071]

[0072] Among them, MultiHead(·) represents the multi-head attention function of the standard Transformer, is the corresponding feature on Similarly, the multi-head attention model for is as follows:

[0073]

[0074] Note that all have the same shape (the above formula finds the multi-head self-attention in three , that is, Q, K, and V all come from ), and then it is encapsulated with AddNorm (residual standard layer). The fused feature information F is as follows:

[0075]

[0076] where LayerNorm represents layer normalization. Finally, the fused feature information F is fed into the FFN module (feed-forward neural network in the Transformer) to generate the encoding result of the UFT;

[0077] Step 4: To alleviate the spurious correlations between the participating visual features and their corresponding words, a Transformer-based decoder structure is constructed. The Transformer-based decoder structure integrates a causal intervention module into each Transformer decoder layer to address the visual and language confusions in image captioning. As Figure 1 shown, a causal intervention module is introduced into the traditional Transformer decoder. The region-based non-entangled representation obtained in the encoder and the text are used as the input to the decoder. The visual and language confusions in the decoding process are eliminated through the causal intervention module to generate the final image caption.

[0078] Causal intervention in image captioning: As Figure 3 shown, by cutting off the two links D2→h1 and D1→V respectively, the backdoor paths V←h1←D2→W and V←D1→W are blocked to capture the true causal effect V→W.

[0079] Use the SCM (Structural Causal Model) to construct the causal relationships among the participating visual feature V, visual context D1, language context D2, participating word feature h1 of the partial generated sentence, fused feature h2, and predicted word W:

[0080] Specifically, the causal effect V→W means that the participating visual features lead to the generation of corresponding words. The causal effect of D1 on V represents D1→V because when a caption model is trained, some frequently occurring visual contexts will seriously affect the participating visual features. The causal effect D1→W means that the visual context directly affects the occurrence frequency of some relevant words in the generated description. D2→h1→V means that the participating word features affected by the language context guide the participating visual features through multi-head cross-attention. h1→h2, V→h2, and h2→W mean that the decoder fuses the visual features and language features and infers the next predicted word W using the fused feature h2. When using the observation probability P(W|V,h1) (the observation probability without causal intervention) as the training objective, due to the confounding factors D1 and D2, the description generation model may learn some spurious correlations between the visual feature V and the predicted word W. To explain the causal intervention in image caption generation, P(W|V,h1) is expressed as:

[0081]

[0082] Among them, the confounding factors usually introduce observation biases through P(d1|V) and P(d2|h1). Using the causal intervention P(W|do(V),do(h1)) instead of the traditional image caption training objective P(W|V,h1) eliminates the causal effect of D1 on the visual feature V and the causal effect of D2 on the participating word feature h1. In this way, the two backdoor paths V←D1→W and h1←D2→W are blocked, and the spurious associations are eliminated. Assuming that the confounding factors D1 and D2 can be stratified respectively, P(W|do(V),do(h1)) can be adjusted according to the backdoor as:

[0083]

[0084] Therefore, according to the adjustment in the formula, the image caption generation model (P(W|do(V),do(h1)) is the predicted output probability of the model) is forced to learn the true causal effect: V→W instead of the spurious associations caused by the visual confounding factor D1 and the language confounding factor D2. Since both D1 and D2 are unobserved and outside the objective of image caption generation, approximate visual confounding factor dictionaries D1 and language confounding factor dictionaries D2 need to be constructed (obtained by linearly projecting the visual features and word embeddings, D1 and D2 are in bold italics and different from D1 and D2). The visual matrix is constructed by setting each entry in the image visual features as the average RoI feature of the objects in each class (the classes in the image classification dataset). where c is the number of classes in the training dataset (a commonly used standard dataset), d v represents the dimension of each RoI feature. At the same time, d e dimensional word embeddings in the predefined vocabulary are used Construct a semantic space, where N is the length of the vocabulary and d e is the dimension of word features; then train a description model to learn two linear projections for the visual matrix V r and the word embedding W e through the formula: D1 = V r P v , D2 = W e P w to be transformed into D1 and D2 respectively, and use the NWGM approximation method to calculate:

[0085] P(W|do(V), do(h1)) ≈ Softmax{g(h2, E D1 [D1], E D2 [D2])},

[0086] where g(·) represents a fully connected layer, and by setting D1 and D2 to fuse the features h2 as a condition to increase the representation ability of ITD;

[0087] Transformer decoder architecture: The Transformer decoder architecture is as Figure 1 shown, where the non-aligned feature Transformer (UFT) encoder consists of FASTER R-CNN (Faster Region Convolutional Neural Network), the intervention object detector IOD, and a standard Transformer encoder (including a multi-head attention layer, a residual standard layer, and a feed-forward neural network layer). The intervention Transformer decoder is formed by inserting a causal intervention CI module after the feed-forward neural network layer module of the standard Transformer decoder. (Corresponding Figure 1The light red layer above the middle dashed box decoder), where the symbol "L×" indicates that the encoding blocks (including the multi-head attention layer, residual standard layer, and feed-forward neural network layer) and decoding blocks (including the masked attention layer, multi-head attention layer, residual standard layer, feed-forward neural network layer, and causal intervention CI module) in the dashed box are stacked L times. Similar to the Transformer encoder, the general Transformer decoder also includes a multi-head attention layer, a residual standard layer, and a feed-forward neural network layer, except that it has an additional masked self-attention layer. It is composed of L identical decoder layers stacked in sequence. We innovate on the basis of the classic Transformer decoder and insert a CI (causal intervention) module after the FFN (feed-forward neural network layer) module of the Transformer decoder. Through the visual dictionary D1 and the language dictionary D2, the CI module combines the fused feature h2 with the expectations of the visual confounding factor D1 and the language confounding factor D2 to predict the next word at each time step (at the beginning of prediction, there is a start token as the text input, and at each subsequent time step, the previously generated word is used as the text input). That is, it actually achieves causal intervention through backdoor adjustment (effectively eliminating the unobserved confounding factors by cutting off the links of the confounding factors). The trained model first extracts the bottom-up features of the input image through FASTER R-CNN, and the intervention-based object detector IOD extracts the deconfounded object features in the RoI candidate regions. The UFT encoder takes the unaligned visual features (referring to the bottom-up features and IOD features) as input and performs alignment and fusion operations. The intervention-based Transformer decoder (ITD) takes the integrated visual features as input and combines the input word information at each time step. The output of the last decoder layer is then projected into an N-dimensional space by the linear embedding layer, where N is the size of the vocabulary. Finally, a softmax operation is used to predict the probabilities of the words in the vocabulary to generate the final predicted word (during training, the word at each time step comes from the true annotated sentence, and during the final prediction, the input is the output word of the previous time step).

[0088] For the pre-training of the CIIC model, this model first uses word-level cross-entropy for pre-training (the training set contains images and corresponding description sentences), and the loss function is:

[0089]

[0090] where θ are all the parameters (including weights and biases) of the model (CIIC model), w* 1:T is the target true sequence. The non-differentiable metric of the model is optimized through reinforcement learning (RL), and a variant of self-supervised sequence training (SCST) is adopted for the beam search sampling sequence, minimizing the negative expected score:

[0091]

[0092] where the reward r(·) is the CIDEr-D score;

[0093] In the test phase, beam search is used to generate sentences word by word in order. The trained model takes the image to be recognized as input. Then the image goes through a series of processes and is input into the decoder. At the first decoding step, the top k candidates are considered. k second words are generated for these k first words. Considering the scores obtained, the top k [first word, second word] combinations are selected. For these k second words, k third words are selected, and the top k [first word, second word, third word] combinations are chosen. Each decoding step is repeated. After k sequences are completed, the sequence with the best comprehensive score is selected, and the sequence with the highest probability in the last beam is obtained.

[0094] To represent the image features, first, the proposed IOD is trained on the MSCOCO dataset to extract 1024-dimensional IOD features of the top 100 objects with the highest confidence. Then, the pre-trained Up-Down model is used to extract 2048-dimensional bottom-up features of the detected objects. Finally, these two features are linearly projected into a model with an input dimension d = 512, and they are input into the UFT encoder. In the experiment, one-hot vectors and pre-trained GloVe word embeddings are used to represent words respectively. Both are linearly projected onto the 512-dimensional input vector of the ITD. To represent the word positions in the sentence, the input vectors and their sine positional encodings are added before the first decoding layer. Words out of the vocabulary are represented as all-zero vectors. The Adam optimizer is used in the training phase with a batch size of 10 and a beam size of 5. A step decay schedule with a warm-up equal to 20000 is used to change the learning rate. All models are first trained with cross-entropy loss for 30 epochs, and then further optimized with CIDEr reward for another 30 epochs with a learning rate of 5×10 -6 of. In the inference phase, we adopt a beam search strategy with a beam size of 3.

[0095] In summary, a new Transformer-based image captioning architecture CIIC is proposed from the perspective of causality, which seamlessly combines causal intervention into object detection and caption generation to jointly alleviate the confounding effect. On the one hand, the proposed IOD effectively untangles visual feature entanglement and promotes the deconfounding of image captioning. On the other hand, the proposed ITD adopts causal intervention to simultaneously handle visual and language confounding factors in the sentence generation process. Experimental results show that this method significantly improves the performance of the Transformer-based image captioning model and achieves a new advanced level in the single-model structure of the MS-COCO dataset.

Claims

1. An automatic image caption generation method based on causal reasoning and Transformer, characterized in that: Construct a causal inference method image captioning CIIC framework based on a detailed causal graph, where the causal graph includes visual confounding factors and text confounding factors; The causal inference method image captioning CIIC framework includes a non-aligned feature Transformer encoder UFT and an intervention Transformer decoder ITD connected in sequence. The non-aligned feature Transformer decoder includes a FASTER R-CNN, an intervention object detector IOD, and a standard Transformer encoder connected in sequence. The intervention Transformer decoder is composed of inserting a causal intervention CI module after the feed-forward neural network layer module of the standard Transformer decoder. The intervention object detector IOD and the intervention Transformer decoder ITD jointly control the visual confounding factors and text confounding factors to first encode the input image and then decode it; Among them, the non-aligned feature Transformer encoder UFT first sends the de-confounded visual features extracted by the IOD and the bottom-up features extracted from the same image into two linear layers to map and generate Q, K, V vectors, integrates them through self-attention and cross-attention, and then performs the AddNorm operation and feed-forward propagation operation in the traditional Transformer. The obtained output is passed to the next encoding block, with a total of L blocks, that is, the encoding is stacked L times; The input of the intervention Transformer decoder ITD is the currently generated sentence part. It performs cross-attention with the final output of the encoding end through the position embedding and masking layer, performs the AddNorm operation and feed-forward propagation operation, eliminates visual and language confusion in the decoding process through the causal intervention CI module, and then performs the AddNorm operation. Similarly, the decoding is repeated L times to obtain the final predicted output; the causal intervention CI module combines the fused visual and language features h2 with the expectations of the visual confounding factor D1 and the language confounding factor D2; The intervention object detector IOD separates region-based visual features by eliminating visual confounding factors: separates the region-of-interest features through the intervention object perceptron, and then combines them with the bottom-up features of the faster region convolutional neural network FASTER R-CNN as the input of the Transformer encoder; the intervention object detector IOD integrates causal inference into the image features extracted by the FASTERR-CNN to cope with the visual confusion extracted by the traditional pre-trained model, so as to obtain a region-based non-entangled representation; inputs the result generated in the decoding stage into the intervention Transformer decoder ITD, and introduces causal intervention into the Transformer decoder used in traditional image caption generation to reduce visual and language confusion in the decoding process; Establish visual and language concepts simultaneously through an encoder and a decoder, mitigate the unobserved confounding factors between the Interventionist Object Detector (IOD) and the Interventionist Transformer Decoder (ITD), eliminate visual and language confusion, effectively eliminate the spurious correlations occurring in visual feature representation and caption generation, and finally generate more realistic image captions.

2. The method for automatically generating image captions based on causal reasoning and Transformer according to claim 1, wherein: The specific steps are as follows: Pass the image for which the caption is to be generated through FASTERR-CNN respectively to extract image features, and use the Interventionist Object Detector (IOD) to eliminate the region features with visual confusion in the image features; Specifically, since the Faster R-CNN object detector uses the likelihood estimation method P(Y|X) as the training objective of the classifier, it results in spurious correlations caused by the confounding factor Z, P(Y|X) = ∑ z P(Y|X, Z = z)P(Z = z|X) where X is the regional visual feature based on the input image, Z is the visual confounding factor of the image, and Y is the class label; Therefore, use causal inference intervention P(Y|do(X)) as the new classifier for object detection. The do operator do(·) plays the role of cutting the link Z→X. Since actual training requires sampling to estimate P(Y|do(X)), the training time is too long. Therefore, by applying the normalized weighted geometric mean (NWGM) approximation, the class probability output by the Interventionist Object Detector is: where concat represents matrix concatenation, is the i-th class label, is the probability output that the x of the pre-trained classifier belongs to the class; x represents the regional features in the specifically input image and y i c represents the feature corresponding to this region, X and Y represent the random variables of x and and x and are represented as specific sample values; Approximate the confounding factors therein as a fixed confounding factor dictionary n represents the class size in the dataset, z i represents the average RoI feature of the i-th one. Each RoI feature is pre-trained by Faster R-CNN. The specific IOD feature extractor method is to first check the region of interest RoI on the feature map, use the faster region convolutional neural network Faster R-CNN to extract the region of interest RoI on the feature map, and use the features of each region of interest RoI to predict the bounding box y B and the class probability output label y with interference from surrounding visual confounding factors C , and according to the class probability output label y C and the confounding dictionary Z, predict the final class label y by performing the do operator I to eliminate the interference of surrounding visual confounding factors; Use the Interventionist Object Detector (IOD) to extract the de-confounded object features from the candidate regions of all Regions of Interest (RoIs) as the features of IOD. Since the bottom-up features extracted have the discriminative ability of different object attributes, send the IOD features and the bottom-up features extracted from the same image into two linear layers for mapping to generate Q, K, V vectors. Here, Q represents the query vector, K represents the vector of the correlation between the query information and other information, and V represents the vector of the information to be queried. Integrate them through self-attention and cross-attention to promote the visual representation of the CIIC model; Since the bottom-up features and the IOD features are not aligned, a multi-view Transformer encoder, namely the Unaligned Feature Transformer Encoder (UFT), is introduced to adjust them. Input the bottom-up features and the IOD features into the UFT encoder for alignment and fusion operations: Let the bottom-up features and IOD features extracted from the image be and respectively, where m≠n and d1≠d2. Use two linear layers constructed in the Transformer network to transform X F and X I into a common d-dimensional space, and represent them with and respectively. Select as the main visual feature, and use the following formula to learn the cross-attention of : where MultiHead(·) represents the multi-head attention function of the standard Transformer, is the corresponding feature on Similarly, a multi-head attention model is established for : Three find the multi-head self-attention, that is, Q, K, and V all come from Therefore, note that all have the same shape, and then use the residual standard layer AddNorm to perform fusion and encapsulation. The fused feature information F is shown as follows: where LayerNorm represents layer normalization. Finally, send the fused feature information F into the FFN module, which is the feed-forward neural network in the Transformer, to generate the encoding result of UFT; To alleviate the spurious correlation between the participating visual features and the words with corresponding meanings, construct a standard Transformer decoder structure. Integrate the causal intervention module CI into each Transformer decoder layer based on the standard Transformer decoder structure. Take the region-based non-entangled representation obtained in the encoder and the text as the input of the decoder, and eliminate the visual and language confusion in the decoding process through the causal intervention module CI to generate the final language description of the image caption.

3. The method for automatically generating image captions based on causal reasoning and Transformer according to claim 2, wherein: Construct the causal relationships among the participating visual features V, visual context D1, language context D2, participating word features h1 of the partial generated sentence, fused feature h2, and predicted word W using the Structural Causal Model (SCM): Among them, the true causal effect is V→W. The visual context D1 and language context D2 affect the visual feature V and predicted word W respectively. The language context D2 affects the visual feature V through the participating word feature h1. The word feature h1 and visual feature V jointly affect the fused feature h2, which ultimately affects the predicted word W. Specifically, the causal effect V→W means that the participating visual features lead to the generation of corresponding words. The causal effect of D1 on V represents D1→V because when a caption model is trained, some frequently occurring visual contexts will seriously affect the participating visual features. The causal effect D1→W means that the visual context directly affects the occurrence frequency of some relevant words in the generated description. D2→h1→V means that the participating word features affected by the language context guide the participating visual features through multi-head cross-attention. h1→h2, V→h2, and h2→W mean that the decoder fuses the visual and language features and infers the next predicted word W using the fused feature h2. When using the observational probability P(W|V,h1) without causal intervention as the training objective, due to the confounding factors D1 and D2, the description generation model may learn some spurious correlations between the visual feature V and the predicted word W. Use causal intervention P(W|do(V),do(h1)) instead of the traditional image caption training objective P(W|V,h1) to eliminate the causal effect of D1 on the visual feature V and the causal effect of D2 on the participating word feature h1, thus blocking the two backdoor paths V←D1→W and h1←D2→W and eliminating the spurious associations. Let the confounding factors D1 and D2 be stratified respectively, then P(W|do(V),do(h1)) is adjusted according to the backdoor as follows: Therefore, according to the adjustment in the formula, the image generation description model (P(W|do(V), do(h1)), which is the predicted output probability of the model) is forced to learn the true causal effect: V→W instead of the spurious associations caused by the visual confounding factor D1 and the linguistic confounding factor D2; since both D1 and D2 are unobserved and outside the scope of the image generation description objective, approximate visual confounding factor dictionary D1 and linguistic confounding factor dictionary D2 need to be constructed. The visual matrix is constructed by setting each entry in the image visual features to the average RoI feature of the objects in the category of each image classification dataset where c is the number of classes in the training dataset, and d v represents the dimension of each RoI feature. At the same time, d e -dimensional word embeddings from a predefined vocabulary are used to construct the semantic space, N is the length of the vocabulary, and d e is the word feature dimension; then the description model is trained to learn two linear projections to transform the visual matrix V r and the word embeddings W e into D1 and D2 respectively through the formulas: D1 = V r P v , D2 = W e P w and the NWGM approximation method is used to calculate: P(W|do(V),do(h1))≈Softmax{g(h2,E D1 [D1],E D2 [D2])}, where g(·) represents the fully connected layer, and conditioned on setting D1 and D2 to fuse the feature h2 to increase the representational ability of the interventionist Transformer decoder; do(h1)) represents eliminating the participating word features affected by the language context, P(W|do(V)) represents the probability of predicting and generating a word after eliminating the visual confusion features, and P(W|do(V),do(h1)) represents the predicted output probability after eliminating the language context confusion and visual confusion features.

4. The automatic image caption generation method based on causal reasoning and Transformer according to claim 2, characterized in that The non-aligned feature Transformer encoder includes FASTER R-CNN, the intervention-based object detector IOD, and a standard Transformer encoder including a multi-head attention layer, a residual standard layer, and a feed-forward neural network layer; The intervention-based Transformer decoder is formed by inserting a causal intervention CI module after the feed-forward neural network layer module of the standard Transformer decoder. The standard Transformer decoder includes a masked attention layer, a multi-head attention layer, a residual standard layer, and a feed-forward neural network layer; Among them, the part composed of the multi-head attention layer, the residual standard layer, and the feed-forward neural network layer of the non-aligned feature Transformer encoder is stacked L times; the part composed of the masked attention layer, the multi-head attention layer, the residual standard layer, and the feed-forward neural network layer and the causal intervention CI module of the intervention-based Transformer decoder is stacked L times; Both the Transformer decoder and the Transformer encoder include multi-head attention layers, residual standard layers, and feed-forward neural network layers. The intervention-based Transformer decoder uses the visual dictionary D1 and the language dictionary D2. The causal intervention module CI combines the fused feature h2 with the expectations of the visual confounding factor D1 and the language confounding factor D2 to predict the next word at each time step. At the beginning of the prediction, a start token is used as the text input, and at each subsequent time step, the previously generated word is used as the text input; that is, through backdoor adjustment, cutting off the link of the confounding factor effectively eliminates the unobserved confounding factor to achieve causal intervention; For the trained misaligned feature Transformer encoder and the intervention-based Transformer decoder, first, the input image is used to extract bottom-up features through FASTER R-CNN and de-confounded object features from the RoI candidate regions extracted by the intervention-based object detector IOD. The UFT encoder takes the bottom-up features and IOD features as inputs for alignment and fusion operations. The intervention-based Transformer decoder takes the integrated visual features as inputs and combines the input word information at each time step. The output of the last decoder layer is then projected into an N-dimensional space by a linear embedding layer, where N is the size of the vocabulary; finally, a softmax operation is used to predict the probability of words in the vocabulary to generate the final predicted word, that is, during training, each time step word comes from the true annotated sentence, and during the final prediction, the input is the output word of the previous time step.

5. The method for automatically generating image captions based on causal reasoning and Transformer according to claim 4, wherein Pre-train the causal inference method image captioning CIIC framework: First, pre-train using word-level cross-entropy. The training set contains images and corresponding descriptive sentences, and the loss function is: where θ are all the parameters of the causal inference method image captioning CIIC framework model, including weights and biases, w* 1:T is the target true sequence. The non-differentiable metric of the model is optimized through reinforcement learning (RL), and a variant of self-supervised sequence training (SCST) is adopted for the beam search sampling sequence to minimize the negative expected score: where the reward r(·) is the CIDEr-D score; Test the trained causal inference method image captioning CIIC framework: Use beam search to generate sentences word by word in sequence. The trained model takes the image to be recognized as input, and then the image is input into the decoder through a series of processes. At the first decoding step, consider the top k candidates; generate k second words for these k first words, consider the obtained scores, and select the top k [first word, second word] combinations. For these k second words, select k third words and select the top k [first word, second word, third word] combinations. Repeat each decoding step. After ending k sequences, select the sequence with the best comprehensive score to obtain the sequence with the highest probability in the last beam.

Citation Information

Patent Citations

  • Image description generation method based on relation between external knowledge and targets

    CN113609326A

  • A reference preposition description-based image description generation method

    CN113946706A