Conceptual bottleneck model construction method and device based on variational information bottleneck guidance

By introducing variational information bottlenecks and thinking chain technologies into the conceptual bottleneck model, more relevant concept pools are generated and image text embeddings are aligned, the problem of low interpretability of existing models on large-scale data sets is solved, and higher classification accuracy and interpretability are achieved.

CN120145183AActive Publication Date: 2025-06-13TIANJIN UNIV

Patent Information

Application Number
CN202510199665.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

Existing conceptual bottleneck models are difficult to generalize on large-scale data sets, and the automatically generated concept labels may not accurately match the real content of the input image, resulting in low model interpretation.

Method used

The conceptual bottleneck model construction method based on variational information bottleneck guidance is adopted, and the concept pool with stronger correlation is generated by introducing thinking chain technology, and the variational information bottleneck is used to deeply align the CLIP-encoded image text embeddings to screen out core conceptual bottlenecks.

Benefits of technology

It improves the interpretability and classification accuracy of the model on large-scale data sets, reduces interference from noise information, and enhances the semantic alignment ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145183A_ABST
    Figure CN120145183A_ABST
Patent Text Reader

Abstract

The invention discloses a concept bottleneck model construction method based on variational information bottleneck guidance, and the method comprises the steps: generating a basic concept pool directly related to the content of an input image through a visual language model, generating a descriptive supplementary concept pool through a large language model, and covering the diversified attribute features of a target category; on the basis of a variational information bottleneck principle, concept importance scores are calculated through a cross-modal aligned image-text embedding space, and high-relevance concepts are screened; the basic concept classifier and the supplementary concept classifier form a double-branch concept bottleneck network, and prediction results of the basic concept classifier and the supplementary concept classifier are fused through a classification fusion device; and an interpretability efficiency index optimization model is adopted to balance classification accuracy and concept interpretation efficiency. According to the method, deep alignment is carried out on image text embedding subjected to CLIP coding by utilizing the variational information bottleneck, and the classification accuracy is further improved by utilizing the supplementary concept bottleneck.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning interpretability, and particularly relates to a method and device for constructing a concept bottleneck model guided by variational information bottleneck. Background Art

[0002] Currently, with the wide application of deep neural networks in various practical applications, the need to understand the decision-making process of these "black box" models is increasing continuously. To improve the interpretability of the models, the concept bottleneck model provides an effective solution path for this challenge. The concept bottleneck model introduces a concept bottleneck layer, using human-understandable concepts as intermediate representations, and thus making the final decision prediction of the model based on these concepts. However, this method relies on high-density manual concept annotation, which is not only costly but also has great limitations in practical applications. Especially in some fields, the acquisition or definition of concept annotation may be difficult and unclear. Therefore, the concept bottleneck model was initially limited to small-scale datasets with concept annotations (such as the CUB dataset) and has long been difficult to be extended to larger-scale datasets.

[0003] In recent years, with the development of large language models and vision-language models, the research on the concept bottleneck model has been significantly promoted. Some studies attempt to use large language models to replace manual annotation and automatically generate concept bottlenecks, thereby improving the efficiency of model training. These methods usually generate a concept pool through a large language model and then perform post-processing training in combination with a vision-language model or a neuron-level interpretation tool. This progress not only eliminates the high cost of manual annotation but also enables the concept bottleneck model to be extended to large-scale datasets such as ImageNet. However, existing methods still face challenges in solving the concept prediction and information leakage problems.

[0004] First of all, the concept prediction problem means that the concept bottleneck model relies on intermediate layer concept prediction to generate the final decision result. Therefore, the accuracy of concept prediction directly affects the credibility of model interpretation. However, the concept labels automatically generated by the large language model may not accurately match the true content of the input image, and in the concept screening process, the visual concepts in the image may not be strictly aligned with the generated text concepts, resulting in deviations in the intermediate concept layer of the model. This deviation may cause the model to give misleading information when interpreting the decision, thus affecting the authenticity of the interpretation result. Therefore, how to extract concepts with strong semantic consistency from the automatically generated concept pool has become an important topic in current research.

[0005] Secondly, the problem of information leakage indicates that the predicted values at the concept layer may contain additional information unrelated to the features of the input image. Ideally, concept prediction should only be related to specific features of the input image and should not contain irrelevant content. However, studies have shown that even with a randomly generated set of concept labels, the model can still achieve a high classification accuracy, indicating that the model inadvertently encodes noise information unrelated to the task at the concept layer, thereby reducing interpretability and making the model's reasoning process difficult to conform to human understanding. Solving this problem requires optimizing the construction and screening process of the concept pool to ensure that the concept layer only transmits core information related to the task and avoids interference from noise. Summary of the Invention

[0006] The present invention provides a method and device for constructing a concept bottleneck model guided by variational information bottleneck to solve the technical problems existing in the prior art.

[0007] The technical solution adopted by the present invention to solve the technical problems existing in the prior art is as follows:

[0008] A method for constructing a concept bottleneck model guided by variational information bottleneck, comprising the following method steps:

[0009] Step 1, build a concept bottleneck model according to the following structure: make the concept bottleneck model include a vision-language model, a large language model, a variational information bottleneck model A, a variational information bottleneck model B, a text encoder A of CLIP, a text encoder B of CLIP, an image encoder of CLIP, a dot product multiplier A, a dot product multiplier B, a basic concept classifier, a supplementary concept classifier, and a classification fusion device; connect the vision-language model, the variational information bottleneck model A, and the text encoder A of CLIP in sequence; connect the large language model, the variational information bottleneck model B, and the text encoder B of CLIP in sequence, and connect the output ends of the text encoder A of CLIP and the image encoder of CLIP to the two input ends of the dot product multiplier A correspondingly; connect the output ends of the text encoder B of CLIP and the image encoder of CLIP to the two input ends of the dot product multiplier B correspondingly; connect the output end of the dot product multiplier A to the input end of the basic concept classifier, connect the output end of the dot product multiplier B to the input end of the supplementary concept classifier, and connect the output ends of the basic concept classifier and the supplementary concept classifier to the input end of the classification fusion device;

[0010] Step 2, input the image data into the vision-language model, and use the vision-language model to generate a basic concept pool directly related to the content of the input image; input the prompt information into the large language model, and use the large language model to generate a descriptive supplementary concept pool covering diverse attribute features of the target category;

[0011] Step 3, the variational information bottleneck model A is used to screen out the core basic concepts with relatively high relevance from the basic concept pool, and the variational information bottleneck model B is used to screen out the core supplementary concepts with relatively high relevance from the supplementary concept pool. The core basic concepts and the core supplementary concepts are used to construct a concept bottleneck;

[0012] Step 4, the core basic concepts and the core supplementary concepts are respectively processed through the text encoders A and B of CLIP to obtain the basic concept embeddings and the supplementary concept embeddings correspondingly; the image data is input into the image encoder of CLIP, and the image embeddings are obtained through the processing of the image encoder of CLIP; the basic concept embeddings and the image embeddings are multiplied through the multiplier A, and the multiplication result of the two is input into the basic concept classifier; the supplementary concept embeddings and the image embeddings are multiplied through the multiplier B, and the multiplication result of the two is input into the supplementary concept classifier; the outputs of the basic concept classifier and the supplementary concept classifier are fused through the classification fusion device to obtain the comprehensive classification prediction result;

[0013] Step 5, an interpretability efficiency index is used to evaluate the concept bottleneck model to balance the classification accuracy and the efficiency of concept interpretation.

[0014] Further, in Step 2, a chain of thought prompt template is set, and the vision-language model adopts the Instruct BLIP model; the Instruct BLIP model extracts various visual features from the image; the chain of thought prompt template is used to guide the Instruct BLIP model to gradually generate concepts related to the image from multiple angles and levels, and finally form a fine-grained basic concept pool;

[0015] The large language model adopts the GPT-3.5 Turbo model; the GPT-3.5 Turbo model generates diverse text descriptions according to the prompts given by the chain of thought prompt template to ensure that all comprehensive attribute features of the target category are covered.

[0016] Further, in Step 3, the variational information bottleneck models A and B are used to align the visual concepts in the text description with the image regions without relying on concept labels; the information bottleneck of the vision-language variant is used for feature attribution analysis, the importance scores of each concept in the text description and the image regions are calculated, and then the comprehensive scores of each concept are obtained according to the frequency accumulation of concept occurrences and the concept professionalism. According to the order of the comprehensive scores from high to low, the top K concepts in the basic concept pool are screened to construct the core basic concepts; the core supplementary concepts are constructed from the top V concepts in the supplementary concept pool.

[0017] Further, in Step 4, let: Ψ b (b) be the output of the basic concept classifier, Ψ c(c) is to supplement the output of the concept classifier; the classification fusion device fuses the outputs of the basic concept classifier and the supplementary concept classifier according to the following formula:

[0018]

[0019] In the formula:

[0020] represents the prediction result;

[0021] α represents the weight coefficient;

[0022] b represents the basic concept bottleneck;

[0023] c represents the supplementary concept bottleneck;

[0024] f represents the visual features encoded by the CLIP image encoder;

[0025] W represents the basic concept text features obtained by the CLIP text encoder;

[0026] U represents the supplementary concept text features obtained by the CLIP text encoder;

[0027] Ψ b () represents the basic concept linear classifier function;

[0028] Ψ c () represents the supplementary concept linear classifier function.

[0029] Furthermore, when training the basic concept classifier and the supplementary concept classifier, forward and backward propagation are performed twice in each batch to update the parameters; in the first propagation, only the basic concept classifier is optimized; in the second propagation, the parameters of the basic concept classifier are fixed, and the supplementary concept classifier is optimized; in order to avoid the influence of extreme values of the supplementary concept on the results, cosine similarity instead of projection distance is used as the concept bottleneck, and the activation value of each concept is normalized so that its mean is 0 and the standard deviation is 1 to accelerate convergence.

[0030] Furthermore, in step 5, the calculation formula of the interpretability efficiency index is as follows:

[0031]

[0032] In the formula:

[0033] IE represents the interpretability efficiency;

[0034] Acc represents the classification accuracy;

[0035] i represents the concept serial number;

[0036] j represents the category serial number;

[0037] W F represents a linear layer for mapping concepts to categories;

[0038] (W F ) ij represents the weight corresponding to the i-th concept and the j-th category in the linear layer;

[0039] O{(W F ) ij ≠0} represents the number of non-zeros in the linear layer;

[0040] R represents the number of concepts;

[0041] S represents the number of categories;

[0042] represents the average concept length.

[0043] Furthermore, in step 3, for the variational information bottleneck models A and B, using the variational approximation method, optimize the objective functions of both, use the KL divergence between the image and the text to adjust the parameters of both, and combine the pre-trained text encoders A and B of CLIP to assign importance scores to the image regions and text concepts.

[0044] Furthermore, for the variational information bottleneck models A and B, set the following objective functions:

[0045]

[0046] where:

[0047] m represents the image or text modality;

[0048] m′ represents the complement of modality m;

[0049] θ m represents the set of parameters of modality m;

[0050] represents the optimization objective under the set of parameters of modality m;

[0051] Z m represents the output of the information bottleneck for modality m;

[0052] E m′ represents the embedding of modality m′;

[0053] X m represents the input of the information bottleneck for modality m;

[0054] β represents the scaling factor;

[0055] I(Z m ,Em′ θ m ) represents the embedding E of mode m′ under the parameter set of mode m m′ The output Z of mode m with information bottleneck m The mutual information between

[0056] I(Z m ,X m θ m ) represents the information bottleneck for the input X of mode m under the parameter set of mode m. m The output Z of mode m with information bottleneck m The mutual information between

[0057] e m′ represents the embedded distribution variable of mode m′;

[0058] z m represents the output variable of mode m;

[0059] x m represents the input variable of mode m;

[0060] p(e m′ ,z m θ m ) means that under the parameter set of mode m, e m′ and z m The joint probability distribution of

[0061] p(e m′ ) means e m′ The probability distribution of

[0062] p(z m ) represents z m The probability distribution of

[0063] p(e m′ ∣z m ) indicates that at z m Under the conditions, e m′ The probability distribution of

[0064] p(x m ,z m θ m ) means that under the parameter set of mode m, x m and z m The joint probability distribution of

[0065] p(x m ) represents x m The probability distribution of

[0066] p(x m ∣z m ) indicates that at zm Under the condition of, x m Probability distribution of

[0067] Furthermore, in practical applications, through the empirical data distribution, the objective function is converted into the following empirical objective function:

[0068]

[0069] In the formula:

[0070] Represents the variational optimization objective function of the empirical data distribution under the parameter set of modality m;

[0071] N represents the total number of input variables;

[0072] n represents the input variable serial number;

[0073] Represents the nth input variable of modality m;

[0074] Represents that under the parameter set of modality m, the condition is the nth input variable of modality m The output variable z m Probability distribution of

[0075] q(e m′ ∣z m ) represents the probability distribution of the embedded distribution variable e m of modality m' under the condition of the output variable z m′ Probability distribution of

[0076] Represents the KL divergence function of the above two distributions;

[0077] β m Represents the compression factor of modality m;

[0078] Define g m As the mapping function after the bottleneck layer of the vision-language pre-training model for modality m, and for each evaluation point, the final embeddings of each modality are normalized in the embedding dimension; for the normalized g m (z m ) and e m′ , the logarithmic form of the Gaussian probability density q(f m′ (x m′ )∣g m (z m )) is simplified to be proportional to the cosine similarity between f m′ (x m′ ) and g m (z m ), thus obtaining the following final optimization objective:

[0079]

[0080] wherein:

[0081] represents the final optimization objective function of the variational distribution of empirical data under the parameter set of modality m;

[0082] S cosine (·,·) represents the cosine similarity function;

[0083] S cosine (e m′ , g m (z m )) represents the cosine similarity between the embedding e m′ of modality m′ and the mapping output g m (z m ).

[0084] The present invention also provides a device for constructing a concept bottleneck model based on a variational information bottleneck guidance method, including a memory and a processor, where the memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of the method for constructing a concept bottleneck model based on variational information bottleneck guidance as described above when executing the computer program.

[0085] The advantages and positive effects of the present invention are:

[0086] The present invention proposes a method for constructing a concept bottleneck model based on variational information bottleneck guidance. By introducing the thought chain technology, it prompts the vision - language model and the large - language model to generate a concept pool with stronger correlation and richer visual attribute information. It deeply aligns the image - text embeddings encoded by CLIP using variational information bottleneck, and further improves the classification accuracy using a complementary concept bottleneck.

[0087] The present invention applies the information bottleneck theory to the process of concept screening. Essentially, it reduces the irrelevant and redundant information in the text description and only retains the key visual concepts, thereby further improving the interpretability of the concept bottleneck model. Through information compression, while the model retains the concepts useful for image classification, it removes the irrelevant parts in the text description, thus retaining a more concise and effective concept representation.

[0088] The present invention proposes a multi-modal information bottleneck optimization objective. By introducing embedding similarity into the information bottleneck framework, the present invention can effectively establish semantic alignment in image-text pairs, thereby enhancing the attribution and interpretability of the model. This optimization objective combines the advantages of variational inference and self-supervised learning, enabling the model to learn interpretable latent representations without task-specific labels, improving the flexibility and robustness of multi-modal attribution analysis.

[0089] In addition, the present invention proposes a new metric for evaluating the interpretability efficiency of the concept bottleneck model and conducts a complete evaluation under various experimental settings. Compared with the most popular concept bottleneck model methods currently, the model of the present invention provides better interpretability while achieving similar or even better classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1 It is a schematic structural diagram of a concept bottleneck model guided by variational information bottleneck according to the present invention;

[0091] Figure 2 It is the Instruct BLIP prompt template based on the chain of thought technology mentioned in Embodiment 1 of the present invention;

[0092] Figure 3 It is the GPT-3.5 Turbo prompt template based on the chain of thought technology mentioned in Embodiment 1 of the present invention.

[0093] In the figure:

[0094] f represents the visual features encoded by the CLIP image encoder.

[0095] α represents the weight coefficient; it corresponds to the output fusion weight of the basic concept classifier.

[0096] b represents the basic concept bottleneck.

[0097] c represents the supplementary concept bottleneck.

[0098] Φ I represents the CLIP image encoder.

[0099] Φ T represents the CLIP text encoder. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0100] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments. It should be understood that the preferred embodiments described herein are only for illustrating and explaining the present invention and are not used to limit the present invention.

[0101] The Chinese interpretations of the following English words, phrases and abbreviations are as follows:

[0102] Instruct BLIP: A multimodal pre-training model based on instructions that combines visual and language understanding capabilities.

[0103] GPT-3.5 Turbo: A large language model that supports text generation, dialogue interaction, and multitasking.

[0104] CLIP: A contrastive language-image pre-training model that achieves cross-modal understanding through image-text contrastive learning.

[0105] CBM: A Concept Bottleneck Model that improves the interpretability of the model by explicitly learning the "concept layer".

[0106] ImageNet: A large-scale image classification dataset containing 14 million annotated images covering more than 20,000 categories. CIFAR-10: A small-scale image classification benchmark dataset containing 10 categories of objects respectively.

[0107] CIFAR-100: A small-scale image classification benchmark dataset containing 100 categories of objects respectively.

[0108] CUB: The Caltech-UCSD Birds Dataset, used for fine-grained image classification.

[0109] Food-101: A food classification dataset (101 categories), focusing on image recognition tasks in a specific domain.

[0110] Flower-102: A flower classification dataset (102 categories), focusing on image recognition tasks in a specific domain.

[0111] ImageNet-100: A subset of ImageNet, selecting 100 categories for lightweight experiments or resource-constrained scenarios. PCBM: A Probabilistic Concept Bottleneck Model that introduces probabilistic concept representations.

[0112] LF-CBM: A Concept Bottleneck Model that does not require manually annotated concept labels.

[0113] LaBo: A Concept Bottleneck Model based on generating concepts from a large language model and screening submodular functions.

[0114] LM4CV: A Concept Bottleneck Model that learns a small number of visual attributes for image classification.

[0115] ConceptNet: A commonsense knowledge graph containing a large number of concepts.

[0116] Please refer to Figures 1 to 3 , a method for constructing a Concept Bottleneck Model guided by variational information bottleneck, including the following method steps:

[0117] Step 1, build a concept bottleneck model with the following structure: make the concept bottleneck model include a vision-language model, a large language model, variational information bottleneck model A, variational information bottleneck model B, text encoder A of CLIP, text encoder B of CLIP, image encoder of CLIP, dot product multiplier A, dot product multiplier B, basic concept classifier, supplementary concept classifier, and classification fusion device; connect the vision-language model, variational information bottleneck model A, and text encoder A of CLIP in sequence; connect the large language model, variational information bottleneck model B, and text encoder B of CLIP in sequence, and connect the output ends of text encoder A of CLIP and image encoder of CLIP to the two input ends of dot product multiplier A correspondingly; connect the output ends of text encoder B of CLIP and image encoder of CLIP to the two input ends of dot product multiplier B correspondingly; connect the output end of dot product multiplier A to the input end of the basic concept classifier, connect the output end of dot product multiplier B to the input end of the supplementary concept classifier, and connect the output ends of the basic concept classifier and the supplementary concept classifier to the input end of the classification fusion device;

[0118] Step 2, input the image data into the vision-language model, and use the vision-language model to generate a basic concept pool directly related to the content of the input image; make the large language model input prompt information, and use the large language model to generate a descriptive supplementary concept pool covering diverse attribute features of the target category;

[0119] Step 3, screen out relatively highly relevant core basic concepts from the basic concept pool through variational information bottleneck model A, and screen out relatively highly relevant core supplementary concepts from the supplementary concept pool through variational information bottleneck model B. The core basic concepts and core supplementary concepts are used to construct a concept bottleneck;

[0120] Step 4, process the core basic concepts and core supplementary concepts respectively through text encoder A and B of CLIP to obtain basic concept embeddings and supplementary concept embeddings correspondingly; input the image data into the image encoder of CLIP, and process it through the image encoder of CLIP to obtain image embeddings; perform dot product on the basic concept embeddings and image embeddings through dot product multiplier A, and input the dot product result of the two into the basic concept classifier; perform dot product on the supplementary concept embeddings and image embeddings through dot product multiplier B, and input the dot product result of the two into the supplementary concept classifier; fuse the outputs of the basic concept classifier and the supplementary concept classifier through the classification fusion device to obtain a comprehensive classification prediction result;

[0121] Step 5, evaluate the concept bottleneck model using an interpretability efficiency metric to balance classification accuracy and the efficiency of concept interpretation.

[0122] Preferably, in step 2, the vision-language model can adopt the Instruct BLIP model; the Instruct BLIP model extracts various visual features from the image; a chain-of-thought prompt template can be set, and the chain-of-thought prompt template is used to guide the Instruct BLIP model to gradually generate concepts related to the image from multiple perspectives and levels, and finally form a fine-grained basic concept pool.

[0123] Use the vision-language model Instruct BLIP to perform a zero-shot image caption generation task to generate a basic concept pool based on the input image, where the generated concepts can accurately reflect the image content and can provide fine-grained visual attribute information related to the image content; adopt the chain-of-thought technique to design a guidance template to refine the generation process of Instruct BLIP, thereby improving the accuracy and diversity of the generated concept descriptions, especially in the performance of fine-grained image classification tasks.

[0124] By combining visual information and language information, Instruct BLIP can extract rich visual features from the image and generate relevant concepts through interaction with natural language. These generated concepts can be basic visual attributes describing the image, such as color, shape, texture, etc., or high-level semantic concepts related to the image content, such as object categories, scene descriptions, etc.

[0125] To improve the fine-grained description of visual attributes, the chain-of-thought technique is introduced. The chain-of-thought technique guides the model to reason step by step, thereby generating more refined concepts. When generating the prompt template, considering the complexity and diversity of the input image, the chain-of-thought will guide the Instruct BLIP model to gradually generate concepts related to the image from multiple perspectives and levels, and finally form a fine-grained concept pool.

[0126] In this way, a concept pool closely related to the input image and refined to each visual attribute can be obtained, thus providing a more accurate and comprehensive basis for subsequent concept prediction and model interpretation.

[0127] Preferably, in step 2, the large language model can adopt the GPT-3.5 Turbo model; a chain-of-thought prompt template can be set, and the GPT-3.5 Turbo model generates diverse text descriptions according to the prompts given by the chain-of-thought prompt template to ensure that all comprehensive attribute features of the target category are covered.

[0128] GPT-3.5 Turbo has powerful text generation capabilities and can generate diverse text descriptions based on given prompts. To cover the diverse attribute features of the target category, the supplementary concept pool generated by GPT-3.5 Turbo will contain multiple descriptive concepts related to the target category. These concepts can cover various attributes of the target object, such as structural location, local contrast, quantitative description, etc.

[0129] By designing a prompt template based on the chain of thought technique, GPT-3.5 Turbo can generate diverse supplementary concepts for each category, thus enriching the final concept pool and ensuring that it covers the comprehensive attribute features of the target category. This method of generating the supplementary concept pool not only improves the interpretability of the model but also ensures that the model can fully capture the diverse features of the target category, thereby enhancing the accuracy and credibility of concept prediction.

[0130] Using the large language model GPT-3.5 Turbo to generate a supplementary concept pool, the supplementary concept pool includes diverse descriptive concepts of image categories. The designed chain of thought prompt template ensures that the generated concepts accurately describe the image categories and avoids generating noise information unrelated to visual attributes.

[0131] Preferably, in step 3, the variational information bottleneck models A and B can align the visual concepts in the text description with the image regions without relying on concept labels; use the information bottleneck of the vision-language variant for feature attribution analysis, calculate the importance scores of each concept in the text description and the image regions, and then obtain the comprehensive scores of each concept according to the frequency of concept occurrence and concept professionalism. The top K concepts in the basic concept pool can be selected in the order of the comprehensive scores from high to low to construct the core basic concepts; the core supplementary concepts are constructed from the top V concepts of the supplementary concept pool.

[0132] The present invention uses the proposed concept screening method based on variational information bottleneck to extract the top K text concepts with the highest scores from the descriptions generated by Instruct BLIP, which are called core basic concepts. For the mapping from the core basic concepts to the image categories, the image encoder and text encoder of CLIP are respectively used to extract the image embeddings and concept embeddings, calculate the dot product to obtain the K-dimensional concept bottleneck, and then obtain the scores of the categories by learning the basic concept classifier.

[0133] Meanwhile, leveraging the generative capabilities of large language models, a supplementary concept bottleneck is constructed using richer, more detailed, and diverse text descriptions as supplementary information. Using the same concept screening method, the top V text concepts with the highest scores are selected as the core supplementary concepts, which are encoded by the text encoder of CLIP to obtain V vectors with dimensions consistent with the existing concept embeddings. Then, the visual representation is mapped onto these vectors to obtain the supplementary concepts. These concepts are used to supplement the prediction results of the original concepts.

[0134] Preferably, in step 4, it can be set that: Ψ b (b) is the output of the basic concept classifier, and Ψ c (c) is the output of the supplementary concept classifier; the classification fusion device can fuse the outputs of the basic concept classifier and the supplementary concept classifier according to the following formula:

[0135]

[0136] In the formula:

[0137] represents the prediction result;

[0138] α represents the weight coefficient;

[0139] b represents the basic concept bottleneck;

[0140] c represents the supplementary concept bottleneck;

[0141] f represents the visual features encoded by the CLIP image encoder;

[0142] W represents the basic concept text features obtained by the CLIP text encoder;

[0143] U represents the supplementary concept text features obtained by the CLIP text encoder;

[0144] Ψ b () represents the basic concept linear classifier function;

[0145] Ψ c () represents the supplementary concept linear classifier function.

[0146] Preferably, a dual-branch concept bottleneck network is constructed by the basic concept classifier and the supplementary concept classifier. To maximize the accuracy when only using the basic concepts, the following optimization objective of the basic concept classifier can be constructed:

[0147]

[0148] On the other hand, the present invention aims to improve the performance of image classification by introducing supplementary concepts, which can be solved by the branch of supplementary concepts. The optimization objective for constructing the supplementary concept classifier is as follows:

[0149]

[0150] Wherein:

[0151] Φ I () represents the image encoder function of CLIP;

[0152] t represents the input image;

[0153] y represents the class label;

[0154] L() represents the cross-entropy loss function;

[0155] Ω() represents the complexity metric function of the regularized model;

[0156] λ represents the regularization strength;

[0157] represents the expression expectation;

[0158] Ψ b represents the basic concept linear classifier for classification using basic concepts;

[0159] Ψ c represents the supplementary concept linear classifier for classification using supplementary concepts;

[0160] When training the basic concept classifier and the supplementary concept classifier, forward and backward propagation are performed twice in each batch to update the parameters; in the first propagation, only the basic concept classifier is optimized; in the second propagation, the parameters of the basic concept classifier are fixed, and the supplementary concept classifier is optimized; in order to avoid the influence of extreme values of supplementary concepts on the results, cosine similarity rather than projection distance is used as the concept bottleneck, and the activation values of each concept are normalized so that the mean is 0 and the standard deviation is 1 to accelerate convergence.

[0161] Preferably, in step 5, the calculation formula of the interpretability efficiency index can be as follows:

[0162]

[0163] Wherein:

[0164] IE represents the interpretability efficiency;

[0165] Acc represents the classification accuracy;

[0166] i represents the concept serial number;

[0167] j represents the category serial number;

[0168] W F represents the linear layer for mapping concepts to categories;

[0169] (W F ) ij represents the weight corresponding to the i-th concept and the j-th category in the linear layer;

[0170] O{(W F ) ij ≠0} represents the number of non-zeros in the linear layer;

[0171] R represents the number of concepts;

[0172] S represents the number of categories;

[0173] represents the average concept length.

[0174] Preferably, in step 3, for the variational information bottleneck models A and B, the variational approximation method can be used to optimize the objective functions of both. The KL divergence between the image and the text can be used to adjust the parameters of both. Combining the text encoders A and B of the pre-trained CLIP, importance scores are assigned to the image regions and text concepts.

[0175] The present invention applies the information bottleneck theory to the process of concept screening. Essentially, it reduces the irrelevant and redundant information in the text description and only retains the key visual concepts, thereby further improving the interpretability of the concept bottleneck model. Through information compression, while the model retains the concepts useful for image classification, it removes the irrelevant parts in the text description, thus retaining a more concise and effective concept representation.

[0176] In the process of concept screening, the original formula of the information bottleneck theory is not directly used. The present invention only uses the generated text and the given image to learn the latent concept representation without relying on concept labels. This method extracts more accurate and concise concepts through the information bottleneck theory and is used for the construction of the concept bottleneck. In order to construct a variational information bottleneck model for the multi-modal pre-trained model, a loss function similar to the optimization objective of the self-supervised method is constructed, and it only depends on the given image and the description of the concepts contained in the image to perform the representation learning between the image regions and text concepts.

[0177] In multi-modal data, there is a reasonable prior for information correlation: if the image and text inputs are related to each other (for example, the text describes the image content), then a better image encoding should contain information about the text, and vice versa. Based on this intuitive understanding, preferably, for the variational information bottleneck models A and B, the following objective functions can be set:

[0178]

[0179] In the formula:

[0180] m represents the image or text modality;

[0181] m′ represents the complement of modality m;

[0182] θ m represents the parameter set of modality m;

[0183] represents the optimization objective under the parameter set of modality m;

[0184] Z m represents the output of the information bottleneck for modality m;

[0185] E m′ represents the embedding of modality m′;

[0186] X m represents the input of the information bottleneck for modality m;

[0187] β represents the scaling factor;

[0188] I(Z m ,E m′ ;θ m ) represents the mutual information between the embedding E of modality m′ and the output Z of the information bottleneck for modality m under the parameter set of modality m; m′ and the output Z of the information bottleneck for modality m; m between;

[0189] I(Z m ,X m ;θ m ) represents the mutual information between the input X of the information bottleneck for modality m and the output Z of the information bottleneck for modality m under the parameter set of modality m; m and the output Z of the information bottleneck for modality m; m between;

[0190] e m′ represents the embedding distribution variable of modality m′;

[0191] z m represents the output variable of modality m;

[0192] x m represents the input variable of modality m;

[0193] p(e m′ ,z m ;θ m ) represents the joint probability distribution of e m′ and z m under the parameter set of modality m;

[0194] p(e m′ ) represents the probability distribution of e m′ ;

[0195] p(z m ) represents the probability distribution of z m ;

[0196] p(e m′ ∣z m ) represents the probability distribution of e m under the condition of z m′ ;

[0197] p(x m ,z m ; θ m ) represents the joint probability distribution of x m and z m under the parameter set of modality m;

[0198] p(x m ) represents the probability distribution of x m ;

[0199] p(x m ∣z m ) represents the probability distribution of x m under the condition of z m .

[0200] The present invention will utilize the information bottleneck principle of this visual - language variant for attribution analysis. Through this optimization objective, information associations can be constructed between different modalities, ensuring that while the model captures the image content, it can also convey the semantic connection with the text modality, providing a solid foundation for the construction of the subsequent concept bottleneck.

[0201] Preferably, in practical applications, through the empirical data distribution, the objective function can be converted into the following empirical objective function:

[0202]

[0203] where:

[0204] represents the variational optimization objective function of the empirical data distribution under the parameter set of modality m;

[0205] N represents the total number of input variables;

[0206] n represents the input variable serial number;

[0207] represents the n - th input variable of modality m;

[0208] Denote the output variable z under the parameter set of modality m with the condition being the n-th input variable of modality m ; m The probability distribution;

[0209] q(e m′ ∣z m ) represents the probability distribution of the embedding distribution variable e of modality m' under the condition of the output variable z m ; m′ The probability distribution;

[0210] Denote the KL divergence function of the above two distributions;

[0211] β m represents the compression factor of modality m;

[0212] Define g m as the mapping function after the bottleneck layer of the vision-language pre-training model for modality m, and for each evaluation point, the final embeddings of each modality are normalized in the embedding dimension; for the normalized g m (z m ) and e m′ , the logarithmic form of the Gaussian probability density q(f m′ (x m′ )∣g m (z m )) is simplified to be proportional to the cosine similarity between f m′ (x m′ ) and g m (z m ), thus obtaining the following final optimization objective:

[0213]

[0214] In the formula:

[0215] Denote the variational final optimization objective function of the empirical data distribution under the parameter set of modality m;

[0216] S cosine (·,·) represents the cosine similarity function;

[0217] S cosine (e m′ ,g m (z m )) represents the cosine similarity between the embedding e of modality m' m′ and the mapping output g m (z m ).

[0218] Under the above optimization objectives, by introducing embedding similarity into the information bottleneck framework, the present invention can effectively establish semantic alignment in image-text pairs, thereby enhancing the attribution and interpretability of the model. This optimization objective combines the advantages of variational inference and self-supervised learning, enabling the model to learn interpretable latent representations without task-specific labels, improving the flexibility and robustness of multimodal attribution analysis.

[0219] The present invention also provides a device for constructing a concept bottleneck model guided by variational information bottleneck, including a memory and a processor. The memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of the method for constructing a concept bottleneck model guided by variational information bottleneck as described above when executing the computer program.

[0220] The following uses several preferred embodiments of the present invention to further illustrate the working process and working principle of the present invention:

[0221] The present invention introduces the chain of thought technique to prompt the vision-language model and the large language model to generate a concept pool with stronger correlation and richer visual attribute information, deeply aligns the image-text embeddings encoded by CLIP using variational information bottleneck, and further improves the classification accuracy using the supplementary concept bottleneck.

[0222] In this embodiment, six publicly available datasets, namely CIFAR-10, CIFAR-100, CUB, Food-101, Flower-102, and ImageNet-100, are selected as the reference basis for the image classification accuracy. Image preprocessing includes scaling the input image to a maximum side length of 256 pixels, randomly cropping it to 224×224 pixels, normalizing it based on the mean and standard deviation of ImageNet, and applying data augmentation methods such as random horizontal flipping. The division of the training set and the test set is based on the split file given by the official of the dataset.

[0223] Please refer to Figure 1 , the present invention proposes a method for constructing a concept bottleneck model guided by variational information bottleneck, including the following method steps:

[0224] Step 1, build a conceptual bottleneck model with the following structure: make the conceptual bottleneck model include a vision-language model, a large language model, variational information bottleneck model A, variational information bottleneck model B, CLIP's text encoder A, CLIP's text encoder B, CLIP's image encoder, dot product multiplier A, dot product multiplier B, basic concept classifier, supplementary concept classifier, and classification fusion device; connect the vision-language model, variational information bottleneck model A, and CLIP's text encoder A in sequence; connect the large language model, variational information bottleneck model B, and CLIP's text encoder B in sequence, and connect the output ends of CLIP's text encoder A and CLIP's image encoder to the two input ends of dot product multiplier A correspondingly; connect the output ends of CLIP's text encoder B and CLIP's image encoder to the two input ends of dot product multiplier B correspondingly; connect the output end of dot product multiplier A to the input end of the basic concept classifier, connect the output end of dot product multiplier B to the input end of the supplementary concept classifier, and connect the output ends of the basic concept classifier and the supplementary concept classifier to the input end of the classification fusion device; input image data into the vision-language model and CLIP's image encoder, and output the predicted classification result by the classification fusion device.

[0225] Step 2, use the vision-language model Instruct BLIP to generate a basic concept pool directly related to the content of the input image, and optimize the generation of a prompt template through the chain of thought technology to enhance the fine-grainedness of visual attribute description:

[0226] For the construction of the basic concept pool, the present invention hopes to use Instruct BLIP to generate text concepts with obvious correspondence to image regions. At the same time, in order to obtain richer visual attributes, the present invention designs to guide and improve the detail richness. At the same time, in order to guide the vision-language model to discover visual attributes different from other similar categories in a fine-grained dataset, a reasoning process of local comparison with similar categories is further added. For the Instruct BLIP prompt template based on the chain of thought technology, please refer to Figure 2 。

[0227] Use the large language model GPT-3.5Turbo to generate a descriptive supplementary concept pool covering diverse attribute features of the target category:

[0228] For the construction of the supplementary concept pool, the present invention uses GPT-3.5Turbo to generate more abundant descriptive statements for specific image categories, and at the same time restricts the text that has nothing to do with visual attributes, such as the calls and behavioral characteristics of specific birds, to prevent introducing noise into the conceptual bottleneck. For the GPT-3.5Turbo prompt template based on the chain of thought technology, please refer to Figure 3 。

[0229] Step 3: Based on the variational information bottleneck principle, calculate the concept importance scores through the cross-modal aligned image-text embedding space, and screen out highly relevant concepts.

[0230] The information bottleneck principle provides a framework for finding a compressed representation of a neural network model. To obtain a latent representation that can reflect the most relevant information of the input data, the information bottleneck principle aims to express the input source X through a stochastic latent representation Z, where the latent representation Z is defined by the parameterized encoder p Z|X (z∣x; θ), which can maximize the information about the target Y while restricting the mutual information between the latent representation Z and the input X.

[0231] This principle can be formulated as the following optimization problem:

[0232]

[0233] In the expression of the optimization problem:

[0234] X represents the input source;

[0235] Y represents the target;

[0236] Z represents the latent representation;

[0237] θ represents the specific parameter representation under the optimization problem;

[0238] I(·,·; θ) represents the mutual information function;

[0239] I(Z,Y; θ) represents the mutual information between the latent representation Z and the target Y under the parameter θ;

[0240] I(Z,X; θ) represents the mutual information between the latent representation Z and the input source X under the parameter θ;

[0241] represents a compression constraint used to control the upper bound of the mutual information between Z and X.

[0242] Where I(·,·; θ) represents the mutual information function, and is a compression constraint used to control the upper bound of the mutual information between Z and X to ensure that the model only retains the key information in the input.

[0243] This optimization problem can be equivalently transformed into maximizing the following objective function:

[0244]

[0245] In the expression of the objective function:

[0246] θ represents the specific parameter representation under the optimization problem;

[0247] Denote the optimal objective under parameter θ;

[0248] β represents a scaling factor, which is used to balance between maximizing the information of Z about the target Y and maximizing the compressibility of Z with respect to the input X. When β is large, the optimization process pays more attention to information compression; when β is small, the model focuses more on the relevance of the latent representation to the target Y;

[0249] I(·,·; θ) represents the mutual information function;

[0250] I(Z, Y; θ) represents the mutual information between the latent representation Z and the target Y under parameter θ;

[0251] I(Z, X; θ) represents the mutual information between the latent representation Z and the input source X under parameter θ.

[0252] The present invention applies the information bottleneck theory to the process of concept screening, which essentially reduces the irrelevant and redundant information in the text description and only retains the key visual concepts, thereby further improving the interpretability of the concept bottleneck model. Through information compression, the model removes the irrelevant parts in the text description while retaining the concepts useful for image classification, thus retaining a more concise and effective concept representation.

[0253] To construct a variational information bottleneck principle for multimodal pre-training models such as CLIP, a loss function more similar to the optimization objective of self-supervised methods needs to be designed, which only depends on the given image and the description of the concepts contained in the image to perform representation learning between image regions and text concepts. This learning method is essentially different from the supervised learning paradigm in single-modal task settings. For example, for an image X of a dog dog and the corresponding label Y dog ="dog", in the single-modal classification task, the present invention can simply maximize I(Y dog , Z dog ; θ)-βI(X dog , Z dog ; θ), where Z dog represents the latent representation of X dog . In the text-image representation learning, there is usually a text description L dog , such as "This is a photo of a dog running on the grass. The dog has four legs, soft fur, and a wagging tail", rather than a definite and single label. In this setting, both X dog and L dog are inputs and there is no predefined label.

[0254] To obtain a task-agnostic graphical representation, the present invention aims to use both image and text input modalities simultaneously and define a variational information bottleneck principle, while the output closely depends on the specific downstream task. This requires redefining the "fitting term" I(Z, Y; θ) in the classical information bottleneck objective to adapt to this graphical representation learning scenario. In multimodal data, there is a reasonable prior for information correlation: if the image and text inputs are related to each other (e.g., the text describes the image content), then a good image encoding should contain information about the text, and vice versa. Based on this intuitive understanding, an information bottleneck objective for vision-language can be expressed as:

[0255]

[0256] In the expression of the information bottleneck objective for vision-language:

[0257] m represents the image or text modality;

[0258] m′ represents the complement of modality m;

[0259] θ m represents the set of parameters for modality m;

[0260] represents the optimization objective under the set of parameters for modality m;

[0261] Z m represents the output of the information bottleneck for modality m;

[0262] E m′ represents the embedding of modality m′;

[0263] X m represents the input to the information bottleneck for modality m;

[0264] β represents the scaling factor;

[0265] I(Z m ,E m′ ;θ m ) represents the mutual information between the embedding E m′ of modality m′ and the output Z m of the information bottleneck for modality m under the set of parameters for modality m;

[0266] I(Z m ,X m ;θ m ) represents the mutual information between the input X m to the information bottleneck for modality m and the output Z m of the information bottleneck for modality m under the set of parameters for modality m.

[0267] Through the theoretical derivation of the variational information bottleneck feature attribution method, the final optimization objective expression is as follows:

[0268]

[0269] In the expression of the final optimization objective:

[0270] m represents the image or text modality;

[0271] m′ represents the complement of modality m;

[0272] θ m represents the parameter set of modality m;

[0273] represents the variational optimization objective function of the empirical data distribution under the parameter set of modality m;

[0274] N represents the total number of input variables;

[0275] n represents the input variable serial number;

[0276] z m represents the output variable of modality m;

[0277] x m represents the input variable of modality m;

[0278] represents the nth input variable of modality m;

[0279] e m′ represents the embedding distribution variable of modality m′;

[0280] represents the probability distribution of the output variable z under the condition of the nth input variable of modality m m ;

[0281] S cosine (·,·) represents the cosine similarity function;

[0282] S cosine (e m′ ,g m (z m )) represents the cosine similarity between the embedding e m′ of modality m′ and the mapped output g m (z m );

[0283] represents the KL divergence function of the above two distributions;

[0284] β m represents the compression factor of modality m.

[0285] Under this multimodal information bottleneck optimization objective, by introducing embedding similarity into the information bottleneck framework, the present invention can effectively establish semantic alignment in image-text pairs, thereby improving the attribution and interpretability of the model. This optimization objective combines the advantages of variational inference and self-supervised learning, enabling the model to learn interpretable latent representations without the need for task-specific labels, thereby improving the flexibility and robustness of multimodal attribution analysis.

[0286] Step 4: A dual-branch concept bottleneck network is constructed by the basic concept classifier and the supplementary concept classifier, and the prediction results of the two are fused through a classification fusion device.

[0287] First, the proposed concept screening method based on variational information bottleneck is used to extract the top K text concepts with the highest scores from the description generated by Instruct BLIP, which are called "basic concepts". For the mapping of basic concepts to image categories, the method unified with previous work is adopted, and the image encoder and text encoder of CLIP are used to extract image embedding and concept embedding respectively, and the dot product is calculated to obtain the K-dimensional concept bottleneck, and then the category score is obtained by learning the basic concept classifier.

[0288] At the same time, we use the generative power of the large language model and richer, more detailed and diverse text descriptions as supplementary information to build a supplementary concept bottleneck. Using the same concept screening method, we select the top V text concepts with the highest scores and encode them through the text encoder of CLIP to obtain V vectors. The dimensions of these vectors are consistent with the existing concept embeddings, and then map the visual representation to these vectors to obtain supplementary concepts. The prediction results of using these concepts to supplement the original concepts are as follows:

[0289]

[0290] In the prediction result expression:

[0291] Indicates the prediction result;

[0292] α represents the weight coefficient;

[0293] b represents the basic concept bottleneck;

[0294] c represents the supplementary concept bottleneck;

[0295] f represents the visual features encoded using the CLIP image encoder;

[0296] W represents the basic concept text features obtained using the CLIP text encoder;

[0297] U represents the supplementary concept text features obtained using the CLIP text encoder.

[0298] Step 5: Optimize the model using the interpretability efficiency metric to balance the classification accuracy and the efficiency of concept explanation.

[0299] The concept bottleneck model cannot be considered solely from the aspect of classification accuracy. As an interpretable image classification method, the present invention also needs to consider the feasibility and effectiveness of its explanation. In addition to using the classification accuracy as the basic metric for model performance, the present invention proposes a metric for measuring the interpretability efficiency of the concept bottleneck model, which focuses on how many concepts actually participate in the final classification decision and the average number of words in the text concepts used for decision-making. The calculation formula for the interpretability efficiency metric is as follows:

[0300]

[0301] In the formula:

[0302] IE represents the interpretability efficiency;

[0303] Acc represents the classification accuracy;

[0304] i represents the concept serial number;

[0305] j represents the class serial number;

[0306] W F represents the linear layer used to map concepts to classes;

[0307] (W F ) ij represents the weight corresponding to the i-th concept and the j-th class in the linear layer;

[0308] O{(W F ) ij ≠0} represents the number of non-zero values in the linear layer;

[0309] R represents the number of concepts;

[0310] S represents the number of classes;

[0311] represents the average concept length.

[0312] Example 2:

[0313] Based on Example 1, but different in the advanced comparison methods selected, including PCBM, LF-CBM, LaBo, and LM4CV. Among them, PCBM uses the ConceptNet knowledge graph dataset to construct a concept library. It treats categories as nodes and incorporates surrounding neighbor nodes into the concept library. By expanding multiple levels outward, it can include more concepts and maps visual embeddings to categories through residual connections; LF-CBM and LaBo generate concepts using large language models and construct concept bottlenecks using different concept screening methods; LM4CV uses a learning search method to mine a more concise set of concepts in the dataset, greatly reducing the size of the concept bottleneck and making the interpretation results easier to understand. The present invention trains and tests the above methods on the same dataset and plots the results in Table 1.

[0314] As can be seen from Table 1, a concept bottleneck model based on variational information bottleneck guidance proposed by the present invention performs well in image classification tasks, showing classification accuracy similar to or even better than other methods.

[0315] As can be seen from Table 2, the experimental results of each comparison method in terms of interpretability.

[0316] The proposed interpretability efficiency evaluation index of the present invention is used to measure the simplicity of the interpretation output of the concept bottleneck model method. Experiments show that the interpretability efficiency of the present invention reaches the highest level on six datasets.

[0317] Table 1: Experimental results of classification accuracy of each comparison method

[0318]

[0319] Table 2: Experimental results of interpretability efficiency of each comparison method

[0320]

[0321]

[0322] Table 3 shows the influence of different concept screening methods on classification accuracy. The experimental results under all sample quantity settings are the average values of the classification accuracies of six datasets. Among them:

[0323] The randomly selected screening method constructs a concept bottleneck by randomly extracting vocabulary from the concept pool.

[0324] The screening method of similarity ranking calculates the similarity between concepts in the concept pool and the category name respectively, sorts them according to the average similarity size, and selects the top K concepts to construct a concept bottleneck.

[0325] The submodular function is a method for selecting the optimal subset in a high-dimensional space, and LaBo applies it to the process of concept screening.

[0326] The variational information bottleneck is the concept screening method proposed in the present invention. By utilizing the cross-modal ability of CLIP, it strictly aligns the concept parts in the image and text descriptions and suppresses the influence of the noise regions. It can be seen that the concept screening method based on the variational information bottleneck proposed in the present invention achieves the highest classification accuracy under all settings.

[0327] Table 3: Classification accuracy based on different concept screening methods

[0328]

[0329] For the above-mentioned functional modules, systems and algorithms such as the vision-language model, large language model, variational information bottleneck model A, variational information bottleneck model B, text encoder A of CLIP, text encoder B of CLIP, image encoder of CLIP, dot product multiplier A, dot product multiplier B, basic concept classifier, supplementary concept classifier, classification fusion device, variational approximation method, KL divergence, GPT-3.5 Turbo model, InstructBLIP model, thought chain prompt template, etc., applicable functional modules, systems and algorithms in the prior art can be adopted, or applicable functional modules, systems and algorithms in the prior art can be adopted and constructed by using conventional technical means.

[0330] It should be noted that the above are only the preferred embodiments of the present invention and do not constitute a limitation on the protection scope of the present invention. Any equivalent replacement or modification made by those skilled in the art within the technical scope disclosed by the present invention based on the technical solution and inventive concept of the present invention shall be regarded as being included within the protection scope of the present invention.

Claims

1. A method for constructing a concept bottleneck model based on variational information bottleneck guidance, characterized in that: The method comprises the following steps: Step 1, build a concept bottleneck model according to the following structure: the concept bottleneck model includes a visual language model, a large language model, a variational information bottleneck model A, a variational information bottleneck model B, a text encoder A of CLIP, a text encoder B of CLIP, an image encoder of CLIP, a dot multiplier A, a dot multiplier B, a basic concept classifier, a supplementary concept classifier and a classification fusion device; the visual language model, the variational information bottleneck model A, and the text encoder A of CLIP are connected in sequence; the large language model, the variational information bottleneck model B, and the text encoder B of CLIP are connected in sequence, and the output ends of the text encoder A of CLIP and the image encoder of CLIP are connected to the two input ends of the dot multiplier A correspondingly; the output ends of the text encoder B of CLIP and the image encoder of CLIP are connected to the two input ends of the dot multiplier B correspondingly; the output end of the dot multiplier A is connected to the input end of the basic concept classifier, the output end of the dot multiplier B is connected to the input end of the supplementary concept classifier, and the output end of the basic concept classifier and the output end of the supplementary concept classifier are connected to the input end of the classification fusion device; Step 2: Input the image data into the visual language model, and use the visual language model to generate a basic concept pool directly related to the input image content; The large language model is used to input prompt information and generate a descriptive supplementary concept pool to cover the diverse attribute characteristics of the target category. Step 3: select core basic concepts with high relevance from the basic concept pool through the variational information bottleneck model A, and select core supplementary concepts with high relevance from the supplementary concept pool through the variational information bottleneck model B. The core basic concepts and core supplementary concepts are used to construct concept bottlenecks; Step 4: Process the core basic concepts and core supplementary concepts through the text encoder A and B of CLIP respectively to obtain the basic concept embedding and supplementary concept embedding respectively; input the image data into the image encoder of CLIP, and process it through the image encoder of CLIP to obtain the image embedding; perform dot multiplication of the basic concept embedding and the image embedding through the dot multiplier A, and input the dot multiplication result of the two into the basic concept classifier; perform dot multiplication of the supplementary concept embedding and the image embedding through the dot multiplier B, and input the dot multiplication result of the two into the supplementary concept classifier; fuse the outputs of the basic concept classifier and the supplementary concept classifier through the classification fusion to obtain the comprehensive classification prediction result; In step 5, the interpretability efficiency index is used to evaluate the concept bottleneck model to balance the classification accuracy and the efficiency of concept explanation.

2. The method for constructing a concept bottleneck model based on variational information bottleneck guidance according to claim 1 is characterized in that: In step 2, a thought chain prompt template is set, and the visual language model adopts the Instruct BLIP model; the InstructBLIP model extracts a variety of visual features from the image; the thought chain prompt template is used to guide the Instruct BLIP model to gradually generate concepts related to the image from multiple angles and levels, and finally form a fine-grained basic concept pool; The large language model uses the GPT-3.5Turbo model; the GPT-3.5Turbo model generates diverse text descriptions based on the prompts given by the thought chain prompt template, ensuring that comprehensive attribute features of the target category are covered.

3. The method for constructing a concept bottleneck model based on variational information bottleneck guidance according to claim 1, characterized in that: In step 3, variational information bottleneck models A and B are used to align the visual concepts in the text description with the image area without relying on concept labels. The information bottleneck of visual language variants is used for feature attribution analysis to calculate the importance scores of each concept in the text description and the image area. The comprehensive score of each concept is then obtained according to the frequency of concept occurrence and the degree of concept expertise. The top K concepts in the basic concept pool are selected in order of the comprehensive scores to construct the core basic concepts. The core supplementary concepts are constructed from the top V concepts in the supplementary concept pool.

4. The method for constructing a concept bottleneck model based on variational information bottleneck guidance according to claim 1, characterized in that: In step 4, let: b (b) is the output of the basic concept classifier, Ψ c (c) is the output of the supplementary concept classifier; the classification fusion unit fuses the outputs of the basic concept classifier and the supplementary concept classifier according to the following formula: Where: Indicates the prediction result; α represents the weight coefficient; b represents the basic concept bottleneck; c represents the supplementary concept bottleneck; f represents the visual features encoded using the CLIP image encoder; W represents the basic concept text features obtained using the CLIP text encoder; U represents the supplementary concept text features obtained using the CLIP text encoder; Ψ b () represents the basic concept linear classifier function; Ψ c () represents the complementary concept linear classifier function.

5. The method for constructing a concept bottleneck model based on variational information bottleneck guidance according to claim 1, characterized in that: When training the basic concept classifier and the supplementary concept classifier, two forward and backward propagations are performed in each batch to update the parameters; in the first propagation, only the basic concept classifier is optimized; in the second propagation, the parameters of the basic concept classifier are fixed and the supplementary concept classifier is optimized; in order to avoid the influence of extreme values ​​of the supplementary concepts on the results, cosine similarity is used as the concept bottleneck instead of projection distance, and the activation value of each concept is standardized so that its mean is 0 and the standard deviation is 1 to accelerate convergence.

6. The method for constructing a concept bottleneck model based on variational information bottleneck guidance according to claim 1, characterized in that: In step 5, the calculation formula of the interpretability efficiency index is as follows: Where: IE stands for interpretability efficiency; Acc represents the classification accuracy; i represents the concept number; j represents the category number; W F represents a linear layer used to map concepts to categories; (W F ) ij Represents the weight corresponding to the i-th concept and the j-th category in the linear layer; O{(W F ) ij ≠0} indicates the number of non-zero values ​​in the linear layer; R represents the number of concepts; S represents the number of categories; Represents the average concept length.

7. The method for constructing a concept bottleneck model based on variational information bottleneck guidance according to claim 1, characterized in that: In step 3, for the variational information bottleneck models A and B, the variational approximation method is used to optimize the objective functions of the two, and the KL divergence between the image and the text is used to adjust the parameters of the two. Combined with the pre-trained CLIP text encoders A and B, importance scores are assigned to image regions and text concepts.

8. The method for constructing a concept bottleneck model based on variational information bottleneck guidance according to claim 7, characterized in that: For the variational information bottleneck models A and B, the following objective function is set: Where: m indicates image or text modality; m′ represents the complement of mode m; θ m represents the parameter set of mode m; represents the optimization objective under the modal parameter set m; Z m represents the output of the information bottleneck to mode m; E m′ represents the embedding of modality m′; X m represents the input of the information bottleneck to the mode m; β represents the scaling factor; I(Z m ,E m′ θ m ) represents the embedding E of mode m′ under the parameter set of mode m m′ The output Z of mode m with information bottleneck m The mutual information between I(Z m ,X m θ m ) represents the information bottleneck for the input X of mode m under the parameter set of mode m. m The output Z of mode m with information bottleneck m The mutual information between e m′ represents the embedded distribution variable of mode m′; z m represents the output variable of mode m; x m represents the input variable of mode m; p(e m′ ,z m θ m ) means that under the parameter set of mode m, e m′ and z m The joint probability distribution of p(e m′ ) means e m′ The probability distribution of p(z m ) represents z m The probability distribution of p(e m′ ∣z m ) indicates that at z m Under the conditions, e m′ The probability distribution of p(x m ,z m θ m ) means that under the parameter set of mode m, x m and z m The joint probability distribution of p(x m ) represents x m The probability distribution of p(x m ∣z m ) indicates that at z m Under the condition of m The probability distribution of .

9. The method for constructing a concept bottleneck model based on variational information bottleneck guidance according to claim 8, characterized in that: In practical applications, the objective function is converted into the following empirical objective function through empirical data distribution: Where: represents the objective function of variational optimization of empirical data distribution under the parameter set of mode m; N represents the total number of input variables; n represents the input variable number; represents the nth input variable of mode m; Indicates that under the parameter set of mode m, the condition is the nth input variable of mode m The output variable z under m The probability distribution of q(e m′ ∣z m ) indicates that in the output variable z m Under the condition, the embedded distribution variable e of mode m′ m′ The probability distribution of Represents the KL divergence function of the above two distributions; β m represents the compression factor of mode m; Definition m is the mapping function after the bottleneck layer of the visual language pre-training model for modality m, and for each evaluation point, the final embedding of each modality is normalized on the embedding dimension; for the normalized g m (z m ) and e m′ , Gaussian probability density q(f m′ (x m′ )|g m (z m )) simplifies to the logarithmic form of f m′ (x m′ ) and g m (z m ), so the final optimization goal is as follows: Where: represents the final optimization objective function of the empirical data distribution variation under the parameter set of mode m; S cosine (·,·) represents the cosine similarity function; S cosine (e m′ ,g m (z m )) represents the embedding e of the modality m′ m′ With the mapping output g m (z m )’s cosine similarity.

10. A device for constructing a conceptual bottleneck model based on variational information bottleneck guidance, comprising a memory and a processor, characterized in that: The memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of the method for constructing a conceptual bottleneck model based on variational information bottleneck guidance as described in any one of claims 1 to 9 when executing the computer program.

Citation Information

Patent Citations

  • Image classification method based on cross-modal concept discovery and reasoning and intelligent terminal

    CN117115564A

  • Image reconstruction method and apparatus for cross-modal communication system

    WO2023280065A1

  • KR20240172831A

Cited By

  • False news detection method and system based on fact-emotion dual uncertainty

    CN120974384A

  • Image classification method and system based on attribute-guided concept bottleneck continual learning

    CN122551081A

  • Image classification method and system based on attribute-guided concept bottleneck continual learning

    CN122551081B