Interactive semantic-aware self-learning framework and interpretable visual recognition method

Through the interactive semantic perception self-learning framework, the semantic guidance of the teacher module is used to improve the interpretability of the visual recognition of the student module, solving the problem of lack of manual guidance and framework limitations of the model training in the prior art, and achieving higher visual recognition interpretability and discernment ability.

WO2025091768A1PCT designated stage expired Publication Date: 2025-05-08SUN YAT SEN UNIV

Patent Information

Application Number
PCT/CN2024/085289
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-01
Filing Date
2024-04-01
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The existing interpretable visual recognition methods lack manual guidance during model training, resulting in limited interpretability of abstract semantic concepts aligning with specific image areas and limited application frameworks, usually only CNN or ViT is supported.

Method used

An interactive semantic perception self-learning framework is proposed, including teacher modules and student modules. The teacher module helps the student module improve the interpretability of visual recognition through semantic guidance. The framework supports network structures such as CNN, ViT and Swin-Transformer.

Benefits of technology

Through the interactive learning mechanism, the interpretability and discrimination ability of the visual recognition model are improved, semantic-rich slices can be generated, and the transparency and traceability of the model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024085289_08052025_PF_FP_ABST
    Figure CN2024085289_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are an interactive semantic-aware self-learning framework and an interpretable visual recognition method. The interactive semantic-aware self-learning framework comprises a teacher module and a student module, wherein the student module is used for visual recognition, and the teacher module is used for performing semantic instruction on the student module; the student module inputs a calculated shard-feature pair to the teacher module, and the teacher module outputs a calculated semantic-rich slice to the student module; the teacher module comprises a first encoder, a category-concept library and a similarity comparison sub-module; and the student module comprises a shard library, a second encoder and a feature selection sub-module. The present invention accurately captures features of different granularities, has great performance in terms of generalization and accuracy, enhances the interpretability regarding the alignment between an abstract semantic concept and a specific image area, and incorporates networks of different structures.
Need to check novelty before this filing date? Find Prior Art

Description

An interactive semantic perception self-learning framework and explainable visual recognition method Technical Field

[0001] The present invention relates to the field of interpretable visual recognition technology in computers, and more specifically, to an interactive semantic perception self-learning framework and an interpretable visual recognition method. Background Art

[0002] Explainable visual recognition aims to generate interpretable and transparent feature representations, thereby enhancing the interpretability and traceability of visual recognition models. Currently, research in explainable visual recognition focuses on two main areas: self-explanatory models and post-hoc analysis methods. Self-explanatory models are transparent and interpretable, allowing for the extraction of concepts generated by visual recognition models. Post-hoc analysis methods analyze the outputs of visual recognition models to improve their interpretability. Some related research focuses on using advanced convolutional layers to learn concentrated semantic regions of image regions, such as ProtoPNet and ProtoPFormer.

[0003] ProtoPNet is a method that uses specific category prototypes based on convolutional neural networks (CNNs) to accurately perceive and identify the discriminative parts of objects. ProtoPNet compares the input image with the selected specific category prototypes and calculates similarity and other operations for visual interpretability, but this method is highly dependent on the selection of category prototypes, and the selection of category prototypes is relatively coarse-grained. ProtoPFormer uses a prototype-based approach to extend the ViT architecture to achieve interpretable visual recognition. However, due to the lack of human guidance during model training, the interpretability of aligning abstract semantic concepts with specific image regions is still limited, and the extraction of image regions is relatively coarse-grained. At the same time, the applicable frameworks of existing interpretable methods are relatively limited, generally only supporting CNN or ViT.

[0004] Summary of the Invention

[0005] In order to overcome the shortcomings of the above-mentioned prior art in the training process of visual recognition models, such as the lack of manual guidance, the limited interpretability in aligning abstract semantic concepts with specific image areas, and the relatively limited applicable framework of existing interpretable methods, the present invention provides an interactive semantic perception self-learning framework and an interpretable visual recognition method.

[0006] The primary purpose of the present invention is to solve the above technical problems, and the technical solutions of the present invention are as follows:

[0007] The first aspect of the present invention proposes an interactive semantic-aware self-learning framework, which includes a teacher module and a student module; the student module is used for visual recognition, and the teacher module is used to provide semantic guidance to the student module; the student module inputs the calculated slice-feature pairs into the teacher module, and the teacher module outputs the calculated semantically rich slices to the student module; the teacher module includes a first encoder, a category concept library and a similarity comparison submodule; the encoder is used to extract image features; the category concept library is used to store category concepts; the similarity comparison submodule is used to generate semantically rich slices, i.e., semantic guidance; the student module includes a slice library, a second encoder and a feature selection submodule; the slice library is used to store image slices for classification; the feature selection submodule is used to perform feature selection on the image slice features generated by the encoder.

[0008] Furthermore, the encoder is a pre-trained feature extraction network, which can be CNN, ViT and Swin-Transformer, wherein the second encoder of the student module and the first encoder of the teacher module are the same encoder and have the same network structure, and the teacher module and the student module can share one encoder;

[0009] The category concept library stores a global concept feature of each image category, namely, a category concept, and has a dictionary structure including keys and values. The key is the category, and the value is the global feature vector corresponding to the category.

[0010] A second aspect of the present invention provides an interpretable visual recognition method based on an interactive semantic perception self-learning framework, comprising the following steps:

[0011] S1: Slice the image and input it into the teacher module and student module respectively, and initialize the teacher module and student module;

[0012] S2: The student module performs feature selection on the segmented image to generate segment-feature pairs;

[0013] S3: Use the shard-feature pair to query the corresponding index in the teacher module and generate patch-specific feature responses;

[0014] S4: The teacher module compares the feature response with the index pair and identifies and generates semantically rich slices;

[0015] S5: Update the teacher and student modules using semantically rich slices;

[0016] S6: Image classification using the updated student module.

[0017] Furthermore, in step S1, the image is sliced ​​and inputted into the teacher module and the student module respectively, and the teacher module and the student module are initialized. The specific process is as follows:

[0018] The input image is divided into slices of preset size and then forwarded to the student module and teacher module. The student module stores the received slices in the slice library, and the teacher module inputs the received slices into the encoder for feature extraction and uses the obtained features to update the category concept library.

[0019] Furthermore, the student module in step S2 performs feature selection on the sliced ​​image to generate slice-feature pairs. The specific process is as follows:

[0020] The student module inputs the slices in the shard library into the encoder in the student module for feature extraction, and inputs the obtained features into the feature selection submodule to generate slice-feature pairs;

[0021] First, the attention weight matrix is ​​obtained according to the attention mechanism

[0022] where x i is the original input sample, w ps represents the parameters in the encoder, f ps Represents the process of feature selection for the original input sample, N is the number of segments in the original input sample, Indicates the patch selection process, patches are indexed using i and j, and ps represents patch selection;

[0023] Secondly, the top-k elements with the highest attention weights are identified, 1≤k≤N; finally, features corresponding to the row and column indices of the top k values ​​in the attention matrix are selected from the input image.

[0024] Furthermore, step S3 uses the shard-feature pair to query the corresponding index in the teacher module and generate a patch-specific feature response. The specific process is as follows:

[0025] (1) Use the shard-feature pair to query the corresponding index in the category concept library

[0026] The teacher module maintains a category concept library matrix Each line is the feature vector obtained after the corresponding category passes through the teacher module:

[0027] where c(x i ) represents the sample x i In the matrix Category index in y iIs a one-hot column vector, representing x i Category;

[0028] (2) Generate feature responses using shard-feature pairs

[0029] The teacher module is set to know some parameters of the student; the teacher module accepts the selected features from the feature selection submodule As input, and generate feature responses

[0030] w te is the learning parameter of the teacher module, f t () is the feature response generation process.

[0031] Furthermore, the teacher module in step S4 compares the feature response with the index pair to identify and generate semantically rich slices. The specific process is as follows:

[0032] (1) Similarity comparison

[0033] Construct a matrix PS with b rows and N columns, where b is the input batch size and N is the number of patches in each sample, and use cosine similarity to calculate similarity:

[0034] Among them, PS ij represents the jth patch feature of the i-th sample in the input batch Patch features corresponding to category concepts The similarity of , that is, similar patch features, is the characteristic response The jth element in is the eigenvector The jth element of ;

[0035] (2) Semantic patch optimization

[0036] Further select each sample Similar patch features as semantic patches:

[0037] in represents the new sample after semantic patch optimization, represents the sample before semantic patch optimization, Before The patch features with the highest similarity, represents the semantic patch features after selection, Represents the normalized semantic patch features;

[0038] (3) Sample optimization

[0039] Construct a sample similarity vector S, which is a b-dimensional vector, each element represents the current sample feature in the category concept library Category representation of the current sample feature The similarity between them, where b is the batch size of the input samples, and then select the first The sample features with the highest similarity Sample; After the above optimization, the teacher module produces a patch with semantic Selected useful samples, i.e., semantically rich slices.

[0040] Furthermore, in step S5, the semantically rich slices are used to update the teacher module and the student module. The specific process is as follows:

[0041] (1) Update the category concept library using semantically rich slices

[0042] With semantic patches To update the category concept library according to the momentum mechanism:

[0043] in Represents the category concept library matrix, α is a hyperparameter that balances the weights of the corresponding category concept representations in the current sample and the maintained category concept library when updating the category concept library;

[0044] (2) Update the sharding library using semantically rich slices.

[0045] Furthermore, in step S6, the updated student module is used to classify the images. The specific process is as follows:

[0046] Use the semantically rich slices as input data for the student module to predict the label of the image:

[0047] Among them, w s is the learning parameter of the student module, f s represents the patch selection process, and y′ represents the predicted label.

[0048] Furthermore, in the process of using semantically rich slices as input data of the student module to predict the label of the image, contrastive semantic loss is introduced. The specific process is as follows:

[0049] The new loss function can be expressed as:

[0050] Where β is the weight factor; is the original cross entropy loss, is the contrastive learning loss; it is defined as:

[0051] where z i and z j is the processed image after patch selection, Sim(z i ,z j ) is z i and z j The cosine similarity of ,γ represents (z i ,z j ) similarity threshold, y i 、y j are the labels corresponding to the two contrast samples in the contrast loss.

[0052] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0053] The interactive semantic perception self-learning framework described in the present invention includes a teacher module and a student module; the student module is used for visual recognition, and the teacher module is used to provide semantic guidance to the student module; the student module inputs the calculated slice-feature pairs into the teacher module, and the teacher module outputs the calculated semantically rich slices to the student module. By drawing on the two-way communication mechanism of the human cognitive structure, the interactive learning method is conducive to explainable visual recognition; the teacher module includes an encoder, a category concept library and a similarity comparison submodule; the student module includes a slice library, an encoder and a feature selection submodule, and the encoder can be CNN, ViT and Swin-Transformer, which has good compatibility and scalability. The similarity sub-comparison module performs fine-grained optimization of semantic patches and can obtain patches with more semantic concept information, which will help improve the interpretability of the framework. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] FIG1 is a schematic diagram of an interactive semantic-aware self-learning framework provided in this embodiment.

[0055] FIG2 is a flow chart of an explainable visual recognition method based on an interactive semantic perception self-learning framework provided in this embodiment.

[0056] FIG3 is an experimental effect diagram of an explainable visual recognition method based on an interactive semantic perception self-learning framework provided in this embodiment. DETAILED DESCRIPTION

[0057] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that, in the absence of conflict, the embodiments of the present application and the features therein can be combined with each other.

[0058] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0059] Example 1

[0060] As shown in Figure 1, the first aspect of the present invention proposes an interactive semantic-aware self-learning framework, which includes a teacher module and a student module; the student module is used for visual recognition, and the teacher module is used to provide semantic guidance to the student module; the student module inputs the calculated slice-feature pairs into the teacher module, and the teacher module outputs the calculated semantically rich slices to the student module; the teacher module includes a first encoder, a category concept library and a similarity comparison submodule; the encoder is used to extract image features; the category concept library is used to store category concepts; the similarity comparison submodule is used to generate semantically rich slices, i.e., semantic guidance; the student module includes a slice library, a second encoder and a feature selection submodule; the slice library is used to store image slices for classification; the feature selection submodule is used to perform feature selection on the image slice features generated by the encoder.

[0061] More specifically, the encoder is a pre-trained feature extraction network, which can be CNN, ViT and Swin-Transformer, wherein the second encoder of the student module and the first encoder of the teacher module are the same encoder and have the same network structure, and the teacher module and the student module can share one encoder;

[0062] The category concept library stores a global concept feature of each image category, namely, a category concept, and has a dictionary structure including keys and values. The key is the category, and the value is the global feature vector corresponding to the category.

[0063] As shown in FIG2 , the second aspect of the present invention proposes an interpretable visual recognition method based on an interactive semantic perception self-learning framework, comprising the following steps:

[0064] S1: Slice the image and input it into the teacher module and student module respectively, and initialize the teacher module and student module.

[0065] The input image is divided into slices of size 16×16 and then forwarded to the student module and the teacher module. The student module stores the received slices in the slice library, and the teacher module inputs the received slices into the encoder for feature extraction and uses the obtained features to update the category concept library.

[0066] S2: The student module performs feature selection on the segmented image to generate segment-feature pairs.

[0067] The student module inputs the slices in the shard library into the encoder in the student module for feature extraction, and inputs the obtained features into the feature selection submodule to generate slice-feature pairs;

[0068] The interactive semantic perception self-learning framework is compatible with isotropic and pyramid architectures. The Swin Transformer is used as an example to illustrate, which is a pyramid architecture. The Swin Transformer uses a self-attention mechanism to calculate within a local window. The window is set in a way that the image is evenly divided in a non-overlapping manner. The process of roughly selecting image blocks in the semantic set from the Swin Transformer involves several steps. First, the attention weight matrix is ​​obtained according to the attention mechanism of the Swin Transformer.

[0069] where x i is the original input sample, w ps Represents the parameters in Swin Transformer, f ps Represents the process of feature selection for the original input sample, N is the number of segments in the original input sample, Indicates the patch selection process, patches are indexed using i and j, and ps represents patch selection;

[0070] Secondly, the top-k elements with the highest attention weights are identified, 1≤k≤N; finally, features corresponding to the row and column indices of the top k values ​​in the attention matrix are selected from the input image.

[0071] S3: Use the shard-feature pair to query the corresponding index in the teacher module and generate patch-specific feature responses;

[0072] (1) Use the shard-feature pair to query the corresponding index in the category concept library

[0073] The teacher module maintains a category concept library matrix Each line is the feature vector obtained after the corresponding category passes through the teacher module:

[0074] where c(x i ) represents the sample x i In the matrix Category index in y i Is a one-hot column vector, representing x i Category;

[0075] (2) Generate feature responses using shard-feature pairs

[0076] The teacher module is set to know some parameters of the student; the teacher module accepts the selected features from the feature selection submodule As input, and generate feature responses

[0077] w te is the learning parameter of the teacher module, f t () is the feature response generation process.

[0078] S4: The teacher module compares the feature response with the index pair and identifies and generates semantically rich slices;

[0079] Fine-grained semantic patch optimization is performed through the similarity comparison submodule within the teacher module to obtain patches with more semantic concept information, which will help improve the interpretability of the framework; the similarity comparison submodule within the teacher module is also able to select more important samples for sample optimization.

[0080] (1) Similarity comparison

[0081] Construct a matrix PS with b rows and N columns, where b is the input batch size and N is the number of patches in each sample, and use cosine similarity to calculate similarity:

[0082] Among them, PS ij represents the jth patch feature of the i-th sample in the input batch Patch features corresponding to category concepts The similarity of , that is, similar patch features, is the characteristic response The jth element in is the eigenvector The jth element of ;

[0083] (2) Semantic patch optimization

[0084] Further select each sample Similar patch features as semantic patches:

[0085] in represents the new sample after semantic patch optimization, represents the sample before semantic patch optimization, Before The patch features with the highest similarity, represents the semantic patch features after selection, Represents the normalized semantic patch features;

[0086] (3) Sample optimization

[0087] The process of sample optimization is similar to semantic patch optimization. The sample similarity vector S is constructed, which is a b-dimensional vector. Each element represents the current sample feature in the category concept library. Category representation of the current sample feature The similarity between them, where b is the batch size of the input samples, and then select the first The sample features with the highest similarity Sample; After the above optimization, the teacher module produces a patch with semantic Selected useful samples, i.e., semantically rich slices.

[0088] S5: Update the teacher and student modules using semantically rich slices;

[0089] (1) Update the category concept library using semantically rich slices

[0090] With semantic patches To update the category concept library according to the momentum mechanism:

[0091] in Represents the category concept library matrix, α is a hyperparameter that balances the weights of the corresponding category concept representations in the current sample and the maintained category concept library when updating the category concept library;

[0092] (2) Update the sharding library using semantically rich slices.

[0093] S6: Image classification using the updated student module.

[0094] More specifically, strong interpretability requires strong discriminative capabilities as its foundation. Specifically, we introduce contrastive learning loss to achieve better feature extraction from coarse-grained to fine-grained. The specific process is as follows:

[0095] The new loss function can be expressed as:

[0096] Where β is the weight factor; is the original cross entropy loss, is the contrastive learning loss; it is defined as:

[0097] where z i and z j is the processed image after patch selection, Sim(zi ,z j ) is z i and z j The cosine similarity of ,γ represents (z i ,z j ) similarity threshold, y i 、y j are the labels corresponding to the two contrast samples in the contrast loss.

[0098] Finally, use the semantically rich slices as input data for the student module to predict the label of the image:

[0099] Among them, w s is the learning parameter of the student module, f s represents the patch selection process, y′ represents the predicted label, and the experimental effect diagram is shown in Figure 3.

[0100] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. An interactive semantic-aware self-learning framework, characterized in that: The interactive semantic perception self-learning framework includes a teacher module and a student module; the student module is used for visual recognition, and the teacher module is used to provide semantic guidance to the student module; the student module inputs the calculated slice-feature pairs into the teacher module, and the teacher module outputs the calculated semantically rich slices to the student module; the teacher module includes a first encoder, a category concept library, and a similarity comparison submodule; the encoder is used to extract features of the image; the category concept library is used to store category concepts; the similarity comparison submodule is used to generate semantically rich slices, i.e., semantic guidance; the student module includes a slice library, a second encoder, and a feature selection submodule; The slice library is used to store image slices for classification; the feature selection submodule is used to perform feature selection on the image slice features generated by the encoder.

2. An interactive semantic-aware self-learning framework according to claim 1, characterized in that: The encoder is a pre-trained feature extraction network, which can be CNN, ViT and Swin-Transformer, wherein the second encoder of the student module and the first encoder of the teacher module are the same encoder, have the same network structure, and the teacher module and the student module can share one encoder; The category concept library stores a global concept feature of each image category, namely, the category concept, and its structure is a dictionary structure, including keys and values, wherein the key is the category and the value is the global feature vector corresponding to the category.

3. An interpretable visual recognition method based on an interactive semantic perception self-learning framework, characterized in that: The following steps are involved: S1: Slice the image and input it into the teacher module and student module respectively, and initialize the teacher module and student module; S2: The student module performs feature selection on the segmented image to generate segment-feature pairs; S3: Use the shard-feature pair to query the corresponding index in the teacher module and generate patch-specific feature responses; S4: The teacher module compares the feature response with the index pair and identifies and generates semantically rich slices; S5: Update the teacher and student modules using semantically rich slices; S6: Image classification using the updated student module.

4. The interpretable visual recognition method based on an interactive semantic perception self-learning framework according to claim 3 is characterized in that: In step S1, the image is sliced ​​and input into the teacher module and the student module respectively, and the teacher module and the student module are initialized. The specific process is as follows: The input image is divided into slices of preset size and then forwarded to the student module and the teacher module. The student module stores the received slices in the slice library, and the teacher module inputs the received slices into the encoder for feature extraction and uses the obtained features to update the category concept library.

5. The interpretable visual recognition method based on an interactive semantic perception self-learning framework according to claim 3 is characterized in that: In step S2, the student module performs feature selection on the sliced ​​image to generate a slice-feature pair. The specific process is as follows: The student module inputs the slices in the slice library into the encoder in the student module for feature extraction, and inputs the obtained features into the feature selection submodule to generate slice-feature pairs; First, the attention weight matrix is ​​obtained according to the attention mechanism where x i is the original input sample, w ps represents the parameters in the encoder, f ps It represents the process of feature selection for the original input sample, N is the number of segments in the original input sample, Indicates the patch selection process, patches are indexed using i and j, and ps stands for patch selection; Secondly, the top-k elements with the highest attention weights are identified, 1≤k≤N; finally, features corresponding to the row and column indices of the top k values ​​in the attention matrix are selected from the input image.

6. The interpretable visual recognition method based on an interactive semantic perception self-learning framework according to claim 3 is characterized in that: Step S3 uses the shard-feature pair to query the corresponding index in the teacher module and generate a patch-specific feature response. The specific process is as follows: (1) Use the shard-feature pair to query the corresponding index in the category concept library The teacher module maintains a category concept library matrix Each row is the feature vector obtained after the corresponding category passes through the teacher module: where c(x i ) represents the sample x i In the matrix The category index in y i is a one-hot column vector, representing x i Category of; (2) Generate feature responses using shard-feature pairs The teacher module is set to know some parameters of the student; the teacher module accepts the selected features from the feature selection submodule As input, and generate characteristic responses w te is the learning parameter of the teacher module, f t () is the feature response generation process.

7. The interpretable visual recognition method based on an interactive semantic perception self-learning framework according to claim 3 is characterized in that: In step S4, the teacher module compares the feature response with the index pair to identify and generate semantically rich slices. The specific process is as follows: (1) Similarity comparison Construct a matrix PS with b rows and N columns, where b is the input batch size and N is the number of patches in each sample, and use cosine similarity to calculate similarity: Among them, PS ij represents the jth patch feature of the i-th sample in the input batch Patch features with corresponding category concepts The similarity of , that is, similar patch features, is the characteristic response The jth element in is the eigenvector The jth element of ; (2) Semantic patch optimization Further select each sample Similar patch features as semantic patches: in represents the new sample after semantic patch optimization, represents the sample before semantic patch optimization, Before The patch features with the highest similarity, represents the semantic patch features after selection, Represents the normalized semantic patch features; (3) Sample optimization Construct a sample similarity vector S, which is a b-dimensional vector, each element of which represents the current sample feature in the category concept library The category representation of the current sample feature The similarity between them, where b is the batch size of the input samples, and then select the previous The sample features with the highest similarity Sample; After the above optimization, the teacher module generates a patch with semantic Selected useful samples, i.e., semantically rich slices.

8. The interpretable visual recognition method based on an interactive semantic perception self-learning framework according to claim 3 is characterized in that: Step S5 uses the semantically rich slices to update the teacher module and the student module. The specific process is as follows: (1) Update the category concept library using semantically rich slices With semantic patches To update the category concept library according to the momentum mechanism: in represents the category concept library matrix, α is a hyperparameter that balances the weights of the corresponding category concept representations in the current sample and the maintained category concept library when updating the category concept library; (2) Update the sharding library using semantically rich slices.

9. The interpretable visual recognition method based on an interactive semantic perception self-learning framework according to claim 3 is characterized in that: Step S6 uses the updated student module to classify images. The specific process is as follows: Use the semantically rich slices as input data for the student module to predict the label of the image: Among them, w s is the learning parameter of the student module, f s represents the patch selection process and y′ represents the predicted label.

10. The interpretable visual recognition method based on an interactive semantic perception self-learning framework according to claim 9, characterized in that: In the process of using semantically rich slices as input data of the student module to predict the label of the picture, contrastive semantic loss is introduced. The specific process is as follows: The new loss function can be expressed as: Where β is the weight factor; is the original cross entropy loss, is the contrastive learning loss; defined as: where z i and z j is the processed image after patch selection, Sim(z i ,z j ) is z i and z j The cosine similarity of i ,z j ) similarity threshold, y i ,y j are the labels corresponding to the two contrast samples in the contrast loss.

Citation Information

Patent Citations

  • Semantic segmentation method based on NesT model

    CN116030257A

  • Interactive semantic perception self-learning framework and interpretable visual identification method

    CN117541855A

  • Method for semantic segmentation based on knowledge distillation

    US20210334543A1

Cited By

  • Low-bit-rate semi-reference image quality inspection method and system based on self-distillation

    CN120612552A

  • Aquaculture management decision support method based on multi-scale behavior monitoring and interpretable potential characterization

    CN122222421A