Fine-grained image classification method of poisonous mushrooms based on multi-stage ViT and contrastive learning

By using a multi-stage ViT and contrastive learning method, combined with image overlapping partitioning and pooling layers, the problems of high computational overhead and high manual labeling cost in fine-grained image classification are solved, and efficient and accurate recognition of fine-grained mushroom images is achieved.

CN115527064BActive Publication Date: 2025-09-23HUAQIAO UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211152826.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-09-23
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

Existing fine-grained image classification methods have problems such as high computational overhead, image segmentation destroying the discrimination area, high computational overhead of the self-attention mechanism, and high manual labeling cost, making it difficult to effectively perform fine-grained recognition of mushrooms.

Method used

A multi-stage ViT and contrastive learning method is adopted. Through image overlapping partitioning, embedded sequence encoding and feature encoding, combined with pooling layers and classifiers, the computational overhead is reduced and the recognition accuracy is improved. Image classification is performed using weakly supervised annotation.

Benefits of technology

This method improves the accuracy of fine-grained mushroom image classification while reducing computational overhead, lowers manual labeling costs, and enhances the representation and generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115527064B_ABST
    Figure CN115527064B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a fine-grained image classification method for poisonous mushrooms based on multi-stage ViT and contrastive learning, which relates to the field of image recognition technology. The image classification method includes S1 obtaining an image to be identified. S2 performing overlapping image division based on the image to be identified to obtain a plurality of partially overlapping image blocks. S3 obtaining an embedding sequence based on the plurality of partially overlapping image blocks. S4 inputting the embedding sequence into a pre-trained multi-stage ViT encoder based on pooling for encoding to obtain a feature code of the image to be identified. S5 inputting the feature code into a classifier for classification to obtain a recognition result of the image to be identified. The pre-trained multi-stage ViT encoder based on pooling includes spaced sub-encoders and pooling layers. The sub-encoder includes L layers of transformer blocks for encoding the embedding sequence into a feature map. The pooling layer is configured between the sub-encoders to adjust the spatial size of the feature map. The multi-stage ViT encoder based on pooling can greatly reduce computational overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a fine-grained poisonous mushroom image classification method based on multi-stage ViT and contrastive learning. Background Art

[0002] Image classification is a fundamental task in computer vision, primarily identifying broad categories of objects, such as mushrooms, fish, dogs, and cars. This type of classification is considered coarse-grained. However, in everyday life, more refined classification is required, such as identifying whether a mushroom is a type of Amanita scaly-stemmed Amanita or a type of pufferfish is a type of pufferfish. This type of classification involves identifying fine-grained subcategories.

[0003] The difficulty of fine-grained image classification lies in the fact that different subclasses are extremely similar in shape and appearance, with only slight differences, making them difficult to distinguish; and the same class is prone to classification errors due to factors such as the target's posture and shooting angle.

[0004] Traditional fine-grained image classification methods require object component labeling in image data to train the model and achieve object component localization and feature learning. However, component labeling consumes a huge amount of manpower and is not conducive to the application of fine-grained classification technology.

[0005] Weakly supervised fine-grained image classification methods, which use only image-level annotations, are an effective way to reduce annotation costs. Applying transformer architectures with self-attention mechanisms to computer vision, such as the Vision Transformer (ViT)-based fine-grained image classification method, can improve recognition performance. However, ViT suffers from image partitioning, which destroys the target discrimination area, and the high computational overhead caused by the self-attention mechanism.

[0006] In view of this, the applicant filed this application after studying the existing technology. Summary of the Invention

[0007] The present invention provides a fine-grained image classification method for poisonous mushrooms based on multi-stage ViT and contrastive learning to improve at least one of the above technical problems.

[0008] The embodiment of the present invention provides a fine-grained poisonous mushroom image classification method based on multi-stage ViT and contrastive learning, which includes:

[0009] S1. Obtain an image to be recognized.

[0010] S2. Perform image overlapping division according to the image to be identified to obtain a plurality of partially overlapping image blocks.

[0011] S3. Obtain an embedding sequence based on multiple partially overlapping image blocks.

[0012] S4. Input the embedded sequence into the pre-trained pooling-based multi-stage Vi T encoder for encoding to obtain the feature code of the image to be recognized.

[0013] S5. Input the feature code into the classifier for classification to obtain the recognition result of the image to be recognized.

[0014] The pre-trained pooling-based multi-stage ViT encoder consists of spaced sub-encoders and pooling layers. The sub-encoders consist of L transformer blocks that encode the embedding sequence into feature maps. Pooling layers are placed between the sub-encoders to adjust the spatial size of the feature maps.

[0015] By adopting the above technical solution, the present invention can achieve the following technical effects:

[0016] The pooling-based multi-stage ViT encoder of the embodiment of the present invention can greatly reduce the computational overhead while accurately performing fine-grained image classification, and has great practical significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is a flowchart of the fine-grained image classification method for poisonous mushrooms.

[0019] Figure 2 It is a schematic diagram of image overlapping partitioning.

[0020] Figure 3 This is the network structure diagram of the multi-stage ViT encoder based on pooling.

[0021] Figure 4 This is the network structure diagram of the pooling layer.

[0022] Figure 5 This is the network structure diagram of the transformer block. DETAILED DESCRIPTION

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0024] See also Figures 1 to 5 A first embodiment of the present invention provides a method for fine-grained poisonous mushroom image classification based on multi-stage ViT and contrastive learning, which can be performed by a poisonous mushroom fine-grained image classification device (hereinafter referred to as the image classification device). Specifically, one or more processors in the image classification device implement steps S1 to S5.

[0025] S1. Obtain an image to be recognized.

[0026] It is understood that the poisonous mushroom fine-grained image classification device can be an electronic device with computing capabilities, such as a portable notebook computer, desktop computer, server, smartphone, or tablet computer. The image to be identified is an image stored in the image classification device or an image transmitted to the image classification device via a network.

[0027] Specifically, the image to be recognized is a 3-channel RGB image X, represented by (C, H, W), where C is the number of channels, and the original number of channels of the RGB image is C = 3. (H, W) represents the image resolution.

[0028] Preferably, the resolution of the image to be identified is (224, 224) or (448, 448). (224, 224) is a commonly used resolution for image classification networks. (448, 448) is a commonly used resolution for fine-grained image classification networks.

[0029] S2. Perform image overlapping division according to the image to be identified to obtain a plurality of partially overlapping image blocks.

[0030] Specifically, the embodiment of the present invention uses a sliding window to implement image overlapping division, which can effectively reduce the damage to the discrimination area, thereby further improving the accuracy of learning the discrimination area of ​​the target image.

[0031] like Figure 2As shown. The given image is divided into multiple overlapping image patches using a sliding window method. The size of the sliding window is (P, P), and the sliding step is S (0 < S ≤ P). The initial position of the sliding window is the upper left corner of the image, and the image area (P, P) selected by the sliding window frame is an image patch. The sliding window first slides step by step in the horizontal direction with a step of S to the right boundary of the picture, then returns to the left boundary, slides down S in the vertical direction, and then slides step by step in the horizontal direction, repeating in a cycle until it reaches the lower right corner of the image.

[0032] The overlapping area of two adjacent image patches is expressed as P * (P - S).

[0033] After sliding, a total of N patches are obtained. The calculation formula for N is as follows:

[0034] N = N H *N W

[0035]

[0036]

[0037] Based on the above embodiments, in an optional embodiment of the present invention, if the sliding window is implemented by 2D convolution, then step S2 is specifically as follows:

[0038] According to the image to be recognized, perform 2D convolution with a convolution kernel of (P, P) and a stride of S to obtain a three-dimensional block embedding. Where 0 < S ≤ P.

[0039] In the embodiments of the present invention, the sliding window is implemented by 2D convolution. The input channel is 3, the output channel is D, the convolution kernel is (P, P), and the stride is S. Perform 2D convolution on the input image to obtain a 3D patch embedding, denoted as (D, N H , N W ). Preferably, in the current step, the output channel D is 256.

[0040] S3. Obtain an embedding sequence according to multiple partially overlapping image patches.

[0041] Specifically, the image needs to be converted into an embedding sequence that can be recognized and operated by a computer before it can be input into a neural network for operation.

[0042] Based on the above embodiments, in an optional embodiment of the present invention, step S3 specifically includes steps S31 to S33.

[0043] S31. Add the three-dimensional block embedding and the three-dimensional position embedding with the same size to obtain a new three-dimensional block embedding.

[0044] S32: transform the new three-dimensional block embedding into a two-dimensional block embedding to obtain a block embedding sequence.

[0045] S33. Concatenate the block embedding sequence and the classification representation vector with the same number of channels to obtain an embedding sequence.

[0046] Specifically, to preserve the position information of the image patch, a learnable position embedding with the same parameters as the block embedding size is designed, and the original block embedding and the position embedding are added together to obtain the new block embedding. Subsequently, a 3D to 2D transformation is performed, i.e. (D, N H ,N W )→(D,N H *N W ), and get the block embedding sequence. Finally, use the learnable classification token vector of size (D, 1) Concatenate with the block embedding sequence to construct the embedding sequence Z0 and input it into the transformer encoder.

[0047] It is understood that the specific parameters of position embedding can be updated based on the parameters of the sliding window and the parameters of the image to be recognized. When both are fixed parameters, the parameters of position embedding can also be fixed. The present invention does not limit the specific values ​​of position embedding. 3D block embedding and 2D block embedding have different dimensions. The conversion process between the two is common knowledge in the art and will not be further described in this invention.

[0048] S4. Input the embedded sequence into a pre-trained pooling-based multi-stage ViT encoder for encoding to obtain a feature code of the image to be recognized. The pre-trained pooling-based multi-stage ViT encoder includes spaced sub-encoders and pooling layers. The sub-encoders include L transformer blocks to encode the embedded sequence into a feature map. Pooling layers are placed between the sub-encoders to adjust the spatial size of the feature map.

[0049] Preferably, the number of sub-encoders is 3. The number of pooling layers is 2. The three sub-encoders and the two pooling layers are spaced apart to form a three-stage ViT encoder. Optionally, the number of transformer blocks in the three sub-encoders is 3, 6, and 4, respectively. The 2D convolution output channels D of the three stages are 256, 512, and 1024, respectively. The output spatial size of the three stages is halved successively.

[0050] In this embodiment, two pooling layers are inserted into an encoder containing L layers of transformer blocks, forming a three-stage hierarchical ViT encoder. The number of layers in each stage is {3, 6, 4}. Specifically, the hierarchical encoder can increase the model's representational power and generalization capabilities.

[0051] Furthermore, pooling layers are added to form a hierarchical structure to transform the spatial scale of feature maps and learn hierarchical features of images. Due to the computational properties of the multi-head self-attention mechanism, operations on feature maps of different spatial scales require only minimal computational overhead, which has great practical significance.

[0052] The network structure of transformer block is as follows Figure 5 As shown in the figure, a transformer block consists of two sets of layer normalization (LN), a set of multiheaded self-attention (MHSA), two residual connections, and a set of multi-layer perceptrons (MLP). It is understood that the transformer block encoder is a state-of-the-art technology and will not be described in detail in this invention. The embedded sequence is encoded by the transformer block to obtain the feature map.

[0053] The transformer block's multi-head self-attention performs multiple self-attention operations on the input embedding sequence, calculating the similarity between the query and key for each patch. A softmax is performed to obtain an attention weight matrix representing the similarity between patches. This weighted sum is then added to the value to produce the output of each self-attention module. After concatenation, a linear transformation is performed to obtain the final feature encoding output of the current block's multi-head self-attention. The calculation process is as follows:

[0054] Z′ l =MSA(LN(Z l-1 ))+Z l-1 ,l=1…L

[0055] z l =MLP(LN(′Z l ))+z′ l ,l=1…L

[0056] It should be noted that if Figure 2As shown, in the three-stage sub-encoder, the number of transformer blocks is 3 / 6 / 4, respectively, i.e., depth = 3 / 6 / 4. For each stage, base_dims = [64, 64, 64] and heads = [4, 8, 16]. base_dims is the baseline hidden layer dimension, and heads refers to the transformer's unique multi-head self-attention parallel operation mechanism. This ensures that the output of the attention layer contains encoded representations in different subspaces, thereby enhancing the model's expressive power. For example, if the hidden layer dimension is 256, it is divided into four heads, i.e., four subspaces, each with a dimension of 64.

[0057] S5. Input the feature code into the classifier for classification to obtain the recognition result of the image to be recognized.

[0058] Specifically, the multi-stage ViT encoder output feature classification head based on pooling It is passed into the classifier to obtain the classification prediction label y′, thereby obtaining the recognition result of the image to be recognized.

[0059] It is understandable that the classifier is an existing classifier, and the present invention does not make specific limitations on this. Preferably, in this embodiment, the classifier is a linear classifier.

[0060] self.head=n.Linear(base_dims[-1]*heads[-1],num_classes)

[0061] Where num_classes is the number of categories in the dataset, base_dims[-1]*heads[-1] is the dimension of the last layer of the feature extraction layer, which is 1024 in this embodiment.

[0062] Based on the above embodiment, in an optional embodiment of the present invention, the classifier is trained using a combination of contrast loss and cross entropy loss as a loss function. The expression of the loss function L is:

[0063] L=L con (Z)+L cross (y,y′)

[0064] Where, L con (Z) represents contrast loss, L cross (y,y′) represents the cross entropy loss.

[0065] Specifically, to address the problem of small differences between subclasses and large differences within classes in fine-grained image classification, this paper combines contrastive feature learning with contrastive loss to maximize the differences between different classes (i.e., different class labels) and minimize the differences between the same class (i.e., same class labels) in order to better supervise the feature learning of the model. The contrastive loss is calculated for N samples using the following formula:

[0066]

[0067] Where, F i ,F j is the classification head after L2 normalization, Sim(F i ,F j ) indicates F i ,F j The cosine similarity between them is calculated, and a threshold of 0.4 is set to filter out simple negative sample pairs with similarity less than 0.4.

[0068] The contrast loss and cross entropy loss are combined as the loss function for training the classifier. The loss function of this fine-grained image classification method is expressed as:

[0069] L=L con (Z)+L cross (y,y′)

[0070] In the training phase, the embodiment of the present invention only requires image-level annotation, which belongs to weakly supervised fine-grained image classification, avoids the huge cost of professional manual annotation, and is beneficial to practical application needs.

[0071] By using a sliding window to overlap and partition the image, the team avoids the damage to the fine-grained image discriminative regions caused by direct segmentation, which helps the self-attention model learn the features of the discriminative regions and achieves higher classification accuracy. In prior art, direct image segmentation destroys the discriminative regions of the target, affecting the model's ability to learn the features of important target regions and resulting in insufficient classification accuracy.

[0072] By inserting pooling layers between sub-encoders, the encoder is divided into multiple stages, which helps increase the model's representational power and generalization capabilities. Furthermore, pooling layers reduce the spatial size of feature maps, enabling the model to learn the hierarchical features of the image and addressing the increased computational overhead associated with overlapping image partitioning. In prior art, feature maps maintained the same spatial size during model training, failing to fully represent the image. This reduced classification accuracy and incurred significant computational overhead.

[0073] In addition, contrastive loss is introduced to enable the model to learn detailed features of different subclasses and similar features of the same category, thereby strengthening the model's feature learning of the image discrimination area, which can greatly improve the classification performance.

[0074] Based on the above embodiment, in an optional embodiment of the present invention, a pooling layer is used to perform steps A1 to A5. Specifically, the addition of the pooling layer doubles the number of feature map channels and halves the spatial size. The operation process is as follows: Figure 4 As shown in Figure 3, the core of the pooling layer is the depth-wise convolution operation.

[0075] A1. Split the feature map output by the previous sub-encoder into classification representation and two-dimensional spatial representation.

[0076] It can be understood that in step S33, the classification representation is added to the block embedding sequence, and in this step, the two are transformed separately.

[0077] A2. Transform the spatial representation into a 3D tensor, and then perform depth-wise convolution to obtain a new 3D tensor with reduced size.

[0078] Specifically, step A2 includes step A21 and step A22. A21, transform the spatial representation into a 3D tensor. The size of the 3D tensor is (D, N H ,N W ), where D is the number of channels, N H and N W A22. Based on the 3D tensor, a new 3D tensor is obtained by depth-wise convolution operation with input channel D, output channel 2D, convolution kernel (3,3), and stride 2. The size of the new 3D tensor is

[0079] A3. Transform the new 3D tensor into a new spatial representation.

[0080] Specifically, step A3 includes steps A31 and A32. A31: Add the new 3D tensor to the position embedding of the same size to obtain a new 3D tensor with position information. A32: Transform the new 3D tensor with position information into a new spatial representation.

[0081] A4. Adjust the classification representation to a new classification representation with the same dimension as the new spatial representation.

[0082] Specifically, step A4 includes obtaining a new classification representation having the same number of channels as the new spatial representation through a fully connected layer based on the classification representation.

[0083] A5. Concatenate the new spatial representation and the new classification representation to obtain a new feature map, which is used to input the next sub-encoder.

[0084] In this embodiment, the pooling layer first converts the 2D features obtained by encoding in the previous stage into Split into spatial tokens and a classification token, then transform the spatial tokens into a 3D tensor, perform a depth-wise convolution operation with input channel D, output channel 2D, convolution kernel (3,3), and stride 2, to realize the 3D tensor from Then, transform the new 3D tensor into a size of At the same time, a fully connected layer is used to adjust the classification tokens to match the spatial tokens dimension (2D, 1). Finally, the spatial tokens and classification tokens are reconstructed to form a new embedding sequence This is input to the encoder of the next stage. The above is the complete pooling layer operation process, which increases the number of channels of the embedded sequence and reduces the spatial size.

[0085] Specifically, a pooling layer is added to construct a multi-stage encoder. The pooling layer uses depth-wise convolution operations to achieve embedding sequence dimension transformation, enabling the model to learn the hierarchical features of the image while effectively reducing computational overhead.

[0086] In the several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or actions, or can be implemented using a combination of dedicated hardware and computer instructions.

[0087] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0088] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, electronic device, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk. It should be noted that, in this article, the terms "include", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements inherent to such a process, method, article or device. Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.

[0089] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.

[0090] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0091] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0092] The "first" and "second" mentioned in the embodiments are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the specific order or precedence of "first" and "second" can be interchanged where appropriate. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.

[0093] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A fine-grained poisonous mushroom image classification method based on multi-stage ViT and contrastive learning, characterized by: Include: Obtain the image to be recognized; Performing image overlapping division according to the image to be identified to obtain a plurality of partially overlapping image blocks; Acquire an embedding sequence according to the plurality of partially overlapping image blocks; Inputting the embedded sequence into a pre-trained pooling-based multi-stage ViT encoder for encoding to obtain the feature code of the image to be recognized; Inputting the feature code into a classifier for classification to obtain a recognition result of the image to be recognized; The pre-trained pooling-based multi-stage ViT encoder comprises spaced sub-encoders and pooling layers; the sub-encoders comprise L-layer transformer blocks for encoding the embedding sequence into feature maps; the pooling layers are arranged between the sub-encoders to adjust the spatial size of the feature maps; The pooling layer is used to: Split the feature map output by the previous sub-encoder into a classification representation and a two-dimensional spatial representation; Transform the spatial representation into a 3D tensor, and then perform depth-wise convolution to obtain a new 3D tensor with reduced size; transforming the new 3D tensor into a new spatial representation; Adjusting the classification representation to a new classification representation with the same dimension as the new spatial representation; The new spatial representation and the new classification representation are concatenated to obtain a new feature map; wherein the new feature map is used to input a subsequent sub-encoder.

2. The method for fine-grained poisonous mushroom image classification based on multi-stage ViT and contrastive learning according to claim 1 is characterized in that: The number of the sub-encoders is 3; the number of the pooling layers is 2; the 3 sub-encoders and the 2 pooling layers are arranged at intervals to form a three-stage ViT encoder.

3. The method for fine-grained poisonous mushroom image classification based on multi-stage ViT and contrastive learning according to claim 2 is characterized in that: The number of transformer block layers of the three sub-encoders are 3, 6, and 4 respectively.

4. The method for fine-grained poisonous mushroom image classification based on multi-stage ViT and contrastive learning according to claim 1 is characterized in that: The spatial representation is transformed into a 3D tensor, and then a new 3D tensor with reduced size is obtained through depth-wise convolution, specifically including: The spatial representation is transformed into a 3D tensor; wherein the size of the 3D tensor is , where D is the number of channels, and is the resolution; According to the 3D tensor, the input channel is D, the output channel is 2D, and the convolution kernel is , the stride is The depth-wise convolution operation is performed to obtain the new 3D tensor; wherein the size of the new 3D tensor is ; Transforming the new 3D tensor into a new spatial representation specifically includes: Adding the new 3D tensor to the position embedding of the same size to obtain a new 3D tensor with position information; Transforming the new 3D tensor with position information into the new spatial representation; Adjusting the classification representation to a new classification representation with the same dimension as the new spatial representation specifically includes: According to the classification representation, a new classification representation having the same number of channels as the new spatial representation is obtained through a fully connected layer.

5. The method for fine-grained poisonous mushroom image classification based on multi-stage ViT and contrastive learning according to claim 1 is characterized in that: Performing image overlapping division according to the image to be identified to obtain a plurality of partially overlapping image blocks specifically includes: According to the image to be identified, the convolution kernel is Perform 2D convolution with a stride of S to obtain a three-dimensional block embedding; where .

6. The method for fine-grained poisonous mushroom image classification based on multi-stage ViT and contrastive learning according to claim 5 is characterized in that: Acquiring an embedding sequence according to the plurality of partially overlapping image blocks specifically includes: Adding the 3D block embedding to a 3D position embedding of the same size to obtain a new 3D block embedding; transforming the new three-dimensional block embedding into two dimensions to obtain a block embedding sequence; The block embedding sequence is concatenated with a classification representation vector having the same number of channels as the block embedding sequence to obtain the embedding sequence.

7. The method for fine-grained poisonous mushroom image classification based on multi-stage ViT and contrastive learning according to any one of claims 1 to 6, characterized in that: The classifier is trained using a combination of contrastive loss and cross entropy loss as the loss function; Loss Function The expression is: , where represents the contrast loss, represents the cross entropy loss.