Fine-grained image classification method, device and system, and storage medium

By introducing multi-level structure relational modeling and multi-stage attention area extraction methods, the problem of low accuracy of existing fine-grained image classification methods in complex backgrounds is solved, and classification accuracy and cross-domain adaptability are improved.

CN120510431APending Publication Date: 2025-08-19ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510587952.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Existing fine-grained image classification methods do not perform well when dealing with complex backgrounds or occlusions, and it is difficult to make full use of the spatial relationship information between image blocks, resulting in low classification accuracy.

Method used

Multi-level structural relationship modeling based on Vision Transformer is adopted, combined with multi-stage attention area extraction and contrast learning strategies, a fine-grained classification network model is constructed, and through structural relationship modeling, contrast learning enhancement and multi-stage attention area extraction, model parameters are optimized, and spatial relationship modeling capabilities between image blocks are improved.

Benefits of technology

The accuracy of fine-grained classification is improved, and the generalization ability of the model is enhanced, so that it can adapt to fine-grained classification tasks in different fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510431A_ABST
    Figure CN120510431A_ABST
Patent Text Reader

Abstract

The invention discloses a fine-grained image classification method, device and system, and a storage medium. The method comprises the following steps: constructing a training data set; constructing a fine-grained classification network model based on combination of Vision Transform and a multi-level structural relationship through structural relationship modeling, comparative learning enhancement and multi-stage attention region extraction; training a fine-grained classification network model, and optimizing model parameters through a multi-stage attention mechanism and a contrast learning strategy to obtain a trained fine-grained classification network model; and inputting a to-be-classified image into the trained fine-grained classification network model, and outputting a fine-grained classification result. By adopting the technical scheme of the invention, the feature difference between fine-grained categories is effectively enhanced, and the classification accuracy and generalization ability of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing technology, and in particular relates to a fine-grained image classification method and device, system, and storage medium. Background Art

[0002] Fine-grained image classification (FGIC) is an important research direction in computer vision. It aims to classify subcategories that belong to the same general category but have subtle differences. For example, distinguishing different species of birds, flowers, or vehicles. Traditional convolutional neural networks often have difficulty capturing subtle local features when handling fine-grained classification tasks, resulting in low classification accuracy. In recent years, the Vision Transformer has achieved remarkable results in image classification tasks through its self-attention mechanism, but it still has certain limitations when handling fine-grained classification tasks.

[0003] Existing fine-grained classification methods typically rely on local feature extraction and region localization techniques, but these methods perform poorly when dealing with complex backgrounds or occlusions. Furthermore, existing methods have limited ability to model image structural relationships, making it difficult to fully utilize the spatial relationship information between image patches. Therefore, a fine-grained classification method that effectively combines global features and local structural relationships is urgently needed to improve classification accuracy and generalization. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a fine-grained image classification method and device, system, and storage medium.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A fine-grained image classification method, comprising:

[0007] Step S1, preprocessing fine-grained image data;

[0008] Step S2: constructing a fine-grained classification network model based on Vision Transformer combined with multi-level structural relationships through structural relationship modeling, contrastive learning enhancement, and multi-stage attention region extraction based on the pre-processed fine-grained image data;

[0009] Step S3: Optimize the parameters of the granularity classification network model through a multi-stage attention mechanism and a contrastive learning strategy to obtain a trained fine-grained classification network model;

[0010] Step S4: input the image data to be processed into the trained fine-grained classification network model and output the fine-grained classification results.

[0011] As a preference, the structural relationship modeling in step S2 is specifically as follows: by sending the preprocessed image into VisionTransformer, extracting the attention weight of each layer, adding and averaging the attention weights of the last three layers, and generating an attention weight map; based on the attention weight map, constructing the spatial relationship information between image blocks for GCN through polar coordinates to obtain the structural relationship of the image; combining the structural relationship with the last layer class token of Vision Transformer and sending it to the fully connected layer to obtain the loss function

[0012] As a preference, the contrastive learning enhancement in step S2 is specifically as follows: extracting features from samples of the same category and calculating the similarity between features; enhancing the feature similarity of samples of the same category by contrastive loss function, and obtaining loss function

[0013] As a preference, the multi-stage attention region extraction in step S2 is specifically as follows: extract the important region from the attention weight map, crop and expand the region, and then send it to the Vision Transformer again, repeating the process of extracting the attention weight map and structural relationship to obtain the loss function

[0014] As a preference, in step S3, using and The model parameters are adjusted through the back-propagation algorithm and the optimizer.

[0015] The present invention also provides a fine-grained image classification device, comprising:

[0016] A first processing module, configured to pre-process fine-grained image data;

[0017] The second processing module is used to construct a fine-grained classification network model based on Vision Transformer combined with multi-level structural relationships through structural relationship modeling, contrast learning enhancement and multi-stage attention region extraction according to the pre-processed fine-grained image data;

[0018] The third processing module is used to optimize the parameters of the granular classification network model through a multi-stage attention mechanism and a contrastive learning strategy to obtain a trained fine-grained classification network model;

[0019] The fourth processing module is used to input the image data to be processed into the trained fine-grained classification network model and output the fine-grained classification results.

[0020] The present invention also provides a granularity image classification system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes the granularity image classification method when executed by the processor.

[0021] The present invention also provides a storage medium, on which a computer program is stored, and the computer program executes the granularity image classification method when running.

[0022] The present invention solves the problems of existing fine-grained classification technology, such as poor performance when dealing with complex backgrounds or occlusion situations, limited ability to model image structural relationships, and difficulty in fully utilizing the spatial relationship information between image blocks. By introducing multi-level structural relationships and an optimized Transformer architecture, the accuracy of fine-grained classification is improved, the generalization ability of the model is enhanced, and it can be adapted to different fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0024] Figure 1 This is a flow chart of a fine-grained image classification method according to an embodiment of the present invention;

[0025] Figure 2 A flow chart of generating a global attention graph according to an embodiment of the present invention;

[0026] Figure 3 A flowchart of multi-level structure relationship modeling according to an embodiment of the present invention;

[0027] Figure 4 This is a flow chart of comparative learning according to an embodiment of the present invention;

[0028] Figure 5 This is a flowchart of selecting and cropping important areas according to an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0030] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] Example 1:

[0032] An embodiment of the present invention provides a fine-grained image classification method, comprising the following steps:

[0033] (1) Data preprocessing: First, the CUB-200-2011, NABirds, and Stanford Dogs datasets are partitioned. The partitioned datasets are then divided into training and test sets to obtain the training and test datasets corresponding to the network model. The input images are preprocessed, including image size normalization, data augmentation, and noise removal, to generate image data suitable for model input.

[0034] (2) Structural Relationship Modeling: Spatial Structure Relationship Learning extracts the attention weight of each layer by feeding the preprocessed image into Vision Transformer, adds and averages the attention weights of the last three layers, and generates an attention weight map. Based on the attention weight map, the spatial relationship information between image blocks is constructed for GCN through polar coordinates to obtain the structural relationship of the image. The structural relationship is combined with the last layer class token of Vision Transformer and fed into the fully connected layer to obtain the loss function.

[0035] In order to explore the spatial relationship between all image blocks, polar coordinates are used to represent the node information of the image blocks. First, the block with the largest attention weight is selected as the reference block and used as the origin. Then, the polar coordinates of other image blocks are constructed as the spatial position information, and the origin node information is set as the reference origin. Therefore, the position information is obtained:

[0036] 1: Express the position (x, y) of the image block as polar coordinates (ρ x,y ,θ x,y ), where ρ x,y is the normalized distance from the image patch to the reference patch. x,y is the normalized angle of the image patch relative to the reference patch.

[0037] P x,y =(ρ x,y ,θ x,y )

[0038] 2: Let ρ x,y To normalize the distance, calculate P x,yThe Euclidean distance to the reference block P0, divided by the feature map width N w and height N H Normalize. Where (x0, y0) is the two-dimensional coordinate of the reference block (the block with the largest attention weight), which is used as the polar coordinate origin. (x, y) is the coordinate of the current image block.

[0039]

[0040] 3: Let θ x,y To normalize the angle, calculate P x,y Relative to the direction of the reference block P0, arctan2(y,x) is a four-quadrant inverse tangent function with an output range of (-π,π](-π,π], and the angle is normalized to [0,1] by +π and divided by 2π.

[0041]

[0042] (3) Contrastive learning enhancement: Using contrastive learning strategy, we enhance the feature similarity of samples in the same category and obtain the loss function

[0043] (4) Multi-stage attention region extraction: extract important regions from the attention weight map, crop and expand the region, and then feed it back into the Vision Transformer. Repeat the process of extracting the attention weight map and structural relationship to obtain the loss function.

[0044] First, the pixel average a of the attention weight map is calculated, and the parameter A(x, y) represents the value of the feature point in the attention weight map A. The parameters h and w represent the height and width of the weight map.

[0045] The coefficient τ is used to set the threshold. The optimal τa value is obtained through ablation experiments. When the value A(x,y) in the attention weight map is greater than the threshold, the image patch mask value is set to 1. Otherwise, it is set to 0.

[0046] The region selection is based on the following formula:

[0047] 1: Calculate the average of the attention weights of all positions by summing A(x,y) at different positions and dividing it by h*w Where A(x,y) is the response value at position (x,y) in the attention weight map, reflecting the degree of attention paid by the model to that position. h and w are the height and width of the attention weight map

[0048]

[0049] 2: Increase the attention weight A(x,y) above the threshold The positions of are set to 1 (retained) and the rest are set to 0 (ignored). The resulting binary mask Mark whether to clip the (x,y) position.

[0050]

[0051] Where τ = 0.2.

[0052] (5) Model training: using and The model parameters are adjusted through the back-propagation algorithm and optimizer to optimize the classification performance of the model.

[0053] in, The logarithmic cross entropy loss is used. We leverage semantic relationships for feature enhancement, using contrastive learning to enhance feature similarity within the same category and weaken feature similarity between different categories. To aid model training by mining hard negative pairs, we use a hyperparameter α in the contrastive learning loss. Negative pairs with a lower similarity than positive pairs are filtered out.

[0054] 1: Let S(I) be all image samples in the training set, y be the true label of image I, and pred(I) be the model's predicted probability vector for image I. Finally, the difference between the predicted probability pred(I) and the true label y is calculated according to the above formula.

[0055]

[0056] 2: Let S(I') be the image samples of all training sets S(I) after region cropping, y be the true label of image I', pred(I') be the model's predicted probability vector for image I', and finally calculate the difference between the predicted probability pred(') and the true label y according to the above formula

[0057]

[0058] 3: Calculate the similarity of negative sample pairs Indicator of the difference between the average similarity and the positive sample pair i,j ,like If the average similarity of the negative sample pair is higher than that of the positive sample pair by at least α, the negative sample pair (i.e., Indicator i,j >0), otherwise ignored (Indicaitor i,j =0). is a pair of similar samples (positive sample pairs), y (i) =y (j) . is a heterogeneous sample pair (negative sample pair), y (i) ≠y (j) . sim(·) cosine similarity. α is the threshold, which controls the strictness of negative sample screening. [y (i) =y (j) , i≠j is the number of similar sample pairs in the batch.

[0059]

[0060] 4: First, minimize the distance between similar samples (ideally Loss → 0) to calculate the positive sample loss Then only for difficult negative samples, namely Indicatior i,j >0 maximizes the distance between heterogeneous samples (ideally ) to calculate the negative sample loss Sum the positive sample loss and the negative sample loss and normalize them, that is, divide them by N 2 Balance the impact of different batches, where N 2 is the square of the batch size.

[0061]

[0062] (6) Classification prediction: Input the image data to be processed into the trained model and output the fine-grained classification results.

[0063] As an implementation method of the embodiment of the present invention, the refinement of the data preprocessing in step (1) includes:

[0064] Step 11: Partition the CUB-200-2011, NABirds, and Stanford Dogs datasets into training and test sets according to a preset ratio.

[0065] Step 12: Normalize the size of the input image and adjust the image to a fixed size;

[0066] Step 13: Perform data augmentation operations on the image, including random cropping, rotation, and flipping;

[0067] Step 14: Remove noise from the image and smooth the image using Gaussian filtering or median filtering.

[0068] As an implementation method of an embodiment of the present invention, the implementation method of the structural relationship modeling in step (2) is:

[0069] Step 21: Extract the attention weight of each layer in the Vision Transformer, and add and average the attention weights of the last three layers to generate an attention weight map;

[0070] Step 22: Based on the attention weight map, convert the center point coordinates of the image block into polar coordinates.

[0071] Step 23: Calculate the relative position relationship between image blocks based on polar coordinates and construct the GCN adjacency matrix; use GCN to model the spatial relationship information between image blocks and obtain the structural relationship of the image.

[0072] Step 24: Combine the structural relationship with the last layer class token of Vision Transformer and send it to the fully connected layer to obtain the loss function

[0073] As an implementation method of an embodiment of the present invention, the contrastive learning enhancement in step (3) is implemented as follows:

[0074] Step 31: Extract features from samples of the same category and calculate the similarity between the features;

[0075] Step 32: Enhance the feature similarity of samples of the same category by comparing the loss function to obtain the loss function

[0076] As an implementation method of an embodiment of the present invention, the multi-stage attention area extraction in step (4) is implemented as follows:

[0077] Step 41: Extract the high-weight region from the attention weight map, crop and expand the region;

[0078] Step 42: Send the cropped area back to the Vision Transformer and repeat the process of extracting the attention weight map and structural relationship to obtain the third loss function

[0079] As an implementation method of an embodiment of the present invention, the optimization strategy for model training in step (5) includes:

[0080] Step 51: Utilization and Optimize model parameters;

[0081] Step 52: Adopt a dynamic learning rate adjustment strategy, use a higher learning rate at the beginning of training, and gradually reduce the learning rate to improve training stability;

[0082] Step 53: Use the SGD optimizer to optimize the parameters of Vision Transformer and GCN.

[0083] As an implementation method of an embodiment of the present invention, the classification prediction in step (6) is implemented as follows:

[0084] Step 61: Input the preprocessed image data into the trained model and extract image features through Vision Transformer and GCN;

[0085] Step 62: Combining the structural relationship, contrast learning features and global features, outputting the fine-grained classification result of the image;

[0086] Step 63: Post-process the classification results and output the final classification category and its confidence score.

[0087] The present invention adopts the Vision Transformer architecture and introduces a multi-level structural relationship modeling mechanism to construct a fine-grained image classification model. Specifically, ViT-B / 16 is configured as the basic architecture in the Vision Transformer backbone network, which includes a 12-layer Transformer encoder, which fuses the preprocessed image block sequence with the position code, and adds a learnable class token at the beginning of the sequence to form the input; the encoder layer design includes a 12-head self-attention mechanism and an extended dimension feedforward network (GELU activation). Furthermore, the attention weight matrix is extracted from the last three layers of the Vision Transformer through the structural relationship modeling module to generate a global attention map. The spatial relationship adjacency matrix between image blocks is constructed based on the polar coordinate transformation, the structural relationship is modeled using GCN, and it is spliced with the class token and sent to the classification layer. At the same time, a contrastive learning enhancement module is introduced to project the class token into a low-dimensional space, and the similarity loss between similar samples is calculated to enhance feature discriminability. In addition, a multi-stage attention region extraction mechanism is designed to locate high-weight areas based on the global attention map and crop and enlarge them. The secondary input is Vision Transformer to extract detailed features. Finally, the model's ability to capture fine-grained differences is improved through multi-loss joint optimization.

[0088] The present invention adopts a multi-task joint training strategy to collect fine-grained datasets such as CUB-200-2011, NABirds, and Stanford Dogs, covering complex scene images of various species such as birds and dogs. The model is optimized by a hierarchical loss function: the total loss function is obtained based on the cross-entropy loss, contrastive learning loss, and cropped area classification loss of the structural relationship between the GCN output and the true label. During training, the weights of the ViT-B / 16 model pre-trained on ImageNet-21k are loaded to initialize the encoder, and the SGD optimizer (initial learning rate 0.001) is used. The learning rate is decayed by a multiple, combined with dynamic gradient clipping to prevent overfitting, thereby improving the model's generalization recognition ability for fine-grained features.

[0089] Example 2:

[0090] like Figure 1 As shown, an embodiment of the present invention provides a fine-grained image classification method, including:

[0091] Step 1: Data preprocessing

[0092] (1) Dataset selection and division

[0093] (1.1) Dataset selection

[0094] CUB-200-2011 (11,788 bird images), NABirds (48,562 bird images), and StanfordDogs (20,580 dog images) are selected as benchmark datasets.

[0095] (1.2) Dataset division

[0096] Each dataset is split into training and test sets in a 7:3 ratio. For example, the training set of CUB-200-2011 contains 8,252 images and the test set contains 3,536 images. Ensure that the class distribution is even to avoid data skew.

[0097] (2) Image size normalization

[0098] (2.1) Adjust the size

[0099] All images are uniformly resized to 448×448 pixels using the bilinear interpolation algorithm.

[0100] (2.2) Standardization

[0101] Perform channel normalization on the image, where the mean and variance are pre-computed based on the ImageNet dataset.

[0102] (3) Multimodal feature fusion

[0103] (3.1) Image block division

[0104] The image is divided into 16×16 non-overlapping blocks (784 blocks in total), each block is denoted as P i,j ∈R 16*16*3 .

[0105] (3.2) Positional encoding

[0106] Generate a position embedding vector E for each image patch pos (i,j)∈R d (d=768).

[0107] (3.3) Feature splicing

[0108] The image block feature P i,j With position encoding E pos Concatenate to form the input vector.

[0109] (4) Data loading and batch processing

[0110] (4.1) Batch processing settings

[0111] The batch size is set to 32, and random sampling is used to ensure uniform data distribution.

[0112] (4.2) Memory Mapping

[0113] Use HDF5 format to store preprocessed data to speed up data reading.

[0114] Step 2: Model construction

[0115] (1) Vision Transformer backbone network configuration

[0116] (1.1) Model architecture selection

[0117] ViT-B / 16 is used, which contains a 12-layer Transformer encoder and takes the preprocessed image block sequence as input.

[0118] (1.2) Positional Encoding Fusion

[0119] The image block position encoding is added to the image block features, and a learnable class token is added to the beginning of the sequence to form the final input.

[0120] (1.3) Transformer encoder layer design

[0121] Each layer contains 12 attention heads, two fully connected layers, the activation function is GELU, and the middle dimension is expanded by 4 times

[0122] (2) Spatial Structure Relationship Learning

[0123] (2.1) Attention weight map generation

[0124] like Figure 2 As shown, an embodiment of the present invention provides a method for generating a global attention map as follows:

[0125] Extract the attention weight matrix from the last three layers of Vision Transformer, take the average of the weights of the last three layers, and generate a global attention map.

[0126] (2.2) Establishment of polar coordinate spatial relationship

[0127] like Figure 3 As shown, an embodiment of the present invention provides a method for generating a multi-level structure relationship modeling flow chart as follows:

[0128] The center coordinates of the image block (x i ,y i ) is converted to polar coordinates (γ i ,θ i ), calculate the similarity between blocks based on polar coordinates and generate the GCN adjacency matrix.

[0129] (2.3) Graph convolution operation

[0130] For the image block feature Z patch ∈R 784*768 Perform graph convolution, concatenate the GCN output with the class token of Vision Transformer, and send it to the fully connected layer for classification, and get

[0131] (3) Contrastive Learning Enhancement

[0132] like Figure 4 As shown, the embodiment of the present invention provides a comparative learning method as follows:

[0133] Project the class token of Vision Transformer into the contrastive learning space. In the same batch, for each sample z i , select similar samples z j is a positive pair, and other class samples are negative pairs. The loss is calculated as follows:

[0134]

[0135] (4) Multi-stage attention region extraction

[0136] (4.1) High-weight area positioning

[0137] Through the global attention map generated in (2.1), the vector weights are converted into a spatial heat map, the salient region mask is generated by thresholding, and the bounding box of the largest connected region is extracted.

[0138] (4.2) Secondary feature extraction

[0139] like Figure 5 As shown, an embodiment of the present invention provides a method for selecting and cropping important areas, as follows:

[0140] Map the bounding box back to the original image, crop and upsample it, adjust it to 448×448 pixels and normalize it, input the cropped area into the Vision Transformer, repeat steps 1 to 2, and get

[0141] Step 3: Model training

[0142] (1) Dataset preparation

[0143] We collected CUB-200-2011 (11,788 bird images), NABirds (48,562 bird images), and StanfordDogs (20,580 dog images) as benchmark datasets, covering images of different species, poses, and background complexity.

[0144] (2) Loss weighting

[0145] Structural relationship loss ( ): Cross entropy loss based on the structural features and category labels output by GCN, with a weight coefficient of α = 0.4;

[0146] Contrastive learning loss ( ): scaled contrast loss, weight coefficient β = 0.4;

[0147] Regional feature loss ( ): Cross entropy loss for cropping region classification, weight coefficient γ = 0.2;

[0148] The total loss is:

[0149]

[0150] (3) Optimize training and training configuration

[0151] (3.1) Loading pre-trained model

[0152] Load the pre-trained ViT-B / 16 model weights as encoder initialization parameters.

[0153] (3.2) Optimizer selection

[0154] The SGD optimizer is used with an initial learning rate of 0.001. The learning rate is adjusted at the 15th, 30th, and 50th epochs. Each time the learning rate is adjusted, it is multiplied by 0.1 (i.e., the learning rate decays to 10% of the original value).

[0155] Step 4: Structured Prediction

[0156] (1) Preprocessing and model input

[0157] The trained model feeds the image to be classified into a preprocessed image. The image is resized to 448×448 pixels and split into 16×16 non-overlapping blocks. Each block is concatenated with the positional encoding to form the input sequence, and a learnable class token is added to construct the final input vector.

[0158] (2) Model calculation and prediction result output

[0159] (2.1) Model calculation

[0160] The model extracts global image features through the Vision Transformer encoder, generates an attention weight map, and locates high-weight areas based on the weight mean of the last three layers. It uses Spatial Structure Relationship Learning to model the spatial relationship of the polar coordinates of image blocks and fuses structural features with class tokens. It crops high-weight areas and upsamples them to 448×448 pixels, and then feeds them into the model twice to extract local detail features.

[0161] (2.2) Prediction result output.

[0162] The beneficial effects of the present invention are as follows:

[0163] (1) Improve the accuracy of fine-grained classification: By introducing multi-level structural relationship modeling (Spatial Structure Relationship Learning module) and multi-stage attention region extraction mechanism, it can effectively capture local subtle differences (such as bird feather texture, vehicle wheel features) and global context information. Experiments show that it improves by 5% to 9% compared with traditional CNN methods (such as ResNet-50) and by 1% to 5% compared with existing Vision Transformer methods (such as TransFG).

[0164] (2) Enhanced cross-domain generalization capability: The model has a slightly improved accuracy in cross-dataset testing, and does not require fine-tuning for specific fields. It can be directly applied to fine-grained classification tasks in multiple scenarios such as birds, flowers, and vehicles.

[0165] Example 3:

[0166] An embodiment of the present invention further provides a fine-grained image classification device, comprising:

[0167] A first processing module, configured to pre-process fine-grained image data;

[0168] The second processing module is used to construct a fine-grained classification network model based on Vision Transformer combined with multi-level structural relationships through structural relationship modeling, contrast learning enhancement and multi-stage attention region extraction according to the pre-processed fine-grained image data;

[0169] The third processing module is used to optimize the parameters of the granular classification network model through a multi-stage attention mechanism and a contrastive learning strategy to obtain a trained fine-grained classification network model;

[0170] The fourth processing module is used to input the image data to be processed into the trained fine-grained classification network model and output the fine-grained classification results.

[0171] Example 4:

[0172] The present invention also provides a granularity image classification system, comprising: a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes the granularity image classification method when executed by the processor.

[0173] Example 5:

[0174] The present invention also provides a storage medium, on which a computer program is stored, and the computer program executes the granularity image classification method when running.

[0175] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.

Claims

1. A fine-grained image classification method, characterized in that: include: Step S1, preprocessing fine-grained image data; Step S2: constructing a fine-grained classification network model based on Vision Transformer combined with multi-level structural relationships through structural relationship modeling, contrastive learning enhancement, and multi-stage attention region extraction based on the pre-processed fine-grained image data; Step S3: Optimize the parameters of the granularity classification network model through a multi-stage attention mechanism and a contrastive learning strategy to obtain a trained fine-grained classification network model; Step S4: input the image data to be processed into the trained fine-grained classification network model and output the fine-grained classification results.

2. The fine-grained image classification method according to claim 1, wherein: The structural relationship modeling in step S2 is as follows: the preprocessed image is fed into the Vision Transformer, the attention weight of each layer is extracted, the attention weights of the last three layers are summed and averaged to generate an attention weight map; based on the attention weight map, the spatial relationship information between image blocks is constructed for the GCN through polar coordinates to obtain the structural relationship of the image; the structural relationship is combined with the last layer class token of the Vision Transformer and fed into the fully connected layer to obtain the loss function 3. The fine-grained image classification method according to claim 2, wherein: The contrastive learning enhancement in step S2 is as follows: extract features from samples of the same category and calculate the similarity between features; enhance the feature similarity of samples of the same category by contrastive loss function, and obtain loss function 4. The fine-grained image classification method according to claim 3, wherein: The multi-stage attention region extraction in step S2 is as follows: extract the important region from the attention weight map, crop and expand the region, and then send it to VisionTransformer again, repeating the process of extracting the attention weight map and structural relationship to obtain the loss function 5. The fine-grained image classification method according to claim 4, wherein: In step S3, use and The model parameters are adjusted through the back-propagation algorithm and the optimizer.

6. A fine-grained image classification device, characterized in that: include: A first processing module, configured to pre-process fine-grained image data; The second processing module is used to construct a fine-grained classification network model based on Vision Transformer combined with multi-level structural relationships through structural relationship modeling, contrast learning enhancement and multi-stage attention region extraction according to the pre-processed fine-grained image data; The third processing module is used to optimize the parameters of the granular classification network model through a multi-stage attention mechanism and a contrastive learning strategy to obtain a trained fine-grained classification network model; The fourth processing module is used to input the image data to be processed into the trained fine-grained classification network model and output the fine-grained classification results.

7. A granular image classification system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the granular image classification method according to any one of claims 1 to 5 is executed.

8. A storage medium, characterized in that: The storage medium stores a computer program, which executes the granular image classification method according to any one of claims 1 to 5 when running.