Image classification method and device and electronic equipment

By segmenting and cascading feature extraction of remote sensing images and using attention units to process image classification models, the problem of time-consuming, labor-intensive, and inaccurate traditional chlorosis identification methods is solved, achieving efficient and automated chlorosis identification.

CN121438084APending Publication Date: 2026-01-30HAINAN AEROSPACE INFORMATION RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511021650.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Traditional methods for identifying yellowing disease rely on manual field surveys and laboratory diagnoses, which are time-consuming, labor-intensive, and inaccurate. They also lack automation and intelligence, limiting their widespread adoption and efficiency in agricultural production.

Method used

An image classification method is used to segment remote sensing images into blocks. A cascaded feature extraction module and attention unit are used to enhance the ability to extract subtle features. The image classification model outputs the chlorosis category, reducing manual intervention and complex parameter settings.

Benefits of technology

It improves the accuracy and robustness of chlorosis identification, enhances the ability to capture image features, and improves the accuracy and practicality of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121438084A_ABST
    Figure CN121438084A_ABST
Patent Text Reader

Abstract

The invention provides an image classification method and device and electronic equipment, and can be applied to the technical field of image processing. The method comprises the following steps: partitioning a to-be-classified remote sensing image to obtain a plurality of to-be-classified remote sensing image blocks; a plurality of to-be-classified remote sensing image blocks are input into M cascaded feature extraction modules in the image classification model, an Mth feature map is output, an mth feature extraction module comprises a non-mth channel attention unit and an mth channel attention unit, under the condition that m is larger than 1 and smaller than or equal to M, the mth feature map is obtained by processing an mth intermediate feature map through the mth channel attention unit, and m is larger than 1 and smaller than or equal to M; the mth intermediate feature map is obtained by processing the (m-1) th feature map by using a non-mth channel attention unit, and under the condition that m is equal to 1, the first feature map is obtained by processing a plurality of remote sensing image blocks to be classified by using a first feature extraction module; and inputting the Mth feature map into a classification module in the image classification model, and outputting an image classification result of the to-be-classified remote sensing image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to an image classification method and device and electronic equipment. BACKGROUND

[0002] Traditional recognition methods of yellowing disease mainly rely on manual field investigation and laboratory diagnosis means. However, manual field investigation has limitations such as time-consuming, laborious, strong subjectivity, and low accuracy. Laboratory diagnosis is restricted in popularization and application in agricultural production practice due to the complexity of yellowing disease recognition. Therefore, a method for improving the recognition accuracy of yellowing disease is urgently needed. SUMMARY

[0003] In view of the above problems, the present application provides an image classification method, device and electronic equipment.

[0004] According to a first aspect of the present application, an image classification method is provided, comprising: dividing a to-be-classified remote sensing image into blocks to obtain a plurality of to-be-classified remote sensing image blocks; inputting the plurality of to-be-classified remote sensing image blocks into M feature extraction modules in cascade in an image classification model to output an Mth feature map, wherein the mth feature extraction module includes a non-mth channel attention unit and an mth channel attention unit, in the case of 1

[0005] The second aspect of the present application provides an image classification device, comprising: a blocking module configured to block a remote sensing image to be classified to obtain a plurality of remote sensing image blocks to be classified; an Mth feature map output module configured to input the plurality of remote sensing image blocks to be classified into M feature extraction modules cascaded in an image classification model to output an Mth feature map, wherein the Mth feature extraction module comprises a non-Mth channel attention layer and an Mth channel attention layer, in the case of 1

[0006] The third aspect of the present application provides an electronic device, comprising: one or more processors; a memory configured to store one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method.

[0007] According to the image classification method, device and electronic device provided by the present application, by blocking the remote sensing image to be classified, a plurality of remote sensing image blocks to be classified can be obtained, the plurality of remote sensing image blocks to be classified are input into M feature extraction modules cascaded in the image classification model, and an Mth feature map can be output. The Mth feature extraction module can comprise a non-Mth channel attention unit and an Mth channel attention unit. In the case of 1 BRIEF DESCRIPTION OF DRAWINGS

[0008] The above and other objects, features and advantages of the present application will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:

[0009] Figure 1 An application scenario diagram of the image classification method according to an embodiment of the present application is shown.

[0010] Figure 2 A flow chart of an image classification method according to an embodiment of the present application is shown.

[0011] Figure 3 A network structure diagram of an image classification model according to an embodiment of the present application is shown.

[0012] Figure 4 A network structure diagram of an m-th channel attention unit according to an embodiment of the present application is shown.

[0013] Figure 5 A network structure diagram of a non-m-th channel attention unit according to an embodiment of the present application is shown.

[0014] Figure 6 A network structure diagram of an image generation model according to an embodiment of the present application is shown.

[0015] Figure 7 A structure block diagram of an image classification apparatus according to an embodiment of the present application is shown.

[0016] Figure 8 A block diagram of an electronic device suitable for implementing an image classification method according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0017] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It should be understood, however, that the description which follows is merely exemplary and is not intended to limit the scope of the application. In the following detailed description of embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that one or more embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present application.

[0018] The terms used herein are merely used to describe specific embodiments and are not intended to limit the present application. The terms "include", "comprise", and the like used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0019] All terms used herein, including technical and scientific terms, have the same meanings as those generally understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having meanings consistent with the context of the present specification, and should not be interpreted in an idealized or overly formal manner.

[0020] In the case of using expressions such as "at least one of A, B, and C", it generally means to include at least one of A, at least one of B, at least one of C, or a combination of at least one of A, at least one of B, and at least one of C (for example, "a system having at least one of A, B, and C" shall include, but not be limited to, a system having A alone, a system having B alone, a system having C alone, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having A, B, and C together etc.).

[0021] The traditional identification method of yellowing disease mainly relies on manual field investigation and laboratory diagnosis. However, manual field investigation has limitations such as time-consuming, labor-intensive, strong subjectivity, and low accuracy. Laboratory diagnosis is restricted in popularization and application in agricultural production practice due to the complexity of yellowing disease identification. Moreover, the existing method lacks sufficient automation and intelligence, and manual selection of feature variables and complex parameter setting are required when the external environment changes, which limits the practicality and efficiency of the technology in large-scale investigation. Therefore, an automatic yellowing disease identification method is urgently needed, which can not only solve the problems of time-consuming and labor-intensive in the identification process, but also improve the accuracy of yellowing disease identification.

[0022] Therefore, embodiments of the present application provide an image classification method, which blocks a to-be-classified remote sensing image to obtain a plurality of to-be-classified remote sensing image blocks; inputs the plurality of to-be-classified remote sensing image blocks into M feature extraction modules in series connection of an image classification model, and outputs an Mth feature map, wherein the mth feature extraction module includes a non-mth channel attention unit and an mth channel attention unit, in the case of 1

[0023] It should be noted that the image classification method and the image classification device provided by the present application can be used in the field of image processing technology, and can also be used in any field other than the field of image processing technology, ecological environment protection, and pest monitoring. Therefore, the application field of the image classification method and the image classification device provided by the present application is not limited.

[0024] Figure 1 An application scenario diagram of the image classification method according to an embodiment of the present application is shown.

[0025] As Figure 1As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0026] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication terminal applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email terminals, social media platform software, etc. (for example only).

[0027] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0028] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0029] It should be noted that the image classification method provided in this application embodiment can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. The image classification device provided in this application embodiment can generally be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0030] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0031] The following will be based on Figure 1 The described scene, through Figures 2-6 The image classification method according to the embodiments of this application will be described in detail.

[0032] Figure 2A flowchart of an image classification method according to an embodiment of the present application is shown.

[0033] As shown in the figure, Figure 2 The image classification method 200 of this embodiment includes operations S210-S230.

[0034] In operation S210, the to-be-classified remote sensing image is blocked to obtain a plurality of to-be-classified remote sensing image blocks.

[0035] In operation S220, the plurality of to-be-classified remote sensing image blocks are input into the M feature extraction modules cascaded in the image classification model to output the Mth feature map.

[0036] In operation S230, the Mth feature map is input into the classification module in the image classification model to output the image classification result of the to-be-classified remote sensing image.

[0037] The to-be-classified remote sensing image can represent a remote sensing image of a plant that can suffer from yellowing disease. For example, areca, Chinese rose, green ivy, etc., but not limited thereto, and the embodiments of the present application do not limit the to-be-classified remote sensing image. The to-be-classified remote sensing image is blocked to obtain a plurality of to-be-classified remote sensing image blocks.

[0038] The mth feature extraction module can include a non-mth channel attention unit and an mth channel attention unit, in the case of 1

[0039] The image classification result indicates the yellowing disease class of the to-be-classified remote sensing image. The yellowing disease class can be determined by the percentage of leaf yellowing area to the entire plant canopy area. The yellowing disease class can be divided into three levels: healthy (<1%), mild (1-10%), and severe (≥10%). The Mth feature map is input into the classification module in the image classification model to output the image classification result of the to-be-classified remote sensing image.

[0040] The image classification model in the embodiments of the present application can realize end-to-end learning, directly learning from the to-be-classified remote sensing image to the final classification result without manual feature extraction, reducing the need for manual intervention and complex parameter settings, and improving the practicality and efficiency.

[0041] Figure 3A network structure diagram of an image classification model according to an embodiment of this application is shown.

[0042] like Figure 3 As shown, multiple remote sensing image patches 310 to be classified are input into M cascaded feature extraction modules in the image classification model, which can output the Mth feature map. The first feature extraction module may include non-first channel attention unit 321 and first channel attention unit 322. The m-th feature extraction module may include non-m-th channel attention unit 323 and m-th channel attention unit 324. The M-th feature extraction module may include non-M-th channel attention unit 325 and M-th channel attention unit 326. Processing the (M-1)-th feature map using the non-M-th channel attention unit 325 yields the M-th intermediate feature map. Processing the M-th intermediate feature map using the M-th channel attention unit 326 yields the M-th feature map. Inputting the M-th feature map into the classification module 330 of the image classification model outputs the image classification result of the remote sensing image to be classified.

[0043] According to embodiments of this application, by dividing the remote sensing image to be classified into blocks, multiple remote sensing image blocks to be classified can be obtained. These multiple blocks are then input into M cascaded feature extraction modules in an image classification model, outputting the Mth feature map. The mth feature extraction module can include non-mth channel attention units and the mth channel attention unit. When 1 < m ≤ M, processing the (m-1)th feature map using the non-mth channel attention unit yields the mth intermediate feature map, enhancing the extraction capability of subtle features in the remote sensing image to be classified. Processing the mth intermediate feature map using the mth channel attention unit yields the mth feature map, reducing inter-channel differences and improving the contrast of the feature map, thereby strengthening the ability to capture image features. Inputting the Mth feature map into the classification module of the image classification model outputs the image classification result of the remote sensing image to be classified, thus improving the accuracy of chlorosis identification.

[0044] According to an embodiment of this application, the m-th feature map is obtained as follows: Average pooling is performed on the m-th intermediate feature map to obtain the m-th average pooling feature vector; the m-th average pooling feature vector is processed using the m-th first multilayer perceptron layer included in the m-th channel attention unit to obtain the m-th first intermediate feature vector; max pooling is performed on the m-th intermediate feature map to obtain the m-th max pooling feature vector; the m-th max pooling feature vector is processed using the m-th second multilayer perceptron layer included in the m-th channel attention unit to obtain the m-th second intermediate feature vector; the m-th average pooling feature vector and the m-th max pooling feature vector are processed using an activation function to obtain the m-th fused feature vector; and the m-th feature map is obtained based on the m-th fused feature vector and the m-th intermediate feature map.

[0045] Figure 4A network structure diagram of the m-th channel attention unit according to an embodiment of this application is shown.

[0046] like Figure 4 As shown, in the m-th channel attention unit 324, the m-th intermediate feature map 410 can be averaged by pooling 420 to obtain the m-th averaged pooling feature vector; the m-th averaged pooling feature vector can be processed by the m-th first multilayer perceptron layer 430 included in the m-th channel attention unit 324 to obtain the m-th first intermediate feature vector; the m-th intermediate feature map 410 can be max-pooled by max-pooling 440 to obtain the m-th max-pooling feature vector; the m-th second multilayer perceptron layer 450 included in the m-th channel attention unit 324 can be processed by max-pooling feature vector to obtain the m-th second intermediate feature vector; the m-th averaged pooling feature vector and the m-th max-pooling feature vector can be processed by activation function 460 to obtain the m-th fused feature vector; and the m-th feature map can be obtained based on the m-th fused feature vector and the m-th intermediate feature map.

[0047] The m-th channel attention unit is constructed based on the Channel Attention (CA) mechanism. CA is used to extract the correlation between different channels, improving the expressive power of the feature map, such as... Figure 4 As shown, the expression is as shown in Formula 1.

[0048] (1);

[0049] in, This is the m-th intermediate feature map; This is the m-th feature map; Let m be the weight matrix of the first multilayer perceptron layer. Let m be the weight matrix of the second multilayer perceptron layer; It is the Sigmoid activation function; AP is average pooling; MP is max pooling; This represents matrix multiplication.

[0050] In complex agricultural environments, crop characteristics may change due to factors such as light and atmosphere. The added channel attention mechanism can further enhance the model's attention to different feature channels, enabling the model to extract and utilize key features more effectively and improve classification accuracy.

[0051] According to an embodiment of this application, the m-th intermediate feature map is obtained as follows: The (m-1)-th feature map is processed using a first m-th normalized attention layer included in a non-m-th channel attention unit to obtain the (m-1)-th normalized feature map, wherein the first m-th normalized attention layer includes a cascaded first m-th instance normalization sub-layer and a first m-th window-based multi-head self-attention sub-layer; the (m-1)-th normalized feature map is processed using a first m-th normalized perceptron layer included in a non-m-th channel attention unit to obtain the (m-1)-th normalized fused feature map, wherein the first m-th normalized perceptron layer includes a cascaded second m-th instance normalization sub-layer and a... The first m-th multilayer perceptron sublayer; the (m-1)-th normalized fusion feature map is processed using the second m-th normalized attention layer included in the non-m-th channel attention unit to obtain the (m-1)-th attention fusion feature map, wherein the second m-th normalized attention layer includes a cascaded third m-th instance normalization sublayer and a second m-th window-based multi-head self-attention sublayer; the (m-1)-th attention fusion feature map is processed using the second m-th normalized perceptron layer included in the non-m-th channel attention unit to obtain the m-th intermediate feature map, wherein the second m-th normalized perceptron layer includes a cascaded fourth m-th instance normalization sublayer and a second m-th multilayer perceptron sublayer.

[0052] Figure 5 A network structure diagram of a non-m-th channel attention unit according to an embodiment of this application is shown.

[0053] like Figure 5 As shown, the (m-1)th feature map can be processed using the first m-th normalized attention layer 510 included in the non-m-th channel attention unit to obtain the (m-1)th normalized feature map. The first m-th normalized attention layer 510 may include a cascaded first m-th instance normalization sub-layer 511 and a first m-th window-based multi-head self-attention sub-layer 512. The (m-1)th normalized feature map can be processed using the first m-th normalized perceptron layer 520 included in the non-m-th channel attention unit to obtain the (m-1)th normalized fused feature map. The first m-th normalized perceptron layer 520 may include a cascaded second m-th instance normalization sub-layer 521 and a first m-th multilayer perceptron sub-layer 522. The (m-1)th normalized fusion feature map can be obtained by processing the second m-th normalized attention layer 530 included in the non-m-th channel attention unit. The second m-th normalized attention layer 530 may include a cascaded third m-th instance normalization sub-layer 531 and a second m-th window-based multi-head self-attention sub-layer 532. The (m-1)th attention fusion feature map can be obtained by processing the second m-th normalized perceptron layer 540 included in the non-m-th channel attention unit. The second m-th normalized perceptron layer 540 may include a cascaded fourth m-th instance normalization sub-layer 541 and a second m-th multilayer perceptron sub-layer 542.

[0054] According to an embodiment of this application, the image generation model is trained as follows: a first simulated remote sensing image is generated using a generator included in a generative adversarial network model, wherein the generative adversarial network model further includes a discriminator; the generator and discriminator are trained alternately using the first simulated remote sensing image and a first real remote sensing image to obtain a trained generator as the image generation model; wherein the generator and discriminator include a multi-axis self-attention module that shares model parameters, the multi-axis self-attention module being used to acquire local fine-grained features and global dependencies of the first simulated remote sensing image or to acquire local fine-grained features and global dependencies of the first real remote sensing image.

[0055] The acquisition of the first real remote sensing image can be achieved through ground sampling data collection. Taking the acquisition of real remote sensing images of areca palms as an example, the severity can first be preliminarily determined manually on-site, and then located using a sub-meter high-precision GPS receiver; then, a drone equipped with a high-definition digital camera is used to take vertical downward images from a distance of about 10m above the areca palm canopy.

[0056] Acquire and preprocess UAV remote sensing data. Flight missions should be conducted under good lighting conditions and wind speeds less than level 3. The UAV flight path should cover the entire study area. Before and after the aerial photography experiment, a calibration reflective panel should be placed on the ground, ensuring the camera is as perpendicular to the panel as possible to obtain accurate reflectivity data. After acquiring UAV images, the images are stitched together, and then preprocessing such as geometric correction, radiometric calibration, and cropping can be performed.

[0057] A first real remote sensing image can be combined with a random vector and then divided into blocks to obtain multiple patches of the first real remote sensing image. The generator within a generative adversarial network (GAN) model processes these multiple patches to generate a first simulated remote sensing image. The generator and discriminator are trained alternately using the first simulated and first real remote sensing images, and the trained generator is used as the image generation model. The generator and discriminator can include a multi-axis self-attention module sharing model parameters. This module is used to acquire local fine-grained features and global dependencies of the first simulated remote sensing image or to acquire local fine-grained features and global dependencies of the first real remote sensing image.

[0058] The datasets for each category can be split into training and validation sets in an 8:2 ratio for training and testing the image generation model. Since the task of the multi-axis self-attention blocks in both the generator and discriminator is to extract image features, the parameters of the multi-axis self-attention blocks in both the generator and discriminator share weights to reduce the training cost. During training, the Adam optimizer is used for parameter tuning, employing a minimum loss function and least squares loss, and a progressive training strategy is adopted. That is, at each new resolution stage, the model first undergoes a smooth transition period to prevent mode collapse.

[0059] By using an image generation model to generate the first simulated remote sensing image, the diversity and quantity of the dataset are increased, avoiding the situation where the network tends to favor the class with more samples during training while ignoring the class with fewer samples, thus effectively solving the problem of imbalanced data.

[0060] According to an embodiment of this application, a first simulated remote sensing image is generated using a generator included in a generative adversarial network model, comprising: inputting a first real remote sensing image into a linear embedding layer included in the generator, and outputting an embedding feature map; inputting the embedding feature map into N cascaded multi-axis attention layers included in the generator, and outputting an embedding attention feature map, wherein the nth multi-axis attention layer includes a cascaded nth attention sub-layer and an nth grid attention sub-layer, N is an integer greater than or equal to 1, 1≤n≤N; and inputting the embedding attention feature map into an upsampling layer included in the generator, and outputting the first simulated remote sensing image.

[0061] Figure 6 A network structure diagram of an image generation model according to an embodiment of this application is shown.

[0062] like Figure 6 As shown, the first real remote sensing image 610 is input into the linear embedding layer 621 of the generator 620, and an embedding feature map is output. The embedding feature map is input into the N cascaded multi-axis attention layers of the generator 620, and an embedding attention feature map is output. The embedding attention feature map is input into the upsampling layer 624 of the generator 620, and a first simulated remote sensing image is output.

[0063] The first real remote sensing image 610 is processed by generator 620 to obtain the first simulated remote sensing image 630. The first simulated remote sensing image 630 is processed by discriminator 640, which can output a discrimination probability 650 to evaluate the authenticity of the first simulated remote sensing image 630.

[0064] The nth multi-axis attention layer consists of the cascaded nth attention sublayer 622 and nth grid attention sublayer 623, where N is an integer greater than or equal to 1, and 1≤n≤N.

[0065] The first simulated remote sensing image 630 is input into the linear embedding layer 641 of the discriminator 640, which outputs a linear embedding feature map. The linear embedding feature map is then input into the N cascaded multi-axis attention layers of the discriminator 640, which outputs a multi-axis embedding attention feature map. The multi-axis embedding attention feature map is then input into the multilayer perceptron 644 of the discriminator 640, which outputs a discrimination probability 650. The generator 620 and the discriminator 640 may include a multi-axis self-attention module that shares model parameters.

[0066] To address the class imbalance problem in datasets, an image generation model is used to generate simulated images, increasing the diversity and quantity of the dataset. This image generation model is a generative adversarial network based on multi-axis self-attention vision, capable of generating high-quality, high-resolution, and diverse images. The use of this model effectively avoids the tendency of deep learning classification networks to favor classes with larger sample sizes during training, while neglecting classes with smaller sample sizes, thus resolving the class imbalance problem in luteinization disease sample data.

[0067] According to an embodiment of this application, the image classification model is trained as follows: multiple sample remote sensing images are divided into blocks to obtain multiple sample remote sensing image blocks for each sample remote sensing image. The multiple sample remote sensing images are obtained by processing multiple original sample remote sensing images using an image generation model. The balance of the sample remote sensing images indicating different chlorosis categories in the multiple sample remote sensing images is better than the balance of the original sample remote sensing images indicating different chlorosis categories in the multiple original sample remote sensing images. The multiple sample remote sensing image blocks for each sample remote sensing image are processed using a deep learning model to obtain a first simulated sample feature vector for each sample remote sensing image. The deep learning model is trained using the first simulated sample feature vectors for each sample remote sensing image and the sample remote sensing image classification labels to obtain the image classification model.

[0068] By processing multiple original sample remote sensing images using an image generation model, multiple sample remote sensing images can be obtained. The balance of the sample remote sensing images indicating different chlorosis categories among the multiple sample remote sensing images is better than the balance of the original sample remote sensing images indicating different chlorosis categories. Dividing the multiple sample remote sensing images into blocks yields multiple sample remote sensing image blocks for each of the multiple sample remote sensing images. Processing these multiple sample remote sensing image blocks with a deep learning model yields the first simulated sample feature vector for each of the multiple sample remote sensing images. Training the deep learning model using the first simulated sample feature vectors of each of the multiple sample remote sensing images and the classification labels of the sample remote sensing images yields an image classification model.

[0069] In addition, the entity that performs the training of the image classification model can be the same as or different from the entity that performs the application of the image classification model.

[0070] According to embodiments of this application, feature extraction is performed on multiple sample remote sensing images and multiple original sample remote sensing images to obtain second simulated sample feature vectors for each of the multiple sample remote sensing images and real sample feature vectors for each of the multiple original sample remote sensing images; based on the multiple second simulated sample feature vectors, simulated feature mean and simulated covariance matrix are obtained; based on the multiple real sample feature vectors, real feature mean and real covariance matrix are obtained; based on the simulated feature mean, real feature mean, simulated covariance matrix, and real covariance matrix, the similarity between the multiple sample remote sensing images and the multiple original sample remote sensing images is determined; if the similarity is greater than or equal to a preset threshold, the multiple sample remote sensing images are used as training samples for training a deep learning model.

[0071] To evaluate the quality of multiple sample remote sensing images generated by the image generation model, the Fréchet Inception Distance (FID score) is used as a metric. This metric directly considers the distance between the sample remote sensing image (simulated generated image) and the original sample remote sensing image (real image) at the feature level to measure the similarity between the sample remote sensing image and the original sample remote sensing image.

[0072] By extracting features from multiple sample remote sensing images and multiple original sample remote sensing images, we can obtain the second simulated sample feature vectors for each of the multiple sample remote sensing images and the true sample feature vectors for each of the multiple original sample remote sensing images.

[0073] By fitting multiple second-simulated sample feature vectors, a multidimensional Gaussian distribution of these feature vectors can be obtained. Based on this Gaussian distribution, the simulated feature mean and simulated covariance matrix can then be derived. Similarly, by fitting multiple real sample feature vectors, a multidimensional Gaussian distribution of these real sample feature vectors can be obtained. Based on this Gaussian distribution, the real feature mean and real covariance matrix can then be derived.

[0074] Based on the simulated feature mean, the true feature mean, the simulated covariance matrix, and the true covariance matrix, the similarity between multiple sample remote sensing images and multiple original sample remote sensing images can be determined, as shown in formula (2).

[0075] (2);

[0076] in, The true mean of the features. To simulate the mean of the features, The true covariance matrix, To simulate the covariance matrix, This represents the sum of the elements on the diagonal of the matrix.

[0077] Based on formula (2), the similarity between multiple sample remote sensing images and multiple original sample remote sensing images can be determined. If the similarity is greater than or equal to a preset threshold, the multiple sample remote sensing images can be used as training samples for training a deep learning model.

[0078] According to an embodiment of this application, a deep learning model is trained using the first simulated sample feature vectors of each of the multiple sample remote sensing images and the classification labels of the sample remote sensing images to obtain an image classification model. This includes: obtaining a loss function value based on the first simulated sample feature vectors of each of the multiple sample remote sensing images, the classification labels of each of the multiple sample remote sensing images, and the true class weight vectors corresponding to the classification labels of each of the multiple sample remote sensing images; and adjusting the model parameters of the deep learning model based on the loss function value to obtain the image classification model.

[0079] Based on the loss function, the loss function value can be obtained from the first simulated sample feature vector of each of the multiple sample remote sensing images, the sample remote sensing image classification label of each of the multiple sample remote sensing images, and the real class weight vector corresponding to the classification label of each of the multiple sample remote sensing images, as shown in formula (3).

[0080] (3);

[0081] in, The value of the loss function; The number of samples; For the first Classification labels for remote sensing images of each sample; Let be the magnitude of the feature vector of the first simulated sample; For the first The first simulated sample feature vector of the nth sample and the nth sample The angle between the true class weight vectors of each class Indicates the first The angle between the first simulated sample feature vector and the true class weight vector of each sample; is a hyperparameter representing the angular boundary.

[0082] By adjusting the model parameters of a deep learning model based on the loss function value, an image classification model can be obtained, which improves the tightness of similar samples in the feature space and the difference between samples of different categories.

[0083] Multiple evaluation methods were used to assess the results of image classification models in classifying the severity of yellowing in remote sensing images to be classified. These methods may include confusion matrix, accuracy, recall, precision, and F1 score.

[0084] The confusion matrix evaluates the performance of an image classification model across different categories. Information such as TruePositives (TP), True Negatives (TN), False Positives (FP), and False Negatives (FN) indicates the model's classification results. The parameters of the confusion matrix are shown in Table 1.

[0085] Table 1 Confusion Matrix Parameters

[0086]

[0087] Accuracy represents the proportion of correctly predicted samples out of the total number of samples; recall represents the proportion of correctly predicted positive samples out of the actual number of positive samples; precision represents the proportion of correctly predicted positive samples out of the total number of predicted positive samples; F1-Score is the harmonic mean of accuracy and recall, used to comprehensively evaluate the model's performance. The formulas for calculating these four metrics are as follows:

[0088] (4)

[0089] (5)

[0090] (6)

[0091] (7)

[0092] in, is the weight of the number of categories for the i-th sample; L is the number of categories.

[0093] Multiple evaluation metrics, including confusion matrix, accuracy, precision, recall, and F1 score, are used to comprehensively evaluate the model's performance. This provides a more comprehensive understanding of the image classification results for the remote sensing images to be classified, offering a basis for further optimization and improvement.

[0094] Based on the above image classification method, this application also provides an image classification apparatus. The following will combine... Figure 7 The device is described in detail.

[0095] Figure 7 A structural block diagram of an image classification apparatus according to an embodiment of this application is shown.

[0096] like Figure 7 As shown, the image classification device 700 of this embodiment includes a block segmentation module 710, an Mth feature map output module 720, and an image classification result acquisition module 730.

[0097] The segmentation module 710 is used to segment the remote sensing image to be classified into multiple remote sensing image blocks.

[0098] The Mth feature map output module 720 is used to input multiple remote sensing image patches to be classified into M cascaded feature extraction modules in the image classification model and output the Mth feature map. The mth feature extraction module includes a non-mth channel attention layer and an mth channel attention layer. When 1 < m ≤ M, the mth feature map is obtained by processing the mth intermediate feature map using the mth channel attention layer. The mth intermediate feature map is obtained by processing the (m-1)th feature map using the non-mth channel attention layer. The mth channel attention layer is used to evaluate the importance of each feature channel in the mth intermediate feature map. When m = 1, the first feature map is obtained by processing multiple remote sensing image patches to be classified using the first feature extraction module. M is an integer greater than 1.

[0099] The image classification result acquisition module 730 is used to input the Mth feature map into the classification module of the image classification model and output the image classification result of the remote sensing image to be classified, wherein the image classification result indicates the chlorosis category of the remote sensing image to be classified.

[0100] According to embodiments of this application, by dividing the remote sensing image to be classified into blocks, multiple remote sensing image blocks to be classified can be obtained. These multiple blocks are then input into M cascaded feature extraction modules in an image classification model, outputting the Mth feature map. The mth feature extraction module can include non-mth channel attention units and the mth channel attention unit. When 1 < m ≤ M, processing the (m-1)th feature map using the non-mth channel attention unit yields the mth intermediate feature map, enhancing the extraction capability of subtle features in the remote sensing image to be classified. Processing the mth intermediate feature map using the mth channel attention unit yields the mth feature map, reducing inter-channel differences and improving the contrast of the feature map, thereby strengthening the ability to capture image features. Inputting the Mth feature map into the classification module of the image classification model outputs the image classification result of the remote sensing image to be classified, thus improving the accuracy of chlorosis identification.

[0101] The Mth feature map output module 720 includes: an average pooling submodule, a first processing submodule, a max pooling submodule, a second processing submodule, a third processing submodule, and an mth feature map acquisition submodule.

[0102] The average pooling submodule is used to perform average pooling on the m-th intermediate feature map to obtain the m-th average pooled feature vector.

[0103] The first processing submodule is used to process the m-th average pooling feature vector using the m-th first multilayer perceptron layer included in the m-th channel attention unit to obtain the m-th first intermediate feature vector.

[0104] The max pooling submodule is used to perform max pooling on the m-th intermediate feature map to obtain the m-th max pooled feature vector.

[0105] The second processing submodule is used to process the m-th max pooling feature vector using the m-th second multilayer perceptron layer included in the m-th channel attention unit to obtain the m-th second intermediate feature vector.

[0106] The third processing submodule is used to process the m-th average pooling feature vector and the m-th max pooling feature vector using activation functions to obtain the m-th fused feature vector.

[0107] The m-th feature map is used to obtain the m-th feature map based on the m-th fused feature vector and the m-th intermediate feature map.

[0108] The Mth feature map output module 720 includes: a fourth processing submodule, a fifth processing submodule, a sixth processing submodule, and a seventh processing submodule.

[0109] The fourth processing submodule is used to process the (m-1)th feature map using the first m-th normalized attention layer included in the non-m-th channel attention unit to obtain the (m-1)th normalized feature map. The first m-th normalized attention layer includes a cascaded first m-th instance normalization sublayer and a first m-th window-based multi-head self-attention sublayer.

[0110] The fifth processing submodule is used to process the (m-1)th normalized feature map using the first m-th normalized perceptron layer included in the non-m-th channel attention unit to obtain the (m-1)th normalized fusion feature map. The first m-th normalized perceptron layer includes a cascaded second m-th instance normalization sublayer and a first m-th multilayer perceptron sublayer.

[0111] The sixth processing submodule is used to process the (m-1)th normalized fusion feature map using the second m-th normalized attention layer included in the non-m-th channel attention unit to obtain the (m-1)th attention fusion feature map. The second m-th normalized attention layer includes a cascaded third m-th instance normalization sublayer and a second m-th window-based multi-head self-attention sublayer.

[0112] The seventh processing submodule is used to process the (m-1)th attention fusion feature map using the second m-th normalized perceptron layer included in the non-m-th channel attention unit to obtain the m-th intermediate feature map. The second m-th normalized perceptron layer includes a cascaded fourth m-th instance normalized sublayer and a second m-th multilayer perceptron sublayer.

[0113] The Mth feature map output module 720 includes a block segmentation submodule, a remote sensing image block processing submodule, and a training submodule.

[0114] The segmentation submodule is used to segment multiple sample remote sensing images into blocks, resulting in multiple sample remote sensing image blocks for each sample remote sensing image. The multiple sample remote sensing images are obtained by processing multiple original sample remote sensing images using an image generation model. The balance of the sample remote sensing images indicating different chlorosis categories in the multiple sample remote sensing images is better than the balance of the original sample remote sensing images indicating different chlorosis categories in the multiple original sample remote sensing images.

[0115] The remote sensing image patch processing submodule is used to process multiple sample remote sensing image patches of multiple sample remote sensing images using a deep learning model, and obtain the first simulated sample feature vector of each of the multiple sample remote sensing images.

[0116] The training submodule is used to train a deep learning model using the first simulated sample feature vectors of multiple sample remote sensing images and the classification labels of the sample remote sensing images, so as to obtain an image classification model.

[0117] The segmented sub-modules include: generation units and alternating training units.

[0118] A generation unit is used to generate a first simulated remote sensing image using a generator included in a generative adversarial network model, wherein the generative adversarial network model also includes a discriminator.

[0119] An alternating training unit is used to alternately train the generator and discriminator using a first simulated remote sensing image and a first real remote sensing image to obtain a trained generator as an image generation model.

[0120] The generator and discriminator include a multi-axis self-attention module that shares model parameters. The multi-axis self-attention module is used to obtain local fine-grained features and global dependencies of a first simulated remote sensing image or to obtain local fine-grained features and global dependencies of a first real remote sensing image.

[0121] The generation unit includes: a first input subunit, a second input subunit, and a third input subunit.

[0122] The first input subunit is used to input a first real remote sensing image into the linear embedding layer of the generator and output an embedded feature map.

[0123] The second input subunit is used to input the embedded feature map into the cascaded N multi-axis attention layers of the generator and output the embedded attention feature map. The nth multi-axis attention layer includes the cascaded nth block attention sublayer and the nth grid attention sublayer. N is an integer greater than or equal to 1, and 1≤n≤N.

[0124] The third input subunit is used to input the embedded attention feature map into the upsampling layer included in the generator and output the first simulated remote sensing image.

[0125] The Mth feature map output module 720 also includes: a feature extraction submodule, a simulated mean matrix submodule, a true mean matrix submodule, a similarity submodule, and a training sample submodule.

[0126] The feature extraction submodule is used to extract features from multiple sample remote sensing images and multiple original sample remote sensing images to obtain the second simulated sample feature vector of each of the multiple sample remote sensing images and the real sample feature vector of each of the multiple original sample remote sensing images.

[0127] The simulated mean matrix submodule is used to obtain the simulated feature mean and simulated covariance matrix based on the feature vectors of multiple second simulated samples.

[0128] The True Mean Matrix submodule is used to obtain the true feature mean and the true covariance matrix based on the feature vectors of multiple true samples.

[0129] The similarity submodule is used to determine the similarity between multiple sample remote sensing images and multiple original sample remote sensing images based on the simulated feature mean, the true feature mean, the simulated covariance matrix, and the true covariance matrix.

[0130] The training sample submodule is used to use multiple remote sensing images as training samples for training a deep learning model when the similarity is greater than or equal to a preset threshold.

[0131] The training submodule includes a loss function value unit and an adjustment unit.

[0132] The loss function value unit is used to obtain the loss function value based on the loss function, according to the first simulated sample feature vector of each of the multiple sample remote sensing images, the sample remote sensing image classification label of each of the multiple sample remote sensing images, and the true class weight vector corresponding to the classification label of each of the multiple sample remote sensing images.

[0133] The adjustment unit is used to adjust the model parameters of the deep learning model based on the loss function value to obtain the image classification model.

[0134] According to embodiments of this application, any plurality of modules in the block segmentation module 710, the Mth feature map output module 720, and the image classification result acquisition module 730 can be merged into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the block segmentation module 710, the Mth feature map output module 720, and the image classification result acquisition module 730 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any one of the three implementation methods or a suitable combination of any of them. Alternatively, at least one of the block segmentation module 710, the Mth feature map output module 720, and the image classification result acquisition module 730 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0135] Figure 8 A block diagram of an electronic device suitable for implementing an image classification method according to an embodiment of this application is shown.

[0136] like Figure 8 As shown, an electronic device 800 according to an embodiment of this application includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage portion 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0137] RAM 803 stores various programs and data required for the operation of electronic device 800. Processor 801, ROM 802, and RAM 803 are interconnected via bus 804. Processor 801 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 802 and / or RAM 803. It should be noted that programs may also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in one or more memories.

[0138] According to embodiments of this application, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to a bus 804. The electronic device 800 may also include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 810 as needed so that computer programs read from it can be installed into the storage section 808 as needed.

[0139] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the image classification method according to the embodiments of this application.

[0140] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803 described above.

[0141] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the image classification method provided in the embodiments of this application.

[0142] When the computer program is executed by the processor 801, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0143] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 809, and / or installed from a removable medium 811. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0144] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0145] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0146] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0147] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

Claims

1. An image classification method, characterized by, The method comprises: blocking the remote sensing image to be classified to obtain a plurality of remote sensing image blocks to be classified; inputting the plurality of remote sensing image blocks to be classified into M feature extraction modules cascaded in an image classification model, and outputting an Mth feature map, wherein the mth feature extraction module comprises a non-mth channel attention unit and an mth channel attention unit, in the case of 1 inputting the Mth feature map into a classification module in the image classification model, and outputting an image classification result of the remote sensing image to be classified, wherein the image classification result indicates the yellowing disease category of the remote sensing image to be classified.

2. The method of claim 1, wherein, The Mth feature map is obtained by the following method: averaging pooling the Mth intermediate feature map to obtain an Mth average pooling feature vector; processing the Mth average pooling feature vector by an Mth first multi-layer perceptron layer included in the Mth channel attention unit to obtain an Mth first intermediate feature vector; max-pooling the Mth intermediate feature map to obtain an Mth max-pooling feature vector; processing the Mth max-pooling feature vector by an Mth second multi-layer perceptron layer included in the Mth channel attention unit to obtain an Mth second intermediate feature vector; processing the Mth average pooling feature vector and the Mth max-pooling feature vector by an activation function to obtain an Mth fusion feature vector; obtaining the Mth feature map according to the Mth fusion feature vector and the Mth intermediate feature map.

3. The method according to claim 1 or 2, characterized in that, The Mth intermediate feature map is obtained by the following method: processing the (m-1)th feature map by a first Mth normalization attention layer included in the non-mth channel attention unit to obtain an (m-1)th normalized feature map, wherein the first Mth normalization attention layer comprises a first Mth instance normalization sub-layer and a first Mth window-based multi-head self-attention sub-layer cascaded; processing the (m-1)th normalized feature map by a first Mth normalized perceptron layer included in the non-mth channel attention unit to obtain an (m-1)th normalized fusion feature map, wherein the first Mth normalized perceptron layer comprises a second Mth instance normalization sub-layer and a first Mth multi-layer perceptron sub-layer cascaded; processing the (m-1)th normalized fusion feature map by a second Mth normalization attention layer included in the non-mth channel attention unit to obtain an (m-1)th attention fusion feature map, wherein the second Mth normalization attention layer comprises a third Mth instance normalization sub-layer and a second Mth window-based multi-head self-attention sub-layer cascaded; The mth intermediate feature map is obtained by processing the (m-1)th attention fusion feature map by using a second mth normalization perception unit included in the non-mth channel attention unit, wherein the second mth normalization perception unit includes a cascaded fourth mth instance normalization sub-layer and a second mth multi-layer perception sub-layer.

4. The method according to claim 1 or 2, characterized in that, The image classification model is trained in the following manner: The multiple sample remote sensing images are divided into blocks to obtain multiple sample remote sensing image blocks of the multiple sample remote sensing images, wherein the multiple sample remote sensing images are obtained by processing multiple original sample remote sensing images by using an image generation model, and the balance of sample remote sensing images indicating different yellowing disease categories in the multiple sample remote sensing images is better than the balance of original sample remote sensing images indicating different yellowing disease categories in the multiple original sample remote sensing images; The first simulation sample feature vector of each of the multiple sample remote sensing images is obtained by processing the multiple sample remote sensing image blocks of each of the multiple sample remote sensing images by using a deep learning model; The image classification model is obtained by training the deep learning model by using the first simulation sample feature vector of each of the multiple sample remote sensing images and a sample remote sensing image classification label.

5. The method of claim 4, wherein, The image generation model is trained in the following manner: A first simulation remote sensing image is generated by using a generator included in a generative adversarial network model, wherein the generative adversarial network model further includes a discriminator; The trained generator is used as the image generation model by alternately training the generator and the discriminator by using the first simulation remote sensing image and a first real remote sensing image; The generator and the discriminator include a multi-axis self-attention module sharing model parameters, and the multi-axis self-attention module is used to obtain local fine-grained features and global dependencies of the first simulation remote sensing image or used to obtain local fine-grained features and global dependencies of the first real remote sensing image.

6. The method of claim 5, wherein, The first simulation remote sensing image is generated by using a generator included in a generative adversarial network model, including: The first real remote sensing image is input into a linear embedding layer included in the generator to output an embedding feature map; The embedding feature map is input into a cascaded N multi-axis attention layer included in the generator to output an embedding attention feature map, wherein an nth multi-axis attention layer includes a cascaded nth block attention sub-layer and an nth grid attention sub-layer, N is an integer greater than or equal to 1, and 1≤n≤N; The embedding attention feature map is input into an upsampling layer included in the generator to output the first simulation remote sensing image.

7. The method of claim 4, wherein, Further comprising: Feature extraction is performed on the multiple sample remote sensing images and the multiple original sample remote sensing images to obtain a second simulation sample feature vector of each of the multiple sample remote sensing images and a real sample feature vector of each of the multiple original sample remote sensing images; A simulation feature mean and a simulation covariance matrix are obtained according to multiple second simulation sample feature vectors; A real feature mean and a real covariance matrix are obtained according to multiple real sample feature vectors; determine a similarity between the plurality of sample remote sensing images and the plurality of original sample remote sensing images according to the simulated feature mean value, the real feature mean value, the simulated covariance matrix and the real covariance matrix; in a case where the similarity is greater than or equal to a preset threshold, use the plurality of sample remote sensing images as training samples for training the deep learning model.

8. The method of claim 4, wherein, training the deep learning model using the first simulated sample feature vector of each of the plurality of sample remote sensing images and the sample remote sensing image classification label to obtain the image classification model, including: obtaining a loss function value based on a loss function according to the first simulated sample feature vector of each of the plurality of sample remote sensing images, the sample remote sensing image classification label of each of the plurality of sample remote sensing images and a real class weight vector corresponding to each of the sample remote sensing image classification labels; adjusting a model parameter of the deep learning model according to the loss function value to obtain the image classification model.

9. An image classification apparatus characterized by comprising: The device comprises: a blocking module configured to block the remote sensing image to be classified to obtain a plurality of remote sensing image blocks to be classified; an Mth feature map output module configured to input the plurality of remote sensing image blocks to be classified into M feature extraction modules cascaded in an image classification model to output an Mth feature map, wherein the Mth feature extraction module comprises a non-Mth channel attention layer and an Mth channel attention layer, in a case where 1 an image classification result obtaining module configured to input the Mth feature map into a classification module in the image classification model to output an image classification result of the remote sensing image to be classified, wherein the image classification result indicates a yellow disease class of the remote sensing image to be classified.

10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1-8.