An interpretable classification method for pest images based on image text generation technology
Through the method based on image text generation technology, Faster-RCNN and Transformer models are used to identify body parts of pest images and generate text descriptions, which solves the problem of low classification accuracy of CNN models when the number of samples is small or the appearance of pests is similar, and high-accurate pest image classification is achieved and interpretable text description is provided.
Patent Information
- Application Number
- CN202310287333.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-23
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-03-23
AI Technical Summary
Existing CNN models have low classification accuracy and are difficult to simulate the diagnostic behavior of agricultural experts when the number of samples per class is small and there is a large total number of categories, or when the appearance of the two pests is very similar.
The pest image interpretability classification method based on image text generation technology is adopted to interpret the pest image classification results by generating text descriptions. The multimodal data set is collected using network crawler technology, combined with Faster-RCNN and Transformer models, the body parts of the pest image are identified and text descriptions are generated, and visual features and text features are fused for classification.
It effectively improves the accuracy of pest image classification, especially when the number of samples is small or the appearance of the pest is similar, it can provide text-level explanations and mimic the diagnostic process of agricultural experts.
Smart Images

Figure CN116310564B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multimodal data processing, and in particular relates to an interpretable classification method for pest images based on image text generation technology. Background Art
[0002] As the scale of agricultural production expands, the probability of crops being attacked by pests is also increasing. Pest identification and control are of great significance to the agricultural economy. Early identification of pests relies on agricultural experts, which is a high-labor-intensity and low-real-time job. At the same time, the number of agricultural experts is not enough to cope with the various types of pests. In recent years, with the development of deep learning, the use of computer vision technology to develop pest image classification models can complete the automatic diagnosis task of agricultural pest images.
[0003] Pest image classification models based on convolutional neural networks (CNNs) perform well. However, good CNN models require a large amount of training data. When the number of samples in each class is small and the total number of classes is large, the classification accuracy of the CNN model will drop rapidly. On the other hand, when the appearance of two pests is very similar, the classification accuracy of the CNN model will also drop. However, agricultural experts can accurately determine the type of pests by carefully observing the pests and describing their physical characteristics. Therefore, designing a model that can simulate the diagnostic behavior of agricultural experts is expected to solve the above problems. Summary of the invention
[0004] In order to solve the above problems, the present invention proposes an interpretable classification method for pest images based on image-text generation technology. By generating text descriptions to explain the pest image classification results to imitate the diagnosis process of agricultural experts, the problem of low classification accuracy of the CNN model can be solved when the number of samples in each class is small and the total number of categories is large or when the appearance of two pests is very similar.
[0005] The technical solution of the present invention is as follows:
[0006] A pest image interpretable classification method based on image-text generation technology explains the pest image classification results by generating text descriptions to imitate the diagnosis process of agricultural experts. The specific steps include:
[0007] Step 1: Use web crawler technology to collect and construct a multimodal dataset with pest images and corresponding text descriptions;
[0008] Step 2: Use the Faster-RCNN model to identify the body parts of the pest image and extract the features of each body part;
[0009] Step 3: Use the Transformer model to generate a text description of the pest image, and fuse the visual features and text features of the pest image to form joint features;
[0010] Step 4: Use the joint features to train a classifier to predict the category label of the pest;
[0011] Step 5: Combine the generated text description and the predicted pest category label to interpret the classification result of the pest at the text level.
[0012] Further, the specific process of Step 1 is as follows:
[0013] Step 1.1: Use web crawler technology to collect pest image data in the Baidu Image Search Engine and Google Image Search Engine, and establish a pest candidate image dataset;
[0014] Step 1.2: Perform image and text preprocessing to obtain the final required agricultural pest multi-modal dataset.
[0015] Further, the specific process of Step 1.2 is as follows:
[0016] Step 1.2.1: Convert all images to the JPEG format, and delete the images that cannot be displayed normally from the candidate image dataset;
[0017] Step 1.2.2: Filter out the images with pixel sizes less than 448*448, and adjust the pixel sizes of all qualified images to 448*448;
[0018] Step 1.2.3: Use the image annotation tool Labelme to annotate the body parts of the pest image to obtain the semantic labels of the parts, and save the annotation file in the Json format;
[0019] Step 1.2.4: Invite 5 experts in the agricultural field to describe the pest based on the color, shape, material, and size of the pest body parts in each image, forming 5 text descriptions;
[0020] Step 1.2.5: Combine each image with its corresponding 5 text descriptions to form an image-text pair, and finally constitute an agricultural pest multi-modal dataset.
[0021] Further, the specific process of Step 2 is as follows:
[0022] Step 2.1: Represent the dataset containing n pest images as V = {V1,..., V i ,..., V n}, and another represents the i-th pest image V im body parts in it, where the j-th body part of the i-th pest image contains two parts of information: (1) represents the label of the body part, and M represents the total number of body part categories; (2) represents the coordinates of the bounding box of the body part;
[0023] Step 2.2. Input V i into the Faster-RCNN model and obtain through supervised training and two mapping functions; used to distinguish the labels of each body part in the pest image; used to identify the bounding box of each body part; this process is expressed by the formula:
[0024]
[0025] where, θ d represents 's parameters, and θ r represents 's parameters;
[0026] Step 2.3. Use and two mapping functions to generate the features of m body parts: represents the feature of each body part.
[0027] Furthermore, the specific process of Step 3 is as follows:
[0028] Step 3.1. Input the features representing the pest body parts into the Encoder module of the Transformer model, and weight each feature through the multi-head self-attention mechanism to obtain the hidden vector representation of the feature: Fh i ∈R m×2048 , this process is expressed by the formula:
[0029] Fh i = Encoder(F i ; θ enc ), Fh i ∈R m×2048 (2);
[0030] where, θ enc represents the parameters of Encoder(·);
[0031] Step 3.2. Design a specific text input for the Decoder module of the Transformer;
[0032] Let represent the text description corresponding to the pest body part P i which is the input of the Decoder module, and fill the [Start] identifier at the start position of the text description; T represents the length of the text description, and L represents the length of the vocabulary;
[0033] Let represent the output of the Decoder module, which is the text description generated by the Transformer model, and fill the [End] identifier at the end position of the text description;
[0034] Step 3.3, input the text description and the hidden vector representation Fh of the pest body part features i into the first layer of the Decoder module, and learn the joint representation of visual features and text features through the multi-head self-attention mechanism This process is expressed by the formula:
[0035]
[0036] where θ dec represents the parameters of Decoder(·);
[0037] Step 3.4, after stacking the Decoder module N times, the output of the previous Decoder module is the input of the next Decoder module, and finally obtain the joint feature Ft after fusing visual features and text features i ;
[0038] Step 3.5, the joint feature Ft i passes through a two-layer fully connected layer module and the Softmax(·) function to obtain the probability distribution of each vocabulary. This process is expressed by the formula:
[0039]
[0040] where θ gen represents the parameters of the fully connected layer module MLP gen (·);
[0041] Step 3.6, query the vocabulary according to the probability distribution of each vocabulary to obtain the final output
[0042] Furthermore, the specific process of Step 4 is as follows:
[0043] Step 4.1, the classifier consists of N stacked ClsDecoder modules; the input of the first layer of the ClsDecoder module consists of two parts. The first part is It is the mean value of the characteristics of each body part of the pest image, used to represent the global characteristics of the entire pest image; the second part is the output of the first-layer Decoder module in step 3.3 This process is expressed by the formula:
[0044]
[0045] Among them, θ clsdec represents the parameters of the fully connected layer module ClsDecoder(·), is the hidden vector representation of the pest image;
[0046] Step 4.2, after stacking the ClsDecoder module N times, the output of the upper-layer ClsDecoder module is the input of the lower-layer ClsDecoder module, and finally the hidden vector representation Hc of the pest image i passes through a two-layer fully connected layer module and the Softmax(·) function to obtain the probability distribution of each category. This process is expressed by the formula:
[0047]
[0048] Among them, θ cls represents the parameters of the fully connected layer module MLP cls (·), and Z represents the total number of categories;
[0049] Step 4.3, obtain the corresponding category label according to the maximum value of the probability distribution of each category.
[0050] Furthermore, in step 5, combine the text description generated in step 3.6 and the pest category label predicted in step 4.3 to construct a sentence pattern to achieve the interpretable classification of the pest image; the specific format of the sentence pattern is: because it is observed that so-and-so, so this pest is predicted to be of so-and-so type.
[0051] The beneficial technical effects brought by the present invention:
[0052] The present invention proposes an interpretable classification method for pest images based on image text generation technology, which can not only identify the body parts of pests, but also generate text descriptions according to the identified body parts to explain the pest classification results, and can effectively solve the problem of low classification accuracy of the CNN model when the number of samples in each category is small and the total number of categories is large or when the appearances of two pests are very similar, and improve the classification accuracy. Brief Description of the Drawings
[0053] Figure 1 is a flowchart of the interpretable classification method for pest images based on image text generation technology of the present invention;
[0054] Figure 2 This is the flowchart of image and text preprocessing for the present invention;
[0055] Figure 3 This is the flowchart of using the Faster-RCNN model to identify each body part of the pest image in the present invention;
[0056] Figure 4 This is the flowchart of using the Transformer model to generate text descriptions for pest images in the present invention;
[0057] Figure 5 This is the flowchart of using combined features to train a classifier to predict the category label of pests in the present invention. Specific embodiments
[0058] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments:
[0059] A method for interpretable classification of pest images based on image-text generation technology can give an explanatory text description while predicting the pest label, and can effectively solve the problem of low classification accuracy of the CNN model when the number of samples in each class is small and the total number of classes is large, or when the appearances of two pests are very similar. As Figure 1 shown, this method explains the classification result of the pest image by generating a text description to imitate the diagnosis process of agricultural experts, and specifically includes the following steps:
[0060] Step 1: Use web crawler technology to collect and construct a multi-modal dataset with pest images and corresponding text descriptions.
[0061] The specific process is as follows:
[0062] Step 1.1: Use web crawler technology to collect pest image data in the Baidu Image Search Engine and the Google Image Search Engine, and establish a pest candidate image dataset.
[0063] There are 28 kinds of pests harmful to cash crops, 19 kinds of pests harmful to food crops, 30 kinds of pests harmful to fruit trees, and 17 kinds of pests harmful to vegetables. Since the characteristics of pests vary greatly in different development stages, the states of pests are divided into two categories: adults and larvae. For example, if a certain pest has 90 adults and 9 larvae, then the image dataset of this pest has 99 categories at this time. When collecting images in the Baidu Image Search Engine and the Google Image Search Engine, the number of images should be no less than 200,000.
[0064] Step 1.2: Perform image and text preprocessing to obtain the final required multi-modal dataset of agricultural pests. As Figure 2 shown, the specific process is as follows:
[0065] Step 1.2.1: Convert all images to the JPEG format and delete the images that cannot be displayed properly from the candidate image dataset;
[0066] Step 1.2.2: Filter out the images with pixel dimensions less than 448*448 and resize all eligible images to a pixel size of 448*448;
[0067] Step 1.2.3: Use the image annotation tool Labelme to annotate the body parts of the pest images to obtain the semantic labels of the parts and save the annotation files in the Json format.
[0068] Step 1.2.4: Invite 5 experts in the agricultural field to describe the pests based on the color, shape, material, and size of the pest body parts in each image, forming 5 text descriptions.
[0069] Step 1.2.5: Combine each image with its corresponding 5 text descriptions to form image-text pairs, ultimately constituting a multi-modal dataset of agricultural pests.
[0070] Step 2: Imitate the "observation" actions of experts, that is, use the Faster-RCNN model to identify the body parts of the pest images and extract the features of each body part. As Figure 3 shown, the specific process is as follows:
[0071] Step 2.1: Represent the dataset containing n pest images as V = {V1,..., V i ,..., V n}. Another use to represent the m body parts in the i-th pest image V i , where the j-th body part in the i-th pest image contains two parts of information: (1) represents the label of the body part, and M represents the total number of body part categories; (2) represents the coordinates of the bounding box of the body part.
[0072] Step 2.2: Input V i into the Faster-RCNN model and use supervised training to obtain and two mapping functions. is used to distinguish the labels of each body part in the pest image. is used to identify the bounding box of each body part. This process can be expressed by the formula:
[0073]
[0074] where θd The parameter represented by , θ r represents the parameter of.
[0075] Step 2.3: Use and two mapping functions to generate the features of m body parts: represents the feature of each body part.
[0076] Step 3: Imitate the "description" action of the expert, that is, use the Transformer model to generate a text description of the pest image, and fuse the visual features and text features of the pest image to form a joint feature. As Figure 4 shown, the specific process is as follows:
[0077] Step 3.1: First, input the feature representing the pest body part into the Encoder module of the Transformer model, and weight each feature through the multi-head self-attention mechanism to obtain the hidden vector representation of the feature: Fh i ∈R m×2048 , and this process can be expressed by the formula:
[0078] Fh i = Encoder(F i ; θ enc ), Fh i ∈R m×2048 (2);
[0079] Among them, θ enc represents the parameter of Encoder(·).
[0080] Step 3.2: In order to enable the Transformer model to obtain the ability to generate text descriptions, a specific text input needs to be designed for the Decoder module of the Transformer.
[0081] Let represent the text description corresponding to the pest body part P i , which is the input of the Decoder module, and fill the [Start] identifier at the starting position of the text description. T represents the length of the text description, and L represents the length of the vocabulary.
[0082] Let Represents the output of the Decoder module, which is the text description generated by the Transformer model, and fills the [End] identifier at the end position of the text description. This way of constructing the identifier is to enable the Decoder module to learn the ability to predict the next word and generate the next word through recursive operations in a loop, ultimately generating the entire text description.
[0083] Step 3.3: The text description and the hidden vector representation Fh of the pest body part features i are input into the first layer of the Decoder module, and the joint representation of visual features and text features is learned through the multi-head self-attention mechanism This process can be expressed by the formula:
[0084]
[0085] where θ dec represents the parameters of Decoder(·).
[0086] Step 3.4: After N times of stacking of the Decoder module, the output of the previous Decoder module is the input of the next Decoder module, and finally the joint feature Ft after fusing visual features and text features is obtained i .
[0087] Step 3.5: The joint feature Ft i passes through a two-layer fully connected layer module and the Softmax(·) function to obtain the probability distribution of each vocabulary. This process can be expressed by the formula:
[0088]
[0089] where θ gen represents the parameters of the fully connected layer module MLP gen (·).
[0090] Step 3.6: Query the vocabulary table according to the probability distribution of each vocabulary to obtain the final output
[0091] Step 4: Imitate the "judgment" action of the expert, that is, use the joint feature to train a classifier to predict the category label of the pest. As Figure 5 shown, the specific process is as follows:
[0092] Step 4.1: The classifier consists of N stacked ClsDecoder modules. The input of the first layer of the ClsDecoder module consists of two parts. The first part is It is the mean value of the features of each body part of the pest image and is used to represent the global features of the entire pest image. The second part is the output of the first-layer Decoder module in step 3.3 This process can be expressed by the formula:
[0093]
[0094] where θ clsdec represents the parameters of the fully connected layer module ClsDecoder(·), is the hidden vector representation of the pest image.
[0095] Step 4.2: After stacking the ClsDecoder module N times, the output of the upper-layer ClsDecoder module is the input of the lower-layer ClsDecoder module. Finally, the hidden vector representation Hc of the pest image i passes through a two-layer fully connected layer module and the Softmax(·) function to obtain the probability distribution of each category. This process can be expressed by the formula:
[0096]
[0097] where θ cls represents the parameters of the fully connected layer module MLP cls (·), and Z represents the total number of categories.
[0098] Step 4.3: The corresponding category label can be obtained according to the maximum value of the probability distribution of each category.
[0099] Step 5: Combine the generated text description and the predicted pest category label to interpret the classification result of the pest at the text level. The specific process is as follows:
[0100] Combine the text description generated in step 3.6 and the pest category label predicted in step 4.3 to construct a sentence pattern of "because XX is observed, so the pest is predicted to be XX type" to achieve the purpose of interpretable classification of the pest image.
[0101] To prove the feasibility and superiority of the model trained by the method of the present invention, comparative experiments are conducted with the existing technologies FC, Att2in, ShowTell, AdaAtt, UpDown, and AoANet. The experimental results are shown in Table 1:
[0102] Table 1 Performance comparison of each model
[0103]
[0104] The larger the values of Bleu-1, Bleu-2, Bleu-3, Bleu-4, Cider-D, and Rouge, the stronger the model's text generation ability. It can be seen from the experimental results that the method of the present invention achieved the best prediction performance, and Bleu-1, Bleu-2, Bleu-3, Bleu-4, Cider-D, and Rouge are all at their maximum values.
[0105] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions, or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method for interpretable classification of pest images based on image text generation technology, characterized in that, Interpret the classification results of pest images by generating text descriptions to mimic the diagnostic process of agricultural experts, which specifically includes the following steps: Step 1: Use web crawler technology to collect and construct a multimodal dataset with pest images and corresponding text descriptions; Step 2: Use the Faster-RCNN model to identify each body part of the pest image and extract the features of each body part; Step 3: Use the Transformer model to generate a text description of the pest image, and fuse the visual features and text features of the pest image to form joint features; The specific process is as follows: Step 3.1: The features representing the pest body parts F i ∈R m×2048 are input into the Encoder module of the Transformer model, and each feature is weighted through the multi-head self-attention mechanism to obtain the hidden vector representation of the feature: Fh i ∈R m×2048 , and this process is expressed by the formula: Fh i = Encoder(F i ; θ enc ), Fh i ∈R m×2048 (2); where, θ enc represents the parameter of Encoder(·); Step 3.2: Design the text input for the Decoder module of the Transformer; Let represent the text description corresponding to the pest body part P i which is the input to the Decoder module and fills the [Start] identifier at the start position of the text description; T represents the length of the text description, and L represents the length of the vocabulary; Let represent the output of the Decoder module, which is the text description generated by the Transformer model, and fill the [End] identifier at the end position of the text description; Step 3.
3. Input the text description and the hidden vector representation Fh of the pest body part features i into the first-layer Decoder module, and learn the joint representation of visual features and text features through the multi-head self-attention mechanism This process is expressed by the formula as follows: where, θ dec represents the parameter of Decoder(·); Step 3.4: After N times of stacking of the Decoder module, the output of the previous Decoder module is the input of the next Decoder module, and finally the joint feature Ft after the fusion of the visual feature and the text feature is obtained i ; Step 3.5, combined feature Ft i The probability distribution of each word is obtained through a two-layer fully connected layer module and the Softmax(·) function. This process is expressed by the formula as follows: Among them, θ gen represents the parameter of the fully connected layer module MLP gen (·); Step 3.6: Query the vocabulary based on the probability distribution of each word to obtain the final output Step 4: Use the joint features to train a classifier to predict the category label of the pest; Step 5: Combine the generated text description and the predicted pest category label to interpret the classification result of the pest at the text level.
2. The method for interpretable classification of pest images based on the image text generation technology according to claim 1, characterized in that, The specific process of Step 1 is as follows: Step 1.1: Use web crawler technology to collect pest image data in the Baidu Image Search Engine and Google Image Search Engine, and establish a pest candidate image dataset; Step 1.2: Perform image and text preprocessing to obtain the final required multimodal dataset of agricultural pests.
3. The method for interpretable classification of pest images based on image text generation technology according to claim 2, wherein The specific process of Step 1.2 is as follows: Step 1.2.1: Convert all images to the JPEG format and delete the images that cannot be displayed normally from the candidate image dataset; Step 1.2.2: Filter out the images with pixel sizes less than 448*448 and adjust the pixel sizes of all eligible images to 448*448; Step 1.2.3: Use the image annotation tool Labelme to annotate the body parts of the pest images to obtain the semantic labels of the parts, and save the annotation files in the Json format; Step 1.2.4: Invite 5 experts in the agricultural field to describe the pests based on the color, shape, material, and size of the pest body parts in each image, forming 5 text descriptions; Step 1.2.5: Combine each image with its corresponding 5 text descriptions to form an image-text pair, and finally construct a multimodal dataset of agricultural pests.
4. The method for classifying the interpretability of pest images based on the image text generation technology according to claim 1, characterized in that, The specific process of Step 2 is as follows: Step 2.
1. Represent a dataset containing n pest images as V = {V1,..., V i ,..., V n}, and use to represent the m body parts in the i-th pest image V i . Among them, the j-th body part in the i-th pest image contains two parts of information: (1) represents the label of the body part, and M represents the total number of body part categories; (2) represents the coordinates of the bounding box of the body part. Step 2.2: Input V i into the Faster-RCNN model, and obtain and two mapping functions through supervised training; which are used to identify the labels of each body part in the pest image; and which are used to recognize the bounding box of each body part. This process is expressed by the formula: Among them, θ d represents parameter, θ r represents parameter; Step 2.3: Use and two mapping functions to generate the features of m body parts: represent the features of each body part.
5. The method for classifying the interpretability of pest images based on the image text generation technology according to claim 4, wherein The specific process of Step 4 is as follows: Step 4.1: The classifier consists of N stacked ClsDecoder modules; the input of the first-layer ClsDecoder module consists of two parts. The first part is the mean of the feature of each body part of the pest image, which is used to represent the global feature of the entire pest image; the second part is the output of the first-layer Decoder module in Step 3.3 This process is expressed by the formula: Among them, θ clsdec represents the parameter of the fully connected layer module ClsDecoder(·), which is the hidden vector representation of the pest image; Step 4.2: After stacking the ClsDecoder module N times, the output of the previous ClsDecoder module serves as the input to the next ClsDecoder module, and finally, the hidden vector representation Hc of the pest image is obtained. i The probability distribution of each category is obtained through a two-layer fully connected layer module and the Softmax(·) function, and this process is expressed by the formula: Among them, θ cls represents the parameter of the fully connected layer module MLP cls (·), Z represents the total number of categories; Step 4.3: Obtain the corresponding category label according to the maximum value of the probability distribution of each category.
6. The method for classifying the interpretability of pest images based on the image text generation technology according to claim 5, wherein In Step 5, combine the text description generated in Step 3.6 and the pest category label predicted in Step 4.3 to construct a sentence pattern to achieve interpretable classification of the pest image; the specific format of the sentence pattern is: Because it is observed that so-and-so, so this pest is predicted to be of the so-and-so type.