Method for training disease detection model and related device

By training the disease detection model, using the feature fusion and multi-scale feature fusion technology of RGB and infrared images, the problem of low disease detection accuracy in complex field environments is solved, and high accuracy and robust disease recognition is achieved.

CN120198768APending Publication Date: 2025-06-24SHENZHEN SHULIAN TIANXIA INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184093.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In complex and changeable field environments, factors such as lighting conditions and leaf morphology changes affect the accuracy of crop disease detection, resulting in low detection accuracy and weak generalization ability.

Method used

A method of training disease detection model is adopted, by obtaining RGB and infrared images of crops, and marking the disease category and location, iterative training is performed using the target detection network. The object detection network includes a feature extraction module, a multi-scale feature fusion module and a detector. The feature extraction module fuses the features of RGB and infrared images, and the multi-scale feature fusion module fuses feature maps at different levels.

Benefits of technology

It improves the accuracy and robustness of crop disease detection, can effectively identify crop diseases in complex environments, and reduces the impact of environmental factors on detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198768A_ABST
    Figure CN120198768A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the crossing field of deep learning and agricultural science and technology, and discloses a method for training a disease detection model and a related device.Firstly, a plurality of training samples are obtained, the training samples comprise RGB images and infrared images of crops, and real labels reflecting disease categories and disease positions are marked; and carrying out iterative training on the target detection network by adopting the plurality of training samples until convergence to obtain a disease detection model. Wherein the target detection network comprises a feature extraction module, a multi-scale feature fusion module and a detector which are cascaded, and the feature extraction module is configured to fuse the features of the infrared image in the process of extracting the features of the RGB image and fuse the features of the RGB image in the process of extracting the features of the infrared image. Thus, the disease detection model obtained through training can effectively combine information obtained by different sensing devices, crop diseases can be accurately recognized in a complex environment, and high robustness is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the cross - field of deep learning and agricultural science and technology, and particularly to a method for training a disease detection model and related devices. Background Art

[0002] In recent years, with the development of computer vision technology and deep learning algorithms, automated crop disease detection systems have gradually become a research hotspot. These systems can automatically detect the types of diseases by analyzing the images of crop leaves or fruits taken, providing a scientific basis for precise pesticide application and disease prevention and control.

[0003] Although some research results have been applied to crop disease detection, however, in the complex and changeable field environment, factors such as lighting conditions and changes in leaf morphology will affect the detection effect. For example, there are problems such as low detection accuracy and weak generalization ability. Summary of the Invention

[0004] In view of this, the embodiments of the present application provide a method for training a disease detection model and related devices. The disease detection model trained by using this training method can effectively improve the accuracy of crop disease detection in complex environments.

[0005] In a first aspect, some embodiments of the present application provide a method for training a disease detection model, including:

[0006] Obtain a plurality of training samples, where the training samples include RGB images and infrared images of crops, and are labeled with true labels reflecting the disease category and disease location;

[0007] Iteratively train a target detection network with the plurality of training samples until convergence to obtain a disease detection model;

[0008] Among them, the target detection network includes a cascaded feature extraction module, a multi - scale feature fusion module, and a detector. The feature extraction module is configured to fuse the features of the infrared image during the process of extracting the features of the RGB image, and fuse the features of the RGB image during the process of extracting the features of the infrared image;

[0009] The multi - scale feature fusion module is configured to fuse the feature maps of different levels output by the feature extraction module; the detector is configured to convert the feature maps output by the multi - scale feature fusion module into prediction labels.

[0010] In some embodiments, the feature extraction module includes a first feature extraction branch and a second feature extraction branch. The first feature extraction branch includes N cascaded first feature extraction layers, and the second feature extraction branch includes N cascaded second feature extraction layers; the same - level first feature extraction layer and second feature extraction layer are connected by a cross - modal fusion layer;

[0011] The first feature map output by the first feature extraction layer and the second feature map output by the second feature extraction layer are input into the cross-modal fusion layer for feature fusion, and then the first cross-modal fusion feature map and the second cross-modal fusion feature map are output. The first cross-modal fusion feature map is input into the first feature extraction layer for feature extraction, and the second cross-modal fusion feature map is input into the second feature extraction layer for feature extraction.

[0012] In some embodiments, the cross-modal fusion layer performs feature fusion in the following manner:

[0013] The first feature map and the second feature map are respectively decomposed into preset equal parts along the channel dimension;

[0014] The first part of the equal parts of the first feature map and the second part of the equal parts of the second feature map are fused by channel splicing to obtain the first cross-modal fusion feature map;

[0015] The second part of the equal parts of the first feature map and the first part of the equal parts of the second feature map are fused by channel splicing to obtain the second cross-modal fusion feature map.

[0016] In some embodiments, the cross-modal fusion layer further includes a first convolution module and a second convolution module. The first convolution module is used to extract features from the first cross-modal fusion feature map, and the output feature map is input into the first feature extraction layer; the second convolution module is used to extract features from the second cross-modal fusion feature map, and the output feature map is input into the second feature extraction layer.

[0017] In some embodiments, the cross-modal fusion layer is respectively connected to a first attention fusion module and a second attention fusion module. The first attention fusion module is used to perform fusion processing on the feature map output by the first convolution module, and the output feature map is input into the first feature extraction layer; the second attention fusion module is used to perform fusion processing on the feature map output by the second convolution module, and the output feature map is input into the second feature extraction layer.

[0018] In some embodiments, the loss function configured during the training process includes a position loss and a classification loss. Among them, the position loss is used to calculate the difference in the disease position between the predicted label and the true label; the classification loss is used to calculate the difference in the disease category between the predicted label and the true label;

[0019] The classification loss is configured with a weight function. The weight function provides a first weight for easy samples and a second weight for difficult samples; among them, the first weight is less than the second weight, and the prediction accuracy of easy samples is greater than that of difficult samples.

[0020] In some embodiments, the location loss includes the intersection over union (IoU) loss and the distribution focal loss. The IoU loss is used to calculate the IoU between the disease location in the predicted label and the disease location in the ground truth label. The distribution focal loss is used to calculate the difference in probability distributions between the predicted label and the ground truth label.

[0021] In a second aspect, some embodiments of the present application provide a disease detection method, including:

[0022] Obtaining the RGB image and the infrared image of the crop;

[0023] Inputting the RGB image and the infrared image into the disease detection model to obtain the disease category and the disease location of the crop. The disease detection model is trained by using the method for training the disease detection model as described in the first aspect.

[0024] In a third aspect, some embodiments of the present application provide an electronic device, including:

[0025] At least one processor, and

[0026] A memory communicatively connected to the at least one processor, wherein,

[0027] The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to execute the method of the first aspect or the second aspect.

[0028] In a fourth aspect, some embodiments of the present application provide a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions for causing a computer device to execute the method of the first aspect or the second aspect.

[0029] The beneficial effects of the embodiments of the present application: Different from the prior art, the method for training the disease detection model provided by the embodiments of the present application first obtains a plurality of training samples. The training samples include the RGB image and the infrared image of the crop and are labeled with the ground truth labels reflecting the disease category and the disease location. Using these plurality of training samples to iteratively train the object detection network until convergence to obtain the disease detection model. The object detection network includes a cascaded feature extraction module, a multi-scale feature fusion module, and a detector. The feature extraction module is configured to fuse the features of the infrared image during the process of extracting the features of the RGB image and fuse the features of the RGB image during the process of extracting the features of the infrared image. The multi-scale feature fusion module is configured to fuse the feature maps of different levels output by the feature extraction module. The detector is configured to convert the feature maps output by the multi-scale feature fusion module into predicted labels.

[0030] In this embodiment, the training samples include RGB images and infrared images of crops. The object detection network can learn the disease characteristics of different sensor data, thereby effectively reducing the influence of environmental factors and the like on disease detection. In addition, during the process of extracting the features of the RGB image in the feature extraction module of the object detection network, the features of the infrared image are fused, and during the process of extracting the features of the infrared image, the features of the RGB image are fused. This enables the feature maps of different levels output by the feature extraction module to fully fuse the features of the two modalities, reduces the differences between cross-modal features, and enhances the consistency of the fused features. In this way, the trained disease detection model can effectively combine the information obtained by different sensing devices, accurately identify crop diseases in complex environments, and has high robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] One or more embodiments are illustrated by way of example in the accompanying drawings, which do not constitute a limitation to the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the figures do not constitute a scale limitation.

[0032] Figure 1 It is a schematic structural diagram of a disease detection system in some embodiments of the present application;

[0033] Figure 2 It is a schematic structural diagram of an electronic device in some embodiments of the present application;

[0034] Figure 3 It is a schematic flowchart of a method for training a disease detection model in some embodiments of the present application;

[0035] Figure 4 It is a schematic structural diagram of an object detection network in some embodiments of the present application;

[0036] Figure 5 It is a schematic flowchart of a disease detection method in some embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The present application will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.

[0038] In order to make the purpose, technical solution and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0039] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other, and all are within the protection scope of the present application. In addition, although functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. In addition, the terms "first", "second", "third", etc. used herein do not limit the data and the execution order, but are only used to distinguish the same items or similar items with basically the same functions and effects.

[0040] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in this specification in the description of the present application are only for the purpose of describing specific embodiments and are not used to limit the present application. The term "and / or" used in this specification includes any and all combinations of one or more of the related listed items.

[0041] In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0042] For the convenience of understanding the method provided by the embodiments of the present application, the nouns involved in the embodiments of the present application are first introduced:

[0043] (1) Neural network

[0044] A neural network can be composed of neural units. Specifically, it can be understood as a neural network with an input layer, hidden layers, and an output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the intermediate layers are all hidden layers. Among them, a neural network with many hidden layers is called a deep neural network (DNN). The operation of each layer in the neural network can be described by the mathematical expression y = a(W·x + b). Physically, the operation of each layer in the neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of the matrix) through five operations on the input space. These five operations include: 1. Dimension increase / dimension reduction; 2. Enlargement / shrinkage; 3. Rotation; 4. Translation; 5. "Bending". Among them, the operations of 2 and 3 are completed by "W·x", the operation of 4 is completed by "+b", and the operation of 5 is implemented by "a()". The reason for using the word "space" here is that the object to be classified is not a single thing, but a class of things. Space refers to the set of all individuals of this class of things. Among them, W is the weight matrix of each layer of the neural network, and each value in this matrix represents the weight value of a neuron in this layer. The matrix W determines the space transformation from the input space to the output space described above, that is, W of each layer of the neural network controls how to transform the space. The purpose of training the neural network, that is, to finally obtain the weight matrices of all layers of the trained neural network. Therefore, the training process of the neural network is essentially to learn the way to control the space transformation, and more specifically, to learn the weight matrix.

[0045] It should be noted that in the embodiments of the present application, the models adopted based on machine learning tasks are essentially neural networks. Common components in neural networks include convolutional layers, pooling layers, normalization layers, and transposed convolutional layers, etc. By assembling these common components in the neural network, a model is designed. When the model parameters (weight matrices of each layer) are determined such that the model error satisfies a preset condition or the number of adjusted model parameters reaches a preset threshold, the model converges.

[0046] Among them, the convolutional layer is configured with multiple convolutional kernels, and each convolutional kernel is set with a corresponding stride to perform a convolutional operation on the image. The purpose of the convolutional operation is to extract different features of the input image. The first convolutional layer may only be able to extract some low-level features such as edges, lines, and corners, etc. Deeper convolutional layers can iteratively extract more complex features from the low-level features. The downsampling convolutional layer is used to map a high-dimensional space to a low-dimensional space while maintaining their connection relationships / patterns (here the connection relationship refers to the connection relationship during convolution).

[0047] The transposed convolution layer (also known as the upsampling convolution layer) is used to map a low-dimensional space to a high-dimensional space while maintaining the connection relationship / pattern between them (here the connection relationship refers to the connection relationship during convolution). Similarly, the transposed convolution layer is configured with multiple convolution kernels, and each convolution kernel is set with a corresponding stride to perform deconvolution operations on the image. Generally, the upsample() function is built into the framework library (such as the PyTorch library) used to design neural networks. By calling this upsample() function, the low-dimensional to high-dimensional space mapping can be achieved.

[0048] The pooling layer is used to reduce the dimension of data or represent the image with higher-level features by mimicking the human visual system. Common operations of the pooling layer include max pooling, average pooling, stochastic pooling, median pooling, and combined pooling, etc. Generally, pooling layers are periodically inserted between the convolution layers of a neural network to achieve dimension reduction.

[0049] The normalization layer is used to perform normalization operations on all neurons in the intermediate layer to prevent gradient explosion and gradient disappearance.

[0050] (2) Loss Function

[0051] During the process of training a neural network, since it is desired that the output of the neural network is as close as possible to the value that is truly wanted to be predicted, the predicted value of the current network and the truly wanted target value can be compared, and then the weight matrix of each layer of the neural network can be updated according to the difference between them (of course, there is usually an initialization process before the first update, that is, parameters are pre-configured for each layer in the neural network). For example, if the predicted value of the network is high, the weight matrix is adjusted to make it predict lower, and continuous adjustment is made until the neural network can predict the truly wanted target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the neural network becomes a process of minimizing this loss as much as possible.

[0052] The application potential of crop disease prediction technology is extensive. It can help farmers and agricultural professionals accurately monitor pests and diseases, thereby optimizing agricultural management and decision-making. In addition, crop disease prediction can effectively guide pesticide application and increase crop yields. Machine learning technology plays an important role in crop disease detection. By training a large number of crop image samples, machine learning algorithms can learn and identify the characteristic patterns of crop diseases. Specifically, based on the input feature vectors, target detection of diseases in the image is performed, and the corresponding disease categories and disease locations are output.

[0053] Although some machine learning models have been applied to crop disease detection, however, in the complex and changeable field environment, factors such as lighting conditions and leaf morphology changes will affect the detection effect. For example, there are problems such as low detection accuracy and weak generalization ability.

[0054] To address the above problems, the embodiments of the present application provide a method for training a disease detection model. First, a number of training samples are obtained. The training samples include RGB images and infrared images of crops, and are labeled with ground truth labels reflecting disease categories and disease locations. These training samples are used to iteratively train a target detection network until convergence to obtain a disease detection model. Among them, the target detection network includes a cascaded feature extraction module, a multi-scale feature fusion module, and a detector. The feature extraction module is configured to fuse the features of the infrared image during the process of extracting the features of the RGB image, and fuse the features of the RGB image during the process of extracting the features of the infrared image. The multi-scale feature fusion module is configured to fuse the feature maps of different levels output by the feature extraction module; the detector is configured to convert the feature maps output by the multi-scale feature fusion module into prediction labels.

[0055] In this embodiment, the training samples include RGB images and infrared images of crops. The target detection network can learn the disease characteristics of different sensor data, thereby effectively reducing the influence of environmental factors and the like on recognition. In addition, in the feature extraction module of the target detection network, the features of the infrared image are fused during the process of extracting the features of the RGB image, and the features of the RGB image are fused during the process of extracting the features of the infrared image, so that the feature maps of different levels output by the extraction module fully fuse the features of the two modalities, and can also reduce the differences between cross-modal features and enhance the consistency of the fused features. In this way, the trained disease detection model can effectively combine the information obtained by different sensing devices, accurately identify crop diseases in complex environments, and has high robustness.

[0056] The following describes an exemplary application of the electronic device provided by the embodiments of the present application for training a disease detection model or for crop disease detection. It can be understood that the electronic device can both train a disease detection model and use the disease detection model to identify diseases.

[0057] The electronic device provided by an embodiment of the present application may be a server, such as a server deployed in the cloud. When the server is used to train a disease detection model, according to the training data and neural network provided by other devices or those skilled in the art, the neural network is iteratively trained using the training data to determine the final model parameters. Then, the neural network configures the final model parameters to obtain the disease detection model. When the server is used for disease detection, the built-in disease detection model is called to perform corresponding computational processing on the RGB image and infrared image of the crop provided by other devices or users, so as to obtain the disease category and disease location of the crop. Among them, the images of the crop can be obtained by drones, satellite remote sensing, or devices equipped with visible light cameras and infrared cameras.

[0058] The electronic devices provided by some embodiments of the present application may be various types of terminals such as laptop computers, desktop computers, or mobile devices. When the terminal is used to train a disease detection model, those skilled in the art input the prepared training data into the terminal and design a neural network on the terminal. The terminal iteratively trains the neural network using the training data to determine the final model parameters. Then, the neural network configures the final model parameters to obtain the disease detection model. When the terminal is used for disease detection of crops, the built-in disease detection model is called to perform corresponding computational processing on the input RGB image and infrared image of the crop, so as to obtain the disease category and disease location. The RGB image and infrared image of the crop can be obtained by drones, satellite remote sensing, or devices equipped with visible light cameras and infrared cameras.

[0059] As an example, refer to Figure 1 , Figure 1 which is a schematic diagram of the application scenario of the disease recognition system provided by an embodiment of the present application. The terminal 10 is connected to the server 20 through a network. Among them, the network can be a wide area network, a local area network, or a combination of the two.

[0060] The terminal 10 can be used to obtain training data and construct a neural network. For example, those skilled in the art download the prepared training data on the terminal and build a neural network structure for identifying disease categories and locations. In some examples, the neural network can adopt existing neural networks for object detection, such as common components including convolutional layers, fully connected layers, and softmax function layers.

[0061] It can be understood that the terminal 10 can also be used to obtain the RGB image and infrared image of the crop. For example, the visible light camera and infrared camera of the terminal 10 respectively take pictures of the leaves or fruits of the crop to obtain the RGB image and infrared image of the crop.

[0062] In some embodiments, the terminal 10 locally executes the method for training a disease detection model provided in the embodiments of the present application to complete the training of a designed neural network using training data, determine the final model parameters, so that the neural network configures the final model parameters, and thus a disease detection model can be obtained. In some embodiments, the terminal 10 can also send the training data stored on the terminal by those skilled in the art and the constructed neural network to the server 20 through the network. The server 20 receives the training data and the neural network, trains the designed neural network using the training set, determines the final model parameters, and then sends the final model parameters to the terminal 10. The terminal 10 saves the final model parameters, so that the neural network configuration can use the final model parameters, and thus a disease detection model can be obtained.

[0063] In some embodiments, the terminal 10 locally executes the disease detection method provided in the embodiments of the present application to provide a crop disease identification service for users, calls the built-in disease detection model, and performs corresponding calculation and processing on the input RGB image and infrared image of the crop to obtain the disease category and disease location of the crop. In some embodiments, the terminal 10 can also send the collected RGB image and infrared image of the crop to the server 20 through the network. The server 20 receives the RGB image and infrared image of the crop, calls the built-in disease detection model to perform corresponding calculation and processing on the RGB image and infrared image, and obtains the disease category and disease location of the crop. Among them, the RGB image and infrared image of the crop can be obtained by shooting with a drone, satellite remote sensing, or a device with a visible light camera and an infrared camera.

[0064] The structure of the electronic device in the embodiments of the present application will be described below. Figure 2 FIG. is a schematic structural diagram of an electronic device 500 in the embodiments of the present application. The electronic device 500 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. Each component in the electronic device 500 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear description, in Figure 2 all kinds of buses are labeled as the bus system 540.

[0065] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0066] The user interface 530 includes one or more output devices 531 enabling the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components facilitating user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, and other input buttons and controls.

[0067] The memory 550 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory. Optionally, the memory 550 includes one or more storage devices physically located far from the processor 510.

[0068] In some embodiments, the memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are exemplarily described below.

[0069] The operating system 551 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0070] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;

[0071] The display module 553 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 associated with the user interface 530 (such as a display screen, a speaker, etc.);

[0072] The input processing module 554 is used to detect and translate one or more user inputs or interactions from one of the one or more input devices 532.

[0073] As can be understood from the above, the methods for training a disease detection model and the disease detection methods provided in the embodiments of the present application can be implemented by various types of electronic devices with computing and processing capabilities, such as intelligent terminals and servers, etc.

[0074] The following, in combination with the exemplary applications and implementations of the server provided in the embodiments of the present application, describes the method for training a disease detection model provided in the embodiments of the present application. Refer to Figure 3 , Figure 3 which is a schematic flowchart of the method for training a disease detection model provided in the embodiments of the present application.

[0075] Please refer to Figure 3 , the training method S100 includes but is not limited to the following steps:

[0076] S10: Obtain a plurality of training samples. The training samples include RGB images and infrared images of crops, and are labeled with true labels reflecting the disease category and disease location.

[0077] Among them, the training samples include images of the same crop collected by a visible light camera and an infrared camera respectively. It can be seen that the RGB image and the infrared image respectively reflect the characteristics of the same crop in different modalities. The RGB image and the infrared image include the same part of the plant, for example, it can be the branches, leaves or fruits of the crop, etc.

[0078] Each training sample is labeled with a true label, and the true label includes the disease category and the disease location. Here, those skilled in the art label each training sample. In some embodiments, an image annotation software can be used to mark a bounding box at the disease location in the RGB image and / or infrared image, and the disease location is determined by the center coordinates, width and height of the bounding box. Each bounding box corresponds to a disease category. Among them, the image annotation software can be labelme, V7 or Labelbox, etc.

[0079] The disease category reflects the type of disease of the crop. For the same kind of crop, there may be multiple disease categories. It can be understood that a plurality of training samples cover multiple disease categories of the same kind of crop. Exemplarily, the crop is a kiwifruit tree, and the disease categories include 13 categories such as fruit rot, soft rot, aphids or canker disease, etc.

[0080] In some embodiments, the disease category can be represented by a vector. In the above example, the disease category is a vector with a length of 13, and each element in the vector corresponds to a disease category. If the element is 1, it represents the corresponding disease category, and if the element is 0, it represents not the corresponding disease category.

[0081] It can be understood that the RGB images and infrared images in these plurality of training samples can be obtained by shooting with a drone having a visible light camera and an infrared camera or other devices having a visible light camera and an infrared camera. These image data can be collected and annotated by those skilled in the art on the terminal and then sent to the server. Thus, the server obtains a plurality of training samples.

[0082] In some embodiments, the training samples can be normalized, which can effectively accelerate the convergence speed of the subsequent training of the object detection network and improve the robustness of the model. Exemplarily, the sizes of the RGB images and infrared images in the training samples are uniformly set to 640*640, and the value range of the image pixels is converted from 0-255 to between 0-1.

[0083] Specifically, the following formula is used for normalization:

[0084]

[0085] where x i represents the i-th pixel value of the image, max(x) and min(x) respectively represent the maximum and minimum values of the image pixels, and norm(x) represents the pixel value after normalization in the image.

[0086] In this embodiment, by normalizing the training samples, the convergence speed of the model can be accelerated and the robustness of the model can be improved.

[0087] For the sake of convenience of expression, the training samples after normalization are also collectively referred to as training samples. That is to say, in the following step S20, several training samples used for inputting into the object detection network for training can be the training samples after normalization.

[0088] S20: Use several training samples to iteratively train the object detection network until convergence to obtain a disease detection model.

[0089] Among them, the object detection network is a pre-set neural network with components of a neural network (convolutional layer, deconvolutional layer, pooling layer, activation function, etc.). The basic structure and principle of the neural network have been described in detail in "Noun Introduction (1)", and will not be elaborated here.

[0090] The object detection network can be constructed by those skilled in the art on a neural network design platform on a terminal (such as a computer) and then sent to the server. In some implementable ways, an existing neural network for object detection can also be used as the object detection network. Exemplarily, the object detection network can be an object detection network of the YOLO series, such as YOLO_V8, etc.

[0091] Here, several training samples are used to iteratively train the object detection network, and the model parameters at the time of convergence are configured as the parameters of the object detection network to obtain the disease detection model. Those skilled in the art can understand that model training is prior art, and the specific training and parameter adjustment process will not be elaborated here.

[0092] It can be understood that the detection ability of the disease detection model is determined by the training data (i.e., a number of training samples), the network structure (i.e., the structure of the object detection network), and the loss function. In this embodiment, the training samples include the RGB images and infrared images of crops. The object detection network can learn the disease characteristics of different sensor data, thereby effectively reducing the influence of environmental factors and the like on disease detection. The specific structure of the object detection network is introduced below.

[0093] The object detection network includes a cascaded feature extraction module, a multi-scale feature fusion module, and a detector. As Figure 4 shown, from the input end to the output end of the object detection network, there are successively a feature extraction module, a multi-scale feature fusion module, and a detector.

[0094] For any training sample, the training sample is first input into the feature extraction module. The feature extraction module extracts features from the RGB image and infrared image in the training sample, fuses the features of the infrared image during the process of extracting the features of the RGB image, and fuses the features of the RGB image during the process of extracting the features of the infrared image. The feature extraction module outputs feature maps of different levels, and these feature maps of different levels are then input into the multi-scale feature fusion module for feature fusion. The feature map output by the multi-scale feature fusion module is then input into the detector to obtain the predicted label through conversion. The multi-scale feature fusion module and the detector will be introduced below.

[0095] In this embodiment, in the feature extraction module of the object detection network, the features of the infrared image are fused during the process of extracting the features of the RGB image, and the features of the RGB image are fused during the process of extracting the features of the infrared image, so that the feature maps of different levels output by the extraction module fully fuse the features of the two modalities, can also reduce the differences between cross-modal features, and enhance the consistency of the fused features. In this way, the trained disease detection model can effectively combine the information obtained by different sensing devices, accurately identify crop diseases in complex environments, and has high robustness.

[0096] In some embodiments, the feature extraction module includes a first feature extraction branch and a second feature extraction branch. The first feature extraction branch includes N cascaded first feature extraction layers, and the second feature extraction branch includes N cascaded second feature extraction layers; the first feature extraction layer and the second feature extraction layer of the same level are connected by a cross-modal fusion layer, where 1 ≤ i ≤ N.

[0097] As Figure 4 shown, the first feature extraction branch includes 5 first feature extraction layers for extracting features from the infrared image. The second feature extraction branch includes 5 second feature extraction layers for extracting features from the RGB image.

[0098] The first feature extraction layer and the second feature extraction layer respectively perform downsampling using a convolutional kernel of size 3*3 to extract feature maps of different scales. The feature maps output by the first feature extraction layer and the second feature extraction layer at the same level have the same size.

[0099] For a given input RGB image I RGB , it is input into the first feature extraction branch, and after a series of convolutional downsampling processes, local features are extracted. The process of the first feature extraction branch processing the RGB image I RGB can be represented by the following formula:

[0100] H Ri = ρ i ...(ρ2(ρ1(I RGB )))(i = 1, 2, 3, 4, 5)

[0101] where I RGB represents the RGB image in the training sample, i is the layer number of the first feature extraction layer, ρ1 is the first first feature extraction layer, ρ2 is the second first feature extraction layer, ρ i is the i-th first feature extraction layer, and H Ri is the RGB feature map output by the i-th first feature extraction layer.

[0102] For a given input infrared image I IR , it is input into the second feature extraction branch, and after a series of convolutional downsampling processes, local features are extracted. The process of the second feature extraction branch processing the infrared image I IR can be represented by the following formula:

[0103] H IRi = f i ...(f2(f1(I IR )))(i = 1, 2, 3, 4, 5)

[0104] where I IR represents the infrared image in the training sample, i is the layer number of the second feature extraction layer, f1 is the first second feature extraction layer, f2 is the second second feature extraction layer, ρ i is the i-th second feature extraction layer, and H IRi is the infrared feature map output by the i-th second feature extraction layer.

[0105] Please refer to again Figure 4 , the first feature extraction layer ρ i and the second feature extraction layer f i at the same level are connected by a cross-modal fusion layer Fusion. The i-th first feature map H i output by the i-th first feature extraction layer ρ Riand the i-th second feature extraction layer f i The output second feature map H IRi The input cross-modal fusion layer Fusion performs feature fusion and outputs the first cross-modal fusion feature map and the second cross-modal fusion feature map. The first cross-modal fusion feature map is input into the first feature extraction layer ρ i for feature extraction, and the second cross-modal fusion feature map is input into the second feature extraction layer f i for feature extraction.

[0106] In this embodiment, by setting a cross-modal fusion layer between the first feature extraction layer and the second feature extraction layer at the same level, the RGB feature map and the infrared feature map at each level are fused. In this way, the infrared features are fused level by level during the process of extracting RBG features, and the RGB features are fused level by level during the process of extracting infrared features, thereby being able to improve the fusion effect of the two modal features, being beneficial to reducing the difference between cross-modal features, and enhancing the consistency of the fusion features.

[0107] In some embodiments, the cross-modal fusion layer performs feature fusion in the following manner:

[0108] (1) The first feature map and the second feature map are respectively decomposed into preset equal parts according to the channel dimension.

[0109] (2) The first equal part of the first feature map and the second equal part of the second feature map are subjected to channel splicing fusion to obtain the first cross-modal fusion feature map.

[0110] (3) The second equal part of the first feature map and the first equal part of the second feature map are subjected to channel splicing fusion to obtain the second cross-modal fusion feature map.

[0111] Among them, the preset equal parts can be determined based on the number of channels of each first feature map and second feature map. Exemplarily, the number of channels of the 1st to 5th first feature maps are 18, 36, 72, 144, and 288 in sequence, and the number of channels of the 1st to 5th second feature maps are 18, 36, 72, 144, and 288 in sequence. It can be seen that the number of channels of each first feature map and second feature map is a multiple of 6. Therefore, the preset equal parts can be 6 equal parts.

[0112] That is, the first feature map is decomposed into 6 equal parts according to the channel dimension, and the second feature map is decomposed into 6 equal parts according to the channel dimension. Then, the first equal part of the first feature map and the second equal part of the second feature map are subjected to channel splicing fusion to obtain the first cross-modal fusion feature map. The second equal part of the first feature map and the first equal part of the second feature map are subjected to channel splicing fusion to obtain the second cross-modal fusion feature map.

[0113] Among them, the first part of the equal divisions may include the first, third, and fifth equal divisions among the above-mentioned 6 equal divisions, that is, the first part of the equal divisions includes odd equal divisions. The second part of the equal divisions may include the second, fourth, and sixth equal divisions among the above-mentioned 6 equal divisions, that is, the second part of the equal divisions includes even equal divisions.

[0114] Exemplarily, the first, third, and fifth equal divisions of the first feature map H Ri , and the second, fourth, and sixth equal divisions of the second feature map H IRi , are combined in a way of channel splicing and fusion to form the first cross-modal fusion feature map T Ri . The second, fourth, and sixth equal divisions of the first feature map H Ri , and the first, third, and fifth equal divisions of the second feature map H IRi , are combined in a way of channel splicing and fusion to form the second cross-modal fusion feature map T IRi .

[0115] The first cross-modal fusion feature map T Ri can be represented by the following formula:

[0116] T Ri =(H Ri(1,3,6) ,H IRi(2,4,6) )

[0117] where H Ri(1,3,5) is the first, third, and fifth equal divisions of the first feature map H Ri ; H IRi(2,4,6) is the second, fourth, and sixth equal divisions of the second feature map H IRi .

[0118] The second cross-modal fusion feature map T IRi can be represented by the following formula:

[0119] T IRi =(H IRi(1,3,5) ,H Ri(2,4,6) )

[0120] where H IRi(1,3,5) is the first, third, and fifth equal divisions of the second feature map H IRi , and H Ri(2,4,6) is the second, fourth, and sixth equal divisions of the first feature map H Ri .

[0121] In this embodiment, through the methods of channel decomposition and cross-splicing fusion, both the first cross-modal fusion feature map and the second cross-modal fusion feature map have the features of two modalities. That is, by integrating the information of different channels to construct cross-modal features, it is beneficial to improve the fusion effect of the features of the two modalities.

[0122] In some embodiments, the cross-modal fusion layer further includes a first convolutional module and a second convolutional module. The first convolutional module is used to extract features from the first cross-modal fusion feature map, and the output feature map is input into the first feature extraction layer. The second convolutional module is used to extract features from the second cross-modal fusion feature map, and the output feature map is input into the second feature extraction layer.

[0123] Among them, the first convolutional module includes multiple convolutional layers. Among them, the convolutional kernel configured by the convolutional layer can be a convolutional kernel with a size of 3*3 or a convolutional kernel with a size of 1*1. The first convolutional module further extracts features from the first cross-modal fusion feature map, so that the extracted feature map has a better cross-modal feature fusion effect.

[0124] In some embodiments, the first convolutional module includes a convolutional layer configured with a 3*3 convolutional kernel and a convolutional layer configured with a 1*1 convolutional kernel. The feature extraction process of the first convolutional module can be represented by the following formula:

[0125]

[0126] Among them, T Ri is the first cross-modal fusion feature map, W 3*3 () represents the convolutional process of the convolutional layer with a 3*3 convolutional kernel, and W 1*1 () represents the convolutional process of the convolutional layer with a 1*1 convolutional kernel. is the feature map output by the first convolutional module.

[0127] Similarly, the second convolutional module includes multiple convolutional layers. Among them, the convolutional kernel configured by the convolutional layer can be a convolutional kernel with a size of 3*3 or a convolutional kernel with a size of 1*1. The second convolutional module further extracts features from the second cross-modal fusion feature map, so that the extracted feature map has a better cross-modal feature fusion effect.

[0128] In some embodiments, the second convolutional module includes a convolutional layer configured with a 3*3 convolutional kernel and a convolutional layer configured with a 1*1 convolutional kernel. The feature extraction process of the second convolutional module can be represented by the following formula:

[0129]

[0130] Among them, T IRi is the second cross-modal fusion feature map, W 3*3 () represents the convolutional process of the convolutional layer with a 3*3 convolutional kernel, and W 1*1 () represents the convolutional process of the convolutional layer with a 1*1 convolutional kernel. is the feature map output by the second convolutional module.

[0131] In this embodiment, the first convolution module and the second convolution module are further used to extract features from the two cross-modal fusion feature maps respectively, which is conducive to integrating information from different channels to construct cross-modal features and improving the fusion effect of the two modal features.

[0132] In some embodiments, the cross-modal fusion layer is respectively connected to the first attention fusion module and the second attention fusion module. The first attention fusion module is used to perform fusion processing on the feature map output by the first convolution module, and the output feature map is input into the first feature extraction layer; the second attention fusion module is used to perform fusion processing on the feature map output by the second convolution module, and the output feature map is input into the second feature extraction layer.

[0133] Please refer to again Figure 4 , a first attention fusion module is provided on the transmission loop from the cross-modal fusion layer to the first feature extraction layer. A second attention fusion module is provided on the transmission loop from the cross-modal fusion layer to the second feature extraction layer.

[0134] For the feature map output by the first convolution module It is input into the first attention fusion module for fusion processing. After the processed feature map is input into the first feature extraction layer for feature extraction, it is input into the first feature extraction layer of the next level.

[0135] Among them, the first attention fusion module assigns different weights to different features, and then fuses the features according to the weights, enabling the model to focus on key information (features with large weights). Thus, the sensitivity of the model to key information is improved, which is conducive to the rapid convergence of the model.

[0136] The working mechanism of the first attention fusion module can be expressed by the following formula:

[0137]

[0138] Among them, represents using a convolutional layer with a configured 1*1 convolution kernel to extract features from the input feature map and convert it into the feature map Q Ri . For the feature map Q Ri , GlobalPooling (global pooling layer) is used for spatial compression, reducing the number of channels to a 1-channel feature map. Since the feature map Q Ri is spatially compressed, the Softmax function is used to enhance the information of the feature map Q Ri .

[0139] represents using a convolutional layer with a configured 1*1 convolution kernel to extract features from the input feature map and convert it into the feature map V Ri . The feature map VRi The size remains the same as the original input size.

[0140] Then, the enhanced feature map Q Ri and the feature map V Ri are subjected to matrix multiplication, and then relu operation is performed by the activation function to obtain the feature map In the feature map all parameters are kept between 0 and 1. After the feature map is backpropagated to the corresponding first feature extraction layer for feature extraction, it is input into the first feature extraction layer of the next level.

[0141] Similarly, for the feature map output by the second convolutional module it is input into the second attention fusion module for fusion processing, and the processed feature map is input into the second feature extraction layer for feature extraction and then input into the second feature extraction layer of the next level.

[0142] Among them, the second attention fusion module assigns different weights to different features, and then fuses the features according to the weights, enabling the model to focus on key information (features with large weights), thereby improving the model's sensitivity to key information and facilitating the model's rapid convergence.

[0143] The working mechanism of the second attention fusion module can be expressed by the following formula:

[0144]

[0145] Among them, represents that the convolutional layer with a configured 1*1 convolutional kernel is used to perform feature extraction on the input feature map to convert it into the feature map Q IRi . For the feature map Q IRi , GlobalPooling (global pooling layer) is used for spatial compression, reducing the number of channels to a 1-channel feature map. Since the feature Q IRi is spatially compressed, the Softmax function is used to enhance the information of the feature map Q IRi .

[0146] represents that the convolutional layer with a configured 1*1 convolutional kernel is used to perform feature extraction on the input feature map to convert it into the feature map V IRi . The size of the feature map V IRi remains the same as the original input size.

[0147] Then, the enhanced feature map Q IRi and the feature map V IRiPerform matrix multiplication, and then perform relu operation by the activation function to obtain the feature map In the feature map All parameters are kept between 0 and 1. After the feature map is backpropagated to the corresponding second feature extraction layer for feature extraction, it is input into the second feature extraction layer of the next level.

[0148] In this embodiment, the first attention fusion module and the second attention fusion module are used to further extract the deep fusion features of the fusion features, fully learn the feature information of different sensors, enable the model to focus on key information, thereby improving the sensitivity of the model to key information and facilitating the rapid convergence of the model.

[0149] Please refer to again Figure 4 , the first feature maps output by the 3rd to 6th first feature extraction layers in the first feature extraction branch of the feature extraction module (such as Figure 4 P3, P4, P5 in), that is, the first feature maps at different levels, are input into the multi-scale fusion module to further perform feature extraction, feature fusion, feature enhancement and feature matching. In this way, the accuracy and robustness of object detection can be effectively improved.

[0150] Among them, the multi-scale fusion module includes a Path Aggregation Feature Pyramid Network (PAFPN), which can enhance the information flow between different levels in the feature pyramid and improve the performance of object detection. For the specific structure of the multi-scale fusion module, please refer to the Neck part in the YOLOv8 network, which will not be introduced in detail here.

[0151] It can be understood that the first feature maps based on multiple scales fuse two modalities of features. Therefore, the feature maps at multiple scales output by the multi-scale fusion module still have two modalities of features. Thus, the detector can output prediction labels based on the feature maps at multiple scales output by the multi-scale fusion module, that is, the disease location, disease category and disease confidence. For the specific structure of the detector, please refer to the Head part in the YOLOv8 network, which will not be introduced in detail here.

[0152] Use the loss function to calculate the loss between the true label and the prediction label. The losses corresponding to several training samples and backpropagation are carried out. Based on the loss and continuously optimize the model parameters, constrain the model to develop in the direction of small loss sum until convergence, and obtain the disease detection model.

[0153] In some embodiments, the loss function configured during the training process includes a location loss and a classification loss. Among them, the location loss is used to calculate the difference in the disease location between the predicted label and the true label; the classification loss is used to calculate the difference in the disease category between the predicted label and the true label.

[0154] In some embodiments, the classification loss calculates the cross-entropy between the class probability of each predicted box in each cell and the true class. The classification loss can be characterized by the following formula:

[0155]

[0156] where S is the size of the feature map of the input detector, B is the number of predicted boxes in each cell, δ represents the class weight, represents the binary label indicating whether the j-th predicted box in the i-th cell contains an object. p i,j (c) is the true probability that the j-th predicted box in the i-th cell belongs to the c-th class, is the predicted probability that the j-th predicted box in the i-th cell belongs to the c-th class predicted by the model.

[0157] In some embodiments, the classification loss is configured with a weight function. The weight function provides a first weight for easy samples and a second weight for hard samples; among them, the first weight is less than the second weight, and the prediction accuracy of easy samples is greater than that of hard samples.

[0158] Among them, easy samples refer to training samples for which the object detection network during the training process can accurately identify the disease category and disease location, such as samples with high confidence. Hard samples refer to training samples for which the object detection network during the training process has difficulty accurately identifying the disease category and disease location, such as samples with low confidence. That is, the prediction accuracy of easy samples is greater than that of hard samples.

[0159] Based on the first weight being less than the second weight, thus, a smaller weight is assigned to easy samples and a larger weight is assigned to hard samples, enabling the model to pay more attention to the learning of hard samples on the basis of training easy samples well.

[0160] In some embodiments, the weight function is represented by the following formula:

[0161]

[0162] where δ represents the class weight, σ is the sigmoid function, and IOU i,j is the confidence.

[0163] In the case of y = c, the training sample is a positive sample. When IOU i,j approaches 1, The value at this time is close to 0, indicating that the positive sample is correctly detected during the object detection process, and it is an easy sample. Thus, δ is smaller, and this positive sample has a smaller weight. When the IOU i,j approaches 0, it is close to 1, indicating that the positive sample is incorrectly detected during the object detection process, and it is a difficult sample. Thus, δ is larger, and this positive sample has a larger weight.

[0164] In the case of y!= c, the training sample is a negative sample. When the IOU i,j approaches 1, it indicates that the negative sample is correctly detected during the object detection process, and it is an easy sample. Thus, δ is smaller, and this negative sample has a smaller weight. When the IOU i,j approaches 0, it indicates that the negative sample is incorrectly detected during the object detection process, and it is a difficult sample. Thus, δ is larger, and this negative sample has a larger weight.

[0165] In this embodiment, constructing the above dynamic class weights can enable the model to focus more on learning the features of the disease classes with classification errors, thus facilitating the improvement of the detection accuracy of the model.

[0166] In some embodiments, the location loss includes the intersection over union loss and the distribution focal loss. Among them, the intersection over union loss is used to calculate the intersection over union between the disease location in the predicted label and the disease location in the true label; the distribution focal loss is used to calculate the probability distribution difference between the predicted label and the true label.

[0167] In this embodiment, the location loss can be expressed by the following formula:

[0168]

[0169]

[0170] Among them, S is the size of the feature map input to the detector, B is the number of predicted boxes in each cell, represents the binary label indicating whether the j-th predicted box in the i-th cell contains the target. x i and y i and w i and h i represent the center point coordinates and width and height of the i-th cell, represents the center coordinates and width and height of the predicted box, is the CIOU (Complete Intersection over Union) function, which is used to calculate the IoU value between the predicted box and the true box. λ coord is the weight coefficient of the localization loss, which is used to balance the ratio between the localization loss and the classification loss, and improve the localization accuracy and overall accuracy of object detection.

[0171] Among them, -((y i+1 -y))log(S i )+((y - y i ))log(S i+1 ) is the distribution focal loss. By comparing the predicted probability distribution with the distribution of the true labels, it enables the network to focus on the values near the target y faster and increase their probabilities. In addition, it optimizes the probability at the position closest to the label y in the form of cross-entropy, thus enabling the network to focus faster on the distribution in the neighboring region of the target position.

[0172] In some embodiments, the loss function further includes a confidence loss, which represents the IOU loss between the predicted target bounding box and the true bounding box. The confidence loss can be expressed by the following formula:

[0173]

[0174] Among them, Loss conf is the confidence loss, S is the size of the feature map input to the detector, B is the number of predicted boxes per cell, and IOU i,j is the confidence. Indicates 1 if the j-th predicted box in the i-th cell contains an object, and 0 otherwise. Indicates 0 if the j-th predicted box in the i-th cell contains an object, and 1 otherwise.

[0175] In this embodiment, using the confidence loss to measure the prediction error of the model for the presence of a target in each cell is beneficial for the model to learn how to more accurately predict the probability of the presence of a disease and improve the accuracy of detection.

[0176] In summary, in the method for training a disease detection model provided in the embodiments of the present application, the training samples include RGB images and infrared images of crops. The object detection network can learn the disease characteristics of different sensor data, thereby effectively reducing the influence of environmental factors and the like on disease detection. In addition, during the process of extracting the features of the RGB image, the feature extraction module in the object detection network fuses the features of the infrared image, and during the process of extracting the features of the infrared image, it fuses the features of the RGB image, enabling the feature maps of different levels output by the feature extraction module to fully fuse the features of the two modalities, and also reducing the differences between cross-modal features and enhancing the consistency of the fused features. In this way, the trained disease detection model can effectively combine the information obtained by different sensing devices, accurately identify crop diseases in a complex environment, and has high robustness.

[0177] After training a disease detection model by using the method for training a disease detection model provided in an embodiment of the present application, the disease detection model can be used to intelligently identify the disease category and location of crops. The disease detection method provided in the embodiment of the present application can be implemented by various types of electronic devices with computing and processing capabilities, such as intelligent terminals and servers.

[0178] The following describes the disease detection method provided in the embodiment of the present application in combination with the exemplary applications and implementations of the terminal provided in the embodiment of the present application. Refer to Figure 5 , Figure 5 which is a schematic flowchart of the disease detection method provided in the embodiment of the present application. The method S200 includes the following steps:

[0179] S201: Obtain the RGB image and infrared image of the crop.

[0180] The RGB image and infrared image can be obtained by a terminal (such as a tablet computer) equipped with a visible light camera and an infrared camera, or can be sent to the terminal after being taken by other devices with a visible light camera and an infrared camera. Among them, other devices with a visible light camera and an infrared camera can be an unmanned aerial vehicle (UAV) or the like.

[0181] Thus, the disease detection assistant (application program) built in the terminal (such as a tablet computer) can obtain the RGB image and infrared image of the crop.

[0182] S202: Input the RGB image and infrared image into the disease detection model to obtain the disease category and location of the crop.

[0183] Among them, the disease detection model is trained by using any one of the methods for training a disease detection model in the above embodiments. Input the RGB image and infrared image into the disease detection model, and the disease detection model outputs the corresponding disease category and location.

[0184] It can be understood that the disease detection model is trained by using the method for training a disease detection model in the above embodiments, and has the same structure and function as the disease detection model in the above embodiments, and will not be elaborated here one by one.

[0185] The embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions for causing an electronic device to execute the method for training a disease detection model provided in the embodiment of the present application. For example, as Figure 3-4 shown in the method for training a disease detection model, or the disease detection method provided in the embodiment of the present application. For example, as Figure 5 shown in the disease detection method.

[0186] In some embodiments, the storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0187] In some embodiments, the executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0188] As an example, the executable instructions may or may not correspond to files in the file system, may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hypertext Markup Language (HTML) document, stored in a single file dedicated to the program in question, or stored in multiple cooperating files (such as files that store one or more modules, subroutines, or portions of code).

[0189] As an example, the executable instructions may be deployed to execute on one computing device (including devices such as smart terminals and servers), or on multiple computing devices located at one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.

[0190] The embodiments of the present application also provide a computer-readable storage medium that stores a computer program, where the computer program includes program instructions that, when executed by a computer, cause the computer to execute the method for training a disease detection model or the disease detection method as described in the foregoing embodiments.

[0191] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0192] Through the description of the above embodiments, those of ordinary skill in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Those of ordinary skill in the art can understand that all or part of the processes in the above-described method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, and the steps can be implemented in any order, and there are many other variations in different aspects of the present application as described above. For the sake of brevity, they are not provided in detail; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for training a disease detection model, characterized in that: include: Acquire a number of training samples, wherein the training samples include RGB images and infrared images of crops and are annotated with real labels reflecting disease categories and disease locations; Iteratively training the target detection network using the plurality of training samples until convergence, thereby obtaining the disease detection model; The target detection network includes a cascaded feature extraction module, a multi-scale feature fusion module and a detector, wherein the feature extraction module is configured to fuse the features of the infrared image in the process of extracting the features of the RGB image, and to fuse the features of the RGB image in the process of extracting the features of the infrared image; The multi-scale feature fusion module is configured to fuse feature maps of different levels output by the feature extraction module; and the detector is configured to convert the feature maps output by the multi-scale feature fusion module into predicted labels.

2. The method according to claim 1, characterized in that The feature extraction module includes a first feature extraction branch and a second feature extraction branch, the first feature extraction branch includes N cascaded first feature extraction layers, and the second feature extraction branch includes N cascaded second feature extraction layers; the first feature extraction layer and the second feature extraction layer at the same level are connected by a cross-modal fusion layer; The first feature map output by the first feature extraction layer and the second feature map output by the second feature extraction layer are input into the cross-modal fusion layer for feature fusion to output a first cross-modal fusion feature map and a second cross-modal fusion feature map. The first cross-modal fusion feature map is input into the first feature extraction layer for feature extraction, and the second cross-modal fusion feature map is input into the second feature extraction layer for feature extraction.

3. The method according to claim 2, characterized in that The cross-modal fusion layer performs feature fusion in the following manner: Decomposing the first feature map and the second feature map into preset equal parts according to the channel dimension respectively; Equally divide the first part of the first feature map and the second part of the second feature map, perform channel splicing and fusion, and obtain the first cross-modal fusion feature map; The second part of the first feature map is equally divided into the first part of the second feature map, and channel splicing and fusing are performed to obtain the second cross-modal fusion feature map.

4. The method according to claim 3, characterized in that The cross-modal fusion layer also includes a first convolution module and a second convolution module. The first convolution module is used to perform feature extraction on the first cross-modal fusion feature map, and the output feature map is input into the first feature extraction layer; the second convolution module is used to perform feature extraction on the second cross-modal fusion feature map, and the output feature map is input into the second feature extraction layer.

5. The method according to claim 4, characterized in that The cross-modal fusion layer is respectively connected to a first attention fusion module and a second attention fusion module, wherein the first attention fusion module is used to perform fusion processing on the feature map output by the first convolution module, and the output feature map is input into the first feature extraction layer; The second attention fusion module is used to perform fusion processing on the feature map output by the second convolution module, and the output feature map is input into the second feature extraction layer.

6. The method according to any one of claims 1 to 5, characterized in that: The loss function configured during the training process includes position loss and classification loss, wherein the position loss is used to calculate the difference in disease position between the predicted label and the true label; the classification loss is used to calculate the difference in disease category between the predicted label and the true label; The classification loss is configured with a weight function, which provides a first weight for easy samples and a second weight for difficult samples; wherein the first weight is less than the second weight, and the prediction accuracy of the easy samples is greater than the prediction accuracy of the difficult samples.

7. The method according to claim 6, characterized in that The position loss includes the intersection-and-parallel ratio loss and the distribution focus loss, wherein the intersection-and-parallel ratio loss is used to calculate the intersection-and-parallel ratio between the disease position in the predicted label and the disease position in the true label; the distribution focus loss is used to calculate the probability distribution difference between the predicted label and the true label.

8. A disease detection method, characterized in that: include: Obtain RGB and infrared images of crops; The RGB image and the infrared image are input into a disease detection model to obtain the disease category and disease location of the crop; wherein the disease detection model is trained using any one of the methods for training a disease detection model as described in claims 1-7.

9. An electronic device, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor, wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer device to execute the method according to any one of claims 1 to 8.