Crop monitoring method, electronic equipment and computer readable storage medium

By collecting and analyzing multimodal image sets of crops, the problem of low efficiency in monitoring crop diseases and pests in existing technologies has been solved, enabling earlier and more accurate monitoring of diseases and pests.

CN121811094APending Publication Date: 2026-04-07ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Current technologies for monitoring crop diseases and pests rely on manual field inspections and experience-based judgment, which are inefficient and make it difficult to provide early warnings.

Method used

By collecting multimodal image sets of crops, performing feature extraction and analysis, and determining the changing characteristics of crop status trends, automatic monitoring of pests and diseases can be achieved.

Benefits of technology

This has improved the accuracy and efficiency of crop pest and disease monitoring, enabling earlier and more accurate monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811094A_ABST
    Figure CN121811094A_ABST
Patent Text Reader

Abstract

The invention discloses a crop monitoring method, electronic equipment and a computer readable storage medium. The method comprises the following steps: acquiring at least two image groups acquired from crops; wherein the image group comprises at least one modal image, the different modal images in the same image group are acquired from the crops at the same moment, and the same modal images in the different image groups are acquired from the crops at different moments; performing feature extraction on each image group to obtain image group features corresponding to each image group; analyzing the features of each image group to obtain change features; wherein the change features are used for representing the change trend of the image group features; monitoring the crops based on the change characteristics to obtain a crop monitoring result; wherein the crop monitoring result comprises a first result, and the first result is used for representing whether diseases and pests exist on the crops or not. In this way, the accuracy of crop monitoring can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a crop monitoring method, electronic device, and computer-readable storage medium. Background Technology

[0002] In agricultural production, pests and diseases are one of the main threats to crop yield reduction or even crop failure. Currently, most agricultural pest and disease monitoring is conducted using traditional methods, which rely heavily on manual field inspections and experience-based judgment. These methods are inefficient and make it difficult to provide early warnings. Summary of the Invention

[0003] The main technical problem addressed by this application is to provide a crop monitoring method, electronic device, and computer-readable storage medium that can improve the accuracy of crop monitoring.

[0004] To address the aforementioned technical problems, the first aspect of this application provides a crop monitoring method, comprising: acquiring at least two image sets collected from crops; wherein the image set includes at least one modal image, different modal images in the same image set are collected from crops at the same time, and the same modal images in different image sets are collected from crops at different times; extracting features from each image set to obtain image set features corresponding to each image set; analyzing the features of each image set to obtain variation features; wherein the variation features are used to characterize the variation trend of the image set features; monitoring crops based on the variation features to obtain crop monitoring results; wherein the crop monitoring results include a first result, the first result being used to characterize whether pests or diseases exist on the crops.

[0005] To solve the above-mentioned technical problems, a second aspect of this application provides an electronic device, which includes a memory and a processor. The memory is used to store program instructions, and the processor is used to execute the program instructions to implement the above-mentioned crop monitoring method.

[0006] To address the aforementioned technical problems, a third aspect of this application provides a computer-readable storage medium for storing program instructions that can be executed to implement the aforementioned crop monitoring method.

[0007] The aforementioned technical solution, by acquiring time-series image sets, can determine the changing characteristics representing the trend of crop state changes based on these image sets. This enables the monitoring of crop diseases and pests based on these changing characteristics, improving the accuracy of monitoring and achieving earlier and more accurate automatic monitoring. Furthermore, compared to relying on manual field inspections and experience-based judgment for crop disease and pest monitoring, this solution improves the efficiency of crop disease and pest monitoring. Attached Figure Description

[0008] Figure 1 This is a schematic flowchart of an embodiment of the crop monitoring method provided in this application; Figure 2 This is a schematic diagram of the framework of an embodiment of the crop monitoring model provided in this application; Figure 3 This is a schematic diagram of the framework of an embodiment of the feature analysis and classification module provided in this application; Figure 4 yes Figure 1 The flowchart of step S12 shown is a schematic diagram of one embodiment; Figure 5 This is a schematic diagram of the framework of the first feature extraction module provided in this application; Figure 6 This is a schematic diagram of the framework of an embodiment of the feature attention mechanism module provided in this application; Figure 7 This is a schematic diagram of the structure of an embodiment of the crop monitoring device provided in this application; Figure 8 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application; Figure 9 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation

[0009] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0010] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0011] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0012] Please see Figure 1 , Figure 1 This is a schematic flowchart of an embodiment of the crop monitoring method provided in this application. It should be noted that if substantially the same result is obtained, this embodiment is not necessarily identical. Figure 1The illustrated process sequence is limited. For example... Figure 1 As shown, this embodiment includes: Step S11: Obtain at least two sets of images of the crops.

[0013] In this embodiment, at least two image groups are obtained from the crop; wherein, the image group includes at least one modal image, different modal images in the same image group are obtained from the crop at the same time, and the same modal image in different image groups are obtained from the crop at different times.

[0014] Images of the same modality within different image sets were collected from crops at different times. Therefore, at least two image sets are time-series image sets, meaning at least two image sets are collected from crops at different times. By collecting time-series image sets, subsequent monitoring can be conducted based on the changing characteristics representing the trend of crop state changes. This allows for the monitoring of crop diseases and pests based on these changing characteristics, improving the accuracy of crop disease and pest monitoring and enabling earlier and more accurate automatic monitoring. Furthermore, compared to relying on manual field inspections and experience-based judgment for crop disease and pest monitoring, this method improves the efficiency of crop disease and pest monitoring.

[0015] There is no limit to the number of modal images contained in an image group; it can be set according to actual usage needs. For example, if an image group contains only one type of modal image, then the image group is a single-modal image group. Or, if an image group contains images of multiple modalities, then the image group is a multimodal image group.

[0016] In one embodiment, at least one modal image includes at least one of the following: a visible light image, an infrared image, a humidity distribution thermal image, and a frame spectrogram image, wherein the humidity distribution thermal image is obtained based on statistics of humidity data collected from crops, and the frame spectrogram image is obtained based on statistics of audio data collected from crops.

[0017] In one specific implementation, a conventional camera can be used to capture images of crops to obtain visible light images.

[0018] In one specific implementation, an infrared camera can be used to capture images of crops to obtain infrared images.

[0019] In one specific implementation, humidity data for crops can be obtained using a humidity sensor. The humidity sensor can be a soil humidity sensor or an air humidity sensor, etc., and is not limited thereto.

[0020] In one specific implementation, humidity sensors can be deployed in an n-meter by n-meter grid across the camera's field of view, with a humidity sensor positioned at the center of each grid. This forms a humidity sensor array, which is used to collect humidity data about crops. In other words, the camera's monitoring area is divided into a uniform grid, and a humidity sensor is deployed at the center of each grid. By deploying humidity sensors in an n-meter by n-meter grid across the camera's field of view, humidity data for the entire area can be collected systematically and uniformly.

[0021] In one specific implementation, the humidity data for crops is a combination of N node coordinates and humidity values ​​collected by a humidity sensor array. This allows the node coordinates and corresponding humidity values ​​on the humidity sensor array to be converted into a humidity distribution thermal image that visually represents both the humidity values ​​and their distribution. The definition of the humidity distribution thermal image is as follows:

[0022]

[0023] Where F(x) represents the value of the humidity distribution thermal image at location x; k i This represents the humidity value of the i-th node; x i This represents the position or coordinates of the i-th node; Represent a delta function; This represents the convolution operation; Let the Gaussian kernel function have a standard deviation of . ; The standard deviation of the adaptive kernel is defined as follows: ; β This represents a proportionality constant; This represents the average distance associated with the i-th node. It should be noted that, to adapt to different scenarios, the Gaussian kernel function uses an adaptive kernel. The kernel size varies depending on the scale of the scene window; that is, the kernel size increases with the vertical axis of the image.

[0024] In one specific implementation, audio data about crops can be acquired using a microphone.

[0025] In one specific implementation, a microphone can be deployed at the center of the camera's field of view to collect audio data about the crops. In other words, the microphone is deployed at the center of the camera's monitored area, or more precisely, at the center of the camera's field of view. By deploying the microphone at the center of the camera's field of view, it can be ensured that the recorded audio data primarily originates from the monitored area.

[0026] Step S12: Extract features from each image group to obtain the image group features corresponding to each image group.

[0027] In this embodiment, feature extraction is performed on each image group to obtain the image group features corresponding to each image group. Feature extraction on each image group is performed to gradually refine the original and complex input data into more abstract, compact, and representative features; that is, feature extraction is performed on the image groups to extract effective feature representations—image group features—from each image group.

[0028] In one implementation, such as Figure 2 As shown, Figure 2 This is a schematic diagram of the framework of an embodiment of the crop monitoring model provided in this application. The crop monitoring model includes a first feature extraction module, which is used to extract features from each image group to obtain the image group features corresponding to each image group.

[0029] Step S13: Analyze the features of each image group to obtain the change features.

[0030] In this embodiment, the features of each image group are analyzed to obtain change features; these change features are used to characterize the changing trends of the image group features. In other words, by analyzing the temporal features of the image group, the change features characterizing the changing trends of crop conditions can be determined.

[0031] In one embodiment, feature extraction can be performed on the features of each image group to further extract abstract semantic features, and analysis can be performed based on each semantic feature to obtain change features.

[0032] In one specific implementation, such as Figure 2 As shown, the crop monitoring model also includes a second feature extraction module and a feature analysis and classification module. The second feature extraction module is used to extract features from each image group to further extract abstract semantic features. The feature analysis and classification module is used to analyze based on each semantic feature to obtain change features.

[0033] In one specific implementation, the backbone network of the second feature extraction module is a fully convolutional neural network (FCN). The selectable convolutional neural networks include, but are not limited to, ResNet, U-Net, PSPNet, DeepLab, Inception series networks, VGG series networks, and networks improved based on these networks.

[0034] Of course, in other implementations, the features of each image group can be directly analyzed to obtain the change features.

[0035] In one embodiment, the features of each image group are analyzed to obtain change features. Specifically, the features of each image group are vectorized to obtain a one-dimensional feature vector corresponding to each image group feature; the one-dimensional feature vector corresponding to each image group feature is analyzed to obtain change features. That is, the image group feature sequence is first one-dimensionalized, and then the one-dimensionalized image group feature sequence is analyzed to obtain the change features of the image group features, thereby revealing the change trend of the image group features.

[0036] In one specific implementation, such as Figure 3 As shown, Figure 3 This is a schematic diagram of the framework of an embodiment of the feature analysis and classification module provided in this application. The feature analysis and classification module includes a flatten layer and an LSTM layer. The flatten layer is used to vectorize the features of each image group to obtain a one-dimensional feature vector corresponding to the features of each image group. The LSTM layer is used to analyze the one-dimensional feature vector corresponding to the features of each image group to obtain the change features.

[0037] The LSTM layer includes, but is not limited to, Long Short-term Memory Networks (LSTM), ConvLSTM, bidirectional ConvLSTM, and other LSTM-like networks, etc., and is not limited here.

[0038] Step S14: Based on the change characteristics, monitor the crops and obtain the crop monitoring results.

[0039] In this embodiment, crops are monitored based on their changing characteristics to obtain monitoring results. These results include a first result, which characterizes the presence of pests and diseases on the crops. By acquiring time-series image sets, the changing characteristics representing the trend of crop state changes can be determined based on these image sets. This allows for monitoring of crop pests and diseases based on these changing characteristics, improving the accuracy of pest and disease monitoring and enabling earlier and more accurate automatic monitoring. Furthermore, compared to relying on manual field inspections and experience-based judgment for crop pest and disease monitoring, this method improves the efficiency of pest and disease monitoring.

[0040] In one embodiment, the first result indicates the presence of pests and diseases on the crop, and the crop monitoring results also include a second result, which is used to characterize the types of pests and diseases present on the crop.

[0041] In one specific implementation, such as Figure 2 As shown, the crop monitoring model includes a feature analysis and classification module, which is used to monitor crops based on their changing characteristics and obtain crop monitoring results.

[0042] In one specific implementation, such as Figure 3 As shown, the feature analysis and classification module also includes a classification layer, which is used to monitor crops based on changing features and obtain crop monitoring results.

[0043] The classification layer includes, but is not limited to, a combination of fully connected layers (FC) and softmax classification layers, fully convolutional neural networks (FCN), and support vector machines (SVM), and is not limited here.

[0044] Please see Figure 4 , Figure 4 yes Figure 1 The flowchart shown is a schematic diagram of one embodiment of step S12. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily follow the same pattern. Figure 4 The illustrated process sequence is limited. For example... Figure 4 As shown, the image group includes at least two modalities of images, and this embodiment includes: Step S41: For each image group, perform feature extraction on each modal image in the image group to obtain the first feature map corresponding to each modal image in the image group.

[0045] In this embodiment, for each image group, feature extraction is performed on each modal image in the image group to obtain a first feature map corresponding to each modal image in the image group. For each image group, feature extraction is performed on each modal image in the image group to gradually refine the original and complex input data into more abstract, more compact, and more representative features; that is, feature extraction is performed on each modal image in the image group to extract an effective feature representation—the first feature map—from each modal image.

[0046] In one embodiment, feature extraction is performed on each modal image in the image group to obtain a first feature map corresponding to each modal image in the image group. Specifically, this involves: performing initial feature extraction on each modal image in the image group to obtain a second feature map corresponding to each modal image in the image group; and performing dimensionality upscaling on the second feature maps corresponding to each modal image in the image group to obtain the first feature map corresponding to each modal image in the image group. Performing initial feature extraction on each modal image in the image group captures the basic features of each modal image—the second feature map. Performing dimensionality upscaling on the second feature maps corresponding to each modal image in the image group further expands the dimension of the features, allowing each feature point to encode richer information. Therefore, performing dimensionality upscaling on the second feature maps corresponding to each modal image in the image group can be seen as a further abstraction and refinement of the basic features of each modal image, which helps in learning deeper-level features. Therefore, after performing initial feature extraction on each modal image in the image group to obtain the second feature map corresponding to each modal image in the image group, performing dimensionality upscaling on the second feature map corresponding to each modal image in the image group can be understood as performing hierarchical feature extraction on each modal image in the image group, gradually constructing higher-level or deeper-level feature representations from lower-level features, thereby improving the ability to understand and represent each modal image.

[0047] The dimensionality-upgrading method for the second feature map corresponding to each modality image in the image group can be, but is not limited to, transposed convolution, interpolation, pixel shuffle, unpooling, etc., and is not limited here.

[0048] In one specific implementation, such as Figure 2 and Figure 5 As shown, Figure 5 This is a schematic diagram of the framework of the first feature extraction module provided in this application. The crop monitoring method is executed using a crop monitoring model. The crop monitoring model includes a first feature extraction module, which includes an initial feature extraction module group and a convolution module group connected in sequence. The convolution module group includes at least one first convolution module connected in sequence. Figure 5The initial feature extraction is performed by the convolutional block B1 in the image group to extract the initial features of each modality image in the image group and obtain the second feature map corresponding to each modality image in the image group. The dimensionality increase of the second feature map corresponding to each modality image in the image group to obtain the first feature map corresponding to each modality image in the image group is performed by the convolutional block.

[0049] The number of first convolutional modules included in the convolutional module group can be 1, 2, 3, 5, etc., and is not limited here.

[0050] In one specific implementation, initial feature extraction is performed on each modal image in the image group to obtain a second feature map corresponding to each modal image in the image group. Specifically, feature extraction is performed on each modal image in the image group to obtain a third feature map corresponding to each modal image in the image group; the third feature map corresponding to each modal image in the image group is downsampled to obtain a second feature map corresponding to each modal image in the image group. Downsampling the third feature map corresponding to each modal image in the image group can reduce the spatial size of the third feature map corresponding to each modal image, significantly reducing the number of parameters that the subsequent network needs to process and improving computational efficiency.

[0051] The downsampling method for downsampling the third feature map corresponding to each modality image in the image group can be, but is not limited to, max pooling, average pooling, stride convolution, bilinear interpolation, region aggregation, random downsampling, adaptive downsampling, etc., and is not limited here.

[0052] In one specific implementation, such as Figure 5 As shown, the initial feature extraction module group includes a second convolutional module ( Figure 5 The convolutional block A1 and the downsampling module (in the middle) Figure 5 The second convolution module is used to extract features from each modality image in the image group using max pooling to obtain the third feature map corresponding to each modality image in the image group. The third feature map corresponding to each modality image in the image group is downsampled to obtain the second feature map corresponding to each modality image in the image group using the downsampling module. That is, each modality image in the image group passes through the second convolution module and the downsampling module in sequence to obtain the corresponding second feature map.

[0053] It should be noted that when using max pooling to downsample the third feature map corresponding to each modality image in the image group, the downsampling module is the max pooling module.

[0054] In other specific implementations, feature extraction can be performed on each modal image in the image group to obtain the feature map corresponding to each modal image in the image group, and this feature map can be directly used as the second feature map corresponding to each modal image in the image group.

[0055] In one specific embodiment, the crop monitoring method is executed using a crop monitoring model. The crop monitoring model includes an initial feature extraction module group and a convolution module group connected in sequence. The convolution module group includes at least one first convolution module connected in sequence. The initial feature extraction is performed on each modality image in the image group to obtain the second feature map corresponding to each modality image in the image group. The dimensionality upscaling is performed on the second feature maps corresponding to each modality image in the image group to obtain the first feature map corresponding to each modality image in the image group. The initial feature extraction module group includes a second convolution module and a downsampling module. Each modality image in the image group passes through the second convolution module and the downsampling module in sequence to obtain the corresponding second feature map. The convolution module includes at least one set of convolutional layers, batch normalization layers, and activation function layers.

[0056] Each convolutional module typically consists of convolutional layers, batch normalization layers, and activation function layers in sequence. The convolutional layer uses a learnable filter or kernel to slide across the input data, performing "multiply-accumulate" operations; each filter is responsible for extracting a specific feature. The batch normalization layer standardizes the features of each layer in a batch of training data (making their mean 0 and variance 1), and then scales and translates them. This makes the input data distribution of each layer more stable, allowing for a larger learning rate and significantly speeding up training. The activation function layer introduces non-linearity, determining which features are activated.

[0057] There is no limit to the number of sets of convolutional layers, batch normalization layers, and activation function layers included in the convolution module; these can be set according to actual usage needs. For example, the convolution module may include a set of convolutional layers, batch normalization layers, and activation function layers.

[0058] In other implementations, initial feature extraction may be performed only on each modal image in the image group, and the extracted feature maps may be directly used as the first feature maps corresponding to each modal image in the image group.

[0059] In one specific implementation, such as Figure 5 As shown, the first feature extraction module includes several feature extraction branches. Each feature extraction branch corresponds to a modality image in the image group and is used to extract features from the corresponding modality image to obtain the first feature map corresponding to that modality image. At this point, each feature extraction branch should include an initial feature extraction module group and a convolution module group connected in sequence. The convolution module group includes at least one first convolution module connected in sequence. The initial feature extraction module group is used to perform initial feature extraction on each modality image in the image group to obtain the second feature map corresponding to each modality image in the image group. The convolution module group is used to perform dimensionality upscaling on the second feature maps corresponding to each modality image in the image group to obtain the first feature map corresponding to each modality image in the image group.

[0060] Step S42: Fuse the first feature maps to obtain the target fused feature map.

[0061] In this embodiment, the first feature maps are fused to obtain a target fused feature map. Fusing the first feature maps corresponding to each modality image comprehensively utilizes the image information from different modalities, enhancing feature representation capabilities; therefore, the target fused feature map obtained by fusing the first feature maps corresponding to each modality image is a richer and more discriminative feature representation.

[0062] In one embodiment, the first feature maps are fused to obtain a target fused feature map. Specifically, the first feature maps are concatenated to obtain a concatenated feature map; at least one statistical operation is performed on the feature values ​​in the concatenated feature map to obtain an attention feature map; and based on the attention feature map, the concatenated feature map is weighted to obtain the target fused feature map. An attention mechanism is introduced to intelligently and selectively fuse the data, balancing the importance of different modalities. By weighting the concatenated feature map using the attention feature map, important feature channels are amplified, while unimportant or irrelevant feature channels are suppressed.

[0063] It should be noted that concatenating the first feature maps simply means stacking them along the channel dimension (C) to generate a larger feature map. For example, suppose there are four first feature maps with dimensions [H (height), W (width), C (number of channels)]. Concatenating these four first feature maps generates a larger feature map—the concatenated feature map—with dimensions [H, W, 4C].

[0064] In one specific implementation, at least one statistic includes at least one of the following: global average pooling and global max pooling.

[0065] In one specific implementation, at least one statistical method includes global average pooling and global max pooling. The attention feature map is obtained by performing at least one statistical method on the feature values ​​in the concatenated feature map, specifically: performing global average pooling on the feature values ​​in the concatenated feature map to obtain a first global statistical feature value; and performing global max pooling on the feature values ​​in the concatenated feature map to obtain a second global statistical feature value; generating a first attention weight and a second attention weight based on the first global statistical feature value and the second global statistical feature value, respectively; and fusing the first attention weight and the second attention weight to obtain the attention feature map.

[0066] Global average pooling is applied to the stitched feature maps to obtain the first global statistical feature value, which represents the global average response intensity of each channel. Global max pooling is then applied to the stitched feature maps to obtain the second global statistical feature value, which represents the response intensity of the most salient feature in each channel. In other words, two different pooling methods are used to compress spatial information from both the average and max dimensions, generating two different channel-level statistical descriptors, providing an information basis for subsequent weight calculations.

[0067] In one specific implementation, such as Figure 5 and Figure 6 As shown, Figure 6 This is a schematic diagram of the framework of an embodiment of the feature attention mechanism module provided in this application. The first feature extraction module further includes a feature attention mechanism module, which includes a max pooling module and an average pooling module. The max pooling module is used to perform global max pooling on the feature values ​​in the spliced ​​feature map to obtain a second global statistical feature value. The average pooling module is used to perform global average pooling on the feature values ​​in the spliced ​​feature map to obtain a first global statistical feature value.

[0068] The feature attention mechanism module also includes a multilayer perceptron module, which generates a first attention weight and a second attention weight based on the first global statistical feature value and the second global statistical feature value, respectively, and fuses the first attention weight and the second attention weight to obtain an attention feature map.

[0069] Specifically, the first and second global statistical feature values ​​are input into a shared multilayer perceptron (MLP) module to obtain two attention weight vectors. The two attention weight vectors are then added element by element and normalized by a sigmoid function to fuse attention information from the average and maximum statistical strategies, generating a more comprehensive and robust final channel weight vector.

[0070] The multilayer perceptron module can include two fully connected layers. The first fully connected layer is used to introduce nonlinearity and reduce the number of parameters, and the second fully connected layer is used to limit the output weight values ​​to the range of 0-1.

[0071] In one specific implementation, such as Figure 5 As shown, the feature extraction module also includes a splicing module and a feature attention mechanism module. The splicing module is used to splice the first feature maps to obtain a spliced ​​feature map. The feature attention mechanism module is used to perform at least one statistical operation on the feature values ​​in the spliced ​​feature map to obtain an attention feature map, and based on the attention feature map, to perform weighted processing on the spliced ​​feature map to obtain a target fusion feature map.

[0072] Of course, in other implementations, the stitched feature map obtained by stitching together the first feature maps can also be used as the target fusion feature map. This is a simple and direct feature fusion method, which physically merges the image information of all modal images together and hands it over to the subsequent network to process the relationships between them.

[0073] Step S43: Based on the target fusion feature map, obtain the image group features corresponding to the image group.

[0074] In this embodiment, the image group features corresponding to the image group are obtained based on the target fusion feature map.

[0075] In one embodiment, based on the target fusion feature map, the image group features corresponding to the image group are obtained. Specifically, the target fusion feature map is deconvolved to obtain a monitoring feature map, which serves as the image group features corresponding to the image group. Since the target fusion feature map corresponding to the image group is obtained by extracting features from each modality image in the image group and fusing them based on the first feature maps corresponding to each modality image, the size of the target fusion feature map corresponding to the image group is already much smaller than the original input image. Therefore, deconvolution processing is performed on the target fusion feature map to enlarge the coarse, low-resolution target fusion feature map to a size similar to the original input image or matching the final task requirements, thereby restoring more spatial detail information.

[0076] In one specific implementation, such as Figure 5 As shown, the first feature extraction module also includes a deconvolution block, which is used to perform deconvolution processing on the target fusion feature map to obtain the monitoring feature map, which serves as the image group feature corresponding to the image group.

[0077] Of course, in other implementations, the target fusion feature map can also be directly used as the image group feature corresponding to the image group.

[0078] Please see Figure 7 , Figure 7This is a schematic diagram of an embodiment of the crop monitoring device provided in this application. The crop monitoring device 70 includes an acquisition module 71, a feature extraction module 72, an analysis module 73, and a monitoring module 74. The acquisition module 71 is used to acquire at least two image groups collected from crops. Each image group includes at least one modal image; different modal images within the same image group are collected from crops at the same time, while the same modal images within different image groups are collected from crops at different times. The feature extraction module 72 is used to extract features from each image group to obtain image group features corresponding to each image group. The analysis module 73 is used to analyze the features of each image group to obtain change features. These change features characterize the changing trend of the image group features. The monitoring module 74 is used to monitor crops based on the change features to obtain crop monitoring results. The crop monitoring results include a first result, which characterizes whether pests or diseases exist on the crops.

[0079] The image group includes at least two modal images; the feature extraction module 72 is used to extract features from each image group to obtain the image group features corresponding to each image group, including: for each image group, extracting features from each modal image in the image group to obtain a first feature map corresponding to each modal image in the image group; fusing each first feature map to obtain a target fused feature map; and obtaining the image group features corresponding to the image group based on the target fused feature map.

[0080] The feature extraction module 72 is used to extract features from each modal image in the image group to obtain a first feature map corresponding to each modal image in the image group. This includes: performing initial feature extraction on each modal image in the image group to obtain a second feature map corresponding to each modal image in the image group; and performing dimensionality upscaling on the second feature map corresponding to each modal image in the image group to obtain a first feature map corresponding to each modal image in the image group.

[0081] The feature extraction module 72 is used to perform initial feature extraction on each modal image in the image group to obtain a second feature map corresponding to each modal image in the image group, including: performing feature extraction on each modal image in the image group to obtain a third feature map corresponding to each modal image in the image group; downsampling the third feature map corresponding to each modal image in the image group to obtain a second feature map corresponding to each modal image in the image group; and / or, the above crop monitoring method is executed using a crop monitoring model, which includes a sequentially connected initial feature extraction module group and a convolution module group, the convolution module group including sequentially connected... The process of extracting initial features from each modality image in the image group, resulting in the second feature map corresponding to each modality image, is performed by the initial feature extraction module group. Then, the process of upscaling the second feature maps corresponding to each modality image in the image group to obtain the first feature map corresponding to each modality image is performed by the convolution module group. The initial feature extraction module group includes a second convolution module and a downsampling module. Each modality image in the image group sequentially passes through the second convolution module and the downsampling module to obtain the corresponding second feature map. The convolution module includes at least one set of convolutional layers, batch normalization layers, and activation function layers.

[0082] The feature extraction module 72 is used to fuse the first feature maps to obtain a target fused feature map, including: splicing the first feature maps to obtain a spliced ​​feature map; performing at least one statistical operation on the feature values ​​in the spliced ​​feature map to obtain an attention feature map; and performing weighted processing on the spliced ​​feature map based on the attention feature map to obtain the target fused feature map.

[0083] Among them, at least one of the above statistics includes at least one of the following: global average pooling and global max pooling.

[0084] Wherein, at least one of the above statistics includes global average pooling and global max pooling; the feature extraction module 72 is used to perform at least one statistical operation on the feature values ​​in the concatenated feature map to obtain an attention feature map, including: performing global average pooling on the feature values ​​in the concatenated feature map to obtain a first global statistical feature value; and performing global max pooling on the feature values ​​in the concatenated feature map to obtain a second global statistical feature value; generating a first attention weight and a second attention weight based on the first global statistical feature value and the second global statistical feature value respectively; and fusing the first attention weight and the second attention weight to obtain an attention feature map.

[0085] The feature extraction module 72 is used to obtain the image group features corresponding to the image group based on the target fusion feature map, including: performing deconvolution processing on the target fusion feature map to obtain the monitoring feature map, which is used as the image group features corresponding to the image group.

[0086] The analysis module 73 is used to analyze the features of each image group to obtain the change features, including: vectorizing the features of each image group to obtain the one-dimensional feature vector corresponding to the features of each image group; and analyzing the one-dimensional feature vector corresponding to the features of each image group to obtain the change features.

[0087] The aforementioned at least one modal image includes at least one of the following: visible light image, infrared image, humidity distribution thermal image, and frame spectrogram image, wherein the humidity distribution thermal image is obtained based on statistics of humidity data collected from crops, and the frame spectrogram image is obtained based on statistics of audio data collected from crops; and / or, the aforementioned first result characterizes the presence of pests and diseases on crops, and the crop monitoring result also includes a second result, which is used to characterize the types of pests and diseases present on crops.

[0088] Please see Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the electronic device provided in this application. The electronic device 80 includes a memory 81 and a processor 82 coupled to each other. The processor 82 is used to execute program instructions stored in the memory 81 to implement the steps of any of the above-described crop monitoring method embodiments. In a specific implementation scenario, the electronic device 80 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 80 may also include mobile devices such as laptops and tablets, which are not limited here.

[0089] Specifically, processor 82 controls itself and memory 81 to implement the steps of any of the above-described crop monitoring method embodiments. Processor 82 can also be referred to as a CPU (Central Processing Unit). Processor 82 may be an integrated circuit chip with signal processing capabilities. Processor 82 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 82 can be implemented using integrated circuit chips.

[0090] Please see Figure 9 , Figure 9This is a schematic diagram of the structure of an embodiment of the computer-readable storage medium provided in this application. The computer-readable storage medium 90 of this application embodiment stores program instructions 91. When executed, these program instructions 91 implement the methods provided by any embodiment of the crop monitoring method of this application and any non-conflicting combination thereof. The program instructions 91 can be formed into a program file and stored in the aforementioned computer-readable storage medium 90 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) can execute all or part of the steps of the methods of various embodiments of this application. The aforementioned computer-readable storage medium 90 includes various media capable of storing program code, such as a USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or terminal devices such as computers, servers, mobile phones, and tablets.

[0091] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.

[0092] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for monitoring crops, characterized in that, The method includes: Acquire at least two sets of images of crops; wherein the set of images includes at least one modal image, different modal images in the same set of images are acquired at the same time, and the same modal images in different sets of images are acquired at different times; Feature extraction is performed on each of the image groups to obtain the image group features corresponding to each image group; The features of each image group are analyzed to obtain variation features; wherein, the variation features are used to characterize the variation trend of the image group features; Based on the aforementioned change characteristics, the crops are monitored to obtain crop monitoring results; wherein, the crop monitoring results include a first result, which is used to characterize whether pests or diseases exist on the crops.

2. The method according to claim 1, characterized in that, The image group includes at least two modalities; the step of extracting features from each image group to obtain the image group features corresponding to each image group includes: For each of the image groups, feature extraction is performed on each modal image in the image group to obtain the first feature map corresponding to each modal image in the image group; The first feature maps are fused to obtain the target fused feature map; Based on the target fusion feature map, the image group features corresponding to the image group are obtained.

3. The method according to claim 2, characterized in that, The step of extracting features from each modal image in the image group to obtain a first feature map corresponding to each modal image in the image group includes: Initial feature extraction is performed on each modal image in the image group to obtain the second feature map corresponding to each modal image in the image group; The second feature map corresponding to each modality image in the image group is subjected to dimensionality upscaling to obtain the first feature map corresponding to each modality image in the image group.

4. The method according to claim 3, characterized in that, The step of performing initial feature extraction on each modality image in the image group to obtain a second feature map corresponding to each modality image in the image group includes: Feature extraction is performed on each modal image in the image group to obtain the third feature map corresponding to each modal image in the image group; The third feature map corresponding to each modal image in the image group is downsampled to obtain the second feature map corresponding to each modal image in the image group; And / or, the crop monitoring method is executed using a crop monitoring model, which includes a sequentially connected initial feature extraction module group and a convolution module group. The convolution module group includes at least one sequentially connected first convolution module. The initial feature extraction of each modality image in the image group to obtain the second feature map corresponding to each modality image in the image group is performed using the initial feature extraction module group. The upscaling of the second feature maps corresponding to each modality image in the image group to obtain the first feature map corresponding to each modality image in the image group is performed using the convolution module group. The initial feature extraction module group includes a second convolution module and a downsampling module. Each modality image in the image group passes through the second convolution module and the downsampling module in sequence to obtain the corresponding second feature map. The convolution module includes at least one set of convolutional layers, batch normalization layers, and activation function layers.

5. The method according to claim 2, characterized in that, The step of fusing each of the first feature maps to obtain the target fused feature map includes: The first feature maps are concatenated to obtain a concatenated feature map. Perform at least one statistical analysis on the feature values ​​in the spliced ​​feature map to obtain an attention feature map; Based on the attention feature map, the spliced ​​feature map is weighted to obtain the target fusion feature map.

6. The method according to claim 5, characterized in that, The at least one statistic includes at least one of the following: global average pooling and global max pooling.

7. The method according to claim 6, characterized in that, The at least one statistical method includes global average pooling and global max pooling; the step of performing at least one statistical method on the feature values ​​in the concatenated feature map to obtain the attention feature map includes: Global average pooling is performed on the feature values ​​in the spliced ​​feature map to obtain the first global statistical feature value; Furthermore, global max pooling is performed on the feature values ​​in the spliced ​​feature map to obtain the second global statistical feature value; First attention weights and second attention weights are generated based on the first global statistical feature value and the second global statistical feature value, respectively. The attention feature map is obtained by fusing the first attention weight and the second attention weight.

8. The method according to claim 2, characterized in that, The step of obtaining the image group features corresponding to the image group based on the target fusion feature map includes: The target fusion feature map is deconvolved to obtain the monitoring feature map, which serves as the image group feature corresponding to the image group.

9. The method according to claim 1, characterized in that, The analysis of the features of each of the image groups to obtain the change features includes: The features of each image group are vectorized to obtain a one-dimensional feature vector corresponding to each image group feature. The change features are obtained by analyzing the one-dimensional feature vectors corresponding to the features of each image group.

10. The method according to claim 1, characterized in that, The at least one modal image includes at least one of the following: visible light image, infrared image, humidity distribution thermal image, and frame spectrogram image, wherein the humidity distribution thermal image is obtained based on statistics of humidity data collected from the crop, and the frame spectrogram image is obtained based on statistics of audio data collected from the crop; And / or, the first result indicates the presence of pests and diseases on the crop, and the crop monitoring result further includes a second result, which is used to characterize the type of pests and diseases present on the crop.

11. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory being used to store program instructions, and the processor being used to execute the program instructions to implement the crop monitoring method as described in any one of claims 1-10.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program instructions that can be executed to implement the crop monitoring method as described in any one of claims 1-10.