Multi-modal fusion identification method and system for rice leaf type diseases and insect pests

Through multimodal fusion recognition methods, combined with feature extraction and adaptive gated feature fusion of image and non-imaging multispectral data, the problems of low efficiency and high computational complexity in rice disease and pest identification are solved, and efficient and accurate disease and pest identification is achieved, which is suitable for low-computing power devices.

CN120766129APending Publication Date: 2025-10-10HANGZHOU DIANZI UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510784874.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing rice pest and disease identification technology mainly relies on visual identification by plant protection experts, which is inefficient. In addition, the existing image spectral fusion recognition algorithm has high computational complexity and is difficult to effectively apply on low-computing power devices, and cannot meet the real-time and cost-sensitive needs of agricultural production.

Method used

A multimodal fusion recognition method is adopted to extract features from image and non-imaging multispectral data, utilize multi-scale convolutional attention modules and cross-attention feature interaction, combine adaptive gated feature fusion, dynamically adjust fusion weights, and construct a lightweight neural network model to achieve pest and disease identification.

Benefits of technology

It improves the accuracy and efficiency of identifying rice foliar pests and diseases, reduces computational complexity, is suitable for low-cost equipment, and meets the real-time identification needs of agricultural production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766129A_ABST
    Figure CN120766129A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal fusion identification method and system for rice leaf type diseases and insect pests. The method comprises the following steps: respectively carrying out feature extraction on image data and non-imaging multispectral data of a detected rice area to obtain image features and spectral features; performing cross attention feature extraction on the image features and the spectral features, and performing adaptive gating feature fusion on the obtained features to obtain modal fusion features; according to the gating feature fusion structure provided by the invention, the fusion weight is dynamically adjusted according to different input features so as to improve the flexibility of model information fusion, so that the fusion features containing different modal effective information are obtained, and the accuracy of rice leaf type disease and pest identification is improved. In addition, the spectral reflectivity data and the spectral vegetation index are subjected to feature extraction at the same time, feature encoders corresponding to the spectral reflectivity data and the spectral vegetation index are stacked for multiple times, and full extraction of spectral features is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of rice disease and insect pest detection, and particularly relates to a multimodal fusion identification method and system for rice foliar diseases and insect pests. Background Art

[0002] Rice is one of humanity's most important food crops, and rapid and accurate pest and disease detection plays a crucial role in understanding its occurrence and trends, as well as assisting in plant protection efforts. Currently, rice pest and disease identification relies primarily on visual statistical analysis by plant protection experts, which is relatively inefficient. While some researchers have developed intelligent rice detection devices that can improve identification efficiency, these devices typically capture only visible light images and then apply image processing algorithms for identification. The captured image bands are limited to visible light, recording only information that visually characterizes pest and disease symptoms on rice leaves and canopies. Data quality is significantly affected by sensor resolution, acquisition method, imaging illumination, as well as the shape, size, and background of the plant pests and diseases, leading to discrepancies in identification results in some scenarios. When pests and diseases infect plants, they affect the pigment content, activity, water content, and cell structure of their leaves, resulting in changes in their optical characteristics and a more sensitive approach. Spectroscopic techniques that monitor changes in vegetation reflectance across multiple optical bands (primarily from the visible to the near-infrared) can more sensitively capture changes in plant status caused by pests and diseases. Therefore, by effectively combining image and spectrum data, which have different detection ranges and mechanisms, more complete and comprehensive information can be obtained for pest and disease diagnosis, and building corresponding recognition algorithms can improve the identification effect of pests and diseases.

[0003] However, currently, there are few rice pest and disease identification algorithms that integrate images and spectra. Therefore, a new method is urgently needed that can effectively fuse spectral data reflecting internal rice information with image data reflecting surface information, while addressing data discrepancies and processing complexity. Furthermore, rice pest and disease identification requires high real-time performance. When promoting this technology in cost-sensitive agricultural production scenarios, the hardware performance of the equipment capable of running the algorithm must be considered. This requires the constructed algorithm model to have low computational complexity. However, the few existing image-spectral fusion recognition algorithm models mostly require imaging spectral data as input. Each pixel in the imaging spectrum records continuous spectral information, resulting in an extremely large amount of data. This leads to high computational complexity in constructing the algorithm model, and the high cost of equipment that can meet these requirements makes it difficult to promote its use. Therefore, an algorithm that is suitable for low-computing edge computing scenarios, can cost-effectively fuse image spectral information, and achieve high-precision rice pest and disease identification is urgently needed. This is crucial for providing reliable and efficient rice pest and disease identification technology solutions for plant protection personnel, producers, and others, and for their implementation in production. Summary of the Invention

[0004] The purpose of the present invention is to provide a multimodal fusion identification method and system for rice foliar pests and diseases.

[0005] In its first aspect, the present invention provides a multimodal fusion identification method for rice foliar pests and diseases, comprising: extracting features from image data and non-imaging multispectral data of a measured rice area to obtain image features and spectral features. The spectral feature extraction process includes extracting multiple spectral vegetation indices from the non-imaging multispectral reflectance data. The spectral reflectance data and spectral vegetation indices are then input into multiple stacked encoders to extract features. The resulting features are then concatenated and input into a fully connected layer for nonlinear calculation to obtain spectral features.

[0006] Cross-attention feature extraction is performed on image and spectral features, and the resulting features are then fused using adaptive gated features to obtain modality fusion features. During the adaptive gated feature fusion process, the fusion weights of different features are dynamically adjusted.

[0007] Identification of rice foliar pests and diseases using modal fusion features.

[0008] Preferably, the process of extracting image features includes: performing multi-layer convolution operations and multi-scale convolution attention operations on the image data to obtain image features. The multi-scale convolution attention operation extracts features in parallel under different receptive fields through multiple convolution kernels of different sizes, and the obtained different-scale features are cascaded and then the important feature channels and key areas are enhanced by the channel-space attention layer. Multiple multi-scale convolution attention operations are performed in the process of extracting image features. Each multi-scale convolution attention operation is added at a different depth in the image feature extraction process. The specific structure of the image feature extraction module is: multiple multi-scale convolution attention modules are evenly added at different depths of the neural network backbone. The neural network backbone includes multiple convolution blocks stacked in sequence.

[0009] Preferably, the feature extraction process of the encoder includes performing feature extraction based on a multi-head attention mechanism on the input features, and adding the obtained features to the input features to obtain attention features. The attention features are input into a multi-layer perceptron for multiple nonlinear and activation operations, and the obtained features are added to the attention features to obtain output features of the feature encoding.

[0010] Preferably, the non-imaging multispectral data includes non-imaging multispectral data of 530nm, 570nm, 680nm, 700nm, 740nm, 780nm bands or adjacent bands. The spectral vegetation index includes normalized difference vegetation index, photochemical reflectance index, simple ratio vegetation index, red-edge normalized difference vegetation index, normalized difference red-edge index and modified simple ratio index.

[0011] Preferably, the process of extracting the cross-attention features is as follows: inputting the spectral features and the image features into two interactive attention extraction branches respectively to obtain the spectral attention features and the image attention features containing cross-modal information.

[0012] The adaptive gating feature obtains the modality fusion feature The process is: connect the spectral attention features and image attention features, input the obtained features into the parallel feature weight calculation and feature nonlinear transformation, and then multiply them to obtain the modal fusion features.

[0013] The modality fusion feature The expression is:

[0014] ;

[0015] in, The feature is obtained by concatenating the spectral attention feature and the image attention feature; Representation layer normalization operation; Indicates nonlinear calculation; 、 Indicates activation.

[0016] Preferably, auxiliary classifiers are constructed based on spectral features, image features, spectral attention features, image attention features, and modal fusion features. A main loss is established for the main classifier for rice foliar pest and disease identification based on the input modal fusion features. Auxiliary losses are established for each auxiliary classifier. A weighted summation of the main loss and multiple auxiliary losses is used to generate a loss function for training the multi-stage fusion.

[0017] In a second aspect, the present invention provides a multimodal fusion identification system for rice foliar pests and diseases, comprising an image data acquisition module, a spectral data acquisition module, a spectral feature extraction module, an image feature extraction module, a feature fusion structure, and a classification module. The image data acquisition module is configured to acquire image data from a rice area under investigation. The spectral feature extraction module is configured to acquire non-imaging spectral reflectance data from multiple bands of the rice area under investigation. Multiple spectral vegetation indices are calculated based on the non-imaging spectral reflectance data.

[0018] The spectral feature extraction module includes two feature extraction branches that respectively input non-imaging spectral reflectance data and spectral vegetation index. The feature extraction branches are composed of a plurality of encoder stacks connected in sequence.

[0019] The image feature extraction module includes a neural network backbone and multiple multi-scale convolutional attention modules. These modules are added to the neural network backbone at different depths. The modules include a sequentially connected multi-scale convolutional extraction layer and a channel-spatial attention layer. The multi-scale convolutional feature extraction layer uses convolution kernels of various sizes to extract features in parallel across different receptive fields.

[0020] The feature fusion structure includes a sequentially connected cross-attention feature interaction module and a gated feature fusion module. The gated feature fusion module includes an input cascade layer, a feature weight calculation layer, and a feature nonlinear transformation layer. The output features of the input cascade layer are respectively input to the feature weight calculation layer and the feature nonlinear transformation layer. The outputs of the feature weight calculation layer and the feature nonlinear transformation layer are connected through multiplication and output to the classification module.

[0021] Preferably, the feature extraction branch corresponding to non-imaging spectral reflectance data consists of a stack of four encoders. The feature extraction branch corresponding to spectral vegetation index consists of a stack of three encoders. The encoder includes a first normalization layer, a multi-head attention layer, a second normalization layer, and a nonlinear perceptron layer arranged in sequence, with a residual connection added between the output of the first normalization layer and the output of the nonlinear perceptron layer for addition.

[0022] Preferably, the cross-attention feature interaction module includes two attention extraction branches. Each attention extraction branch includes a first normalization layer, a multi-head cross-attention layer, a second normalization layer, and a nonlinear perceptron layer connected in sequence. A residual connection is added between the output of the first normalization layer and the output of the nonlinear perceptron layer for addition. The multi-head cross-attention layers in the two attention extraction branches perform feature interaction, specifically by exchanging the query vector Q in the attention mechanism.

[0023] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the memory stores the computer program; and the processor executes the aforementioned multimodal fusion recognition method.

[0024] In a fourth aspect, the present invention provides a readable storage medium storing a computer program; when the computer program is executed by a processor, it is used to implement the aforementioned multimodal fusion recognition method.

[0025] The present invention has the following beneficial effects.

[0026] 1. The gated feature fusion structure provided by the present invention dynamically adjusts the fusion weight according to the different input features to improve the flexibility of the model when fusing information, thereby obtaining fusion features containing effective information of different modalities and improving the accuracy of rice foliar pest and disease identification.

[0027] 2. The present invention extracts six spectral vegetation indices based on spectral reflectance data and simultaneously performs feature extraction on both the spectral reflectance data and the spectral vegetation index, achieving full extraction of spectral features. Furthermore, feature encoders are used to calculate features for the spectral data and the spectral vegetation index, respectively, and the number of times the feature encoders are stacked is optimized, improving the spectral feature extraction effect. Optimal spectral feature extraction is achieved by calculating the spectral data using a four-stacked feature encoder and the spectral vegetation index using a three-stacked feature encoder.

[0028] 3. The present invention serially adds multi-scale convolutional attention modules at multiple different depths of the neural network backbone for extracting image features, thereby achieving full extraction of image features.

[0029] 4. The loss function designed in the present invention integrates the main loss corresponding to the main classifier and the auxiliary loss of the auxiliary classifier established based on spectral features, image features, spectral attention features, image attention features and modal fusion features, so that the different classification information provided by the features at different stages complement each other, helping the model to better identify the occurrence of pests and diseases in different samples during the training process. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a flow chart of Example 1 of the present invention.

[0031] Figure 2 This is a system block diagram of the multimodal fusion recognition model adopted in Example 1 of the present invention.

[0032] Figure 3 This is the feature visualization and feature-label mutual information diagram of Example 1 of the present invention.

[0033] Figure 4 This is a diagram showing the operation detection results of Example 1 of the present invention. DETAILED DESCRIPTION

[0034] The present invention will be further explained below with reference to the accompanying drawings;

[0035] like Figure 1 As shown, a multimodal fusion identification method for rice foliar pests and diseases includes the following steps:

[0036] Step 1: Acquisition and preprocessing of rice foliar disease atlas data

[0037] Taking rice bacterial blight, rice stem borer, and rice leaf roller as examples, atlas data of three categories, namely healthy / mildly diseased / severely diseased, were collected in parts of the rice-growing areas in the lower reaches of the Yangtze River, including Yuyao City, Ningbo, Zhejiang, Ninghai County, Ningbo, Zhejiang, Fenghua District, Ningbo, Zhejiang, Cixi City, Ningbo, Zhejiang, and Jiangdu District, Yangzhou, Jiangsu. The image data are visible light images such as RGB images or full-color images, and the spectral data are non-imaging multispectral data in the 530nm, 570nm, 680nm, 700nm, 740nm, 780nm bands or adjacent bands, and the collection field of view of the two types of data is consistent.

[0038] Preprocessing operations include:

[0039] The image was resampled to 224×224 pixels. The processed visible light image was then matched with the corresponding six-band non-imaging spectral reflectance values ​​to form a set of multimodal atlas data, which served as the algorithm input. After processing, all sets of multimodal atlas data were constructed into a single atlas dataset, containing 5,152 sets in total. This dataset was divided into training, validation, and test sets in a 6:2:2 ratio.

[0040] Step 2: Model spectral feature extraction

[0041] The spectral data input requirement of the model is non-imaging multispectral data that can be obtained at low cost, such as Figure 2 As shown in the spectral feature extraction module in Figure 1, after inputting the non-imaging spectral reflectance of six bands, the module calculates six spectral vegetation indices based on the reflectance values: the Normalized Difference Vegetation Index (NDVI), the Photochemical Reflectance Index (PRI), the Simple Ratio Vegetation Index (SRVI), the Normalized Difference Red-Edge Vegetation Index (NDRI), the Normalized Difference Red-Edge Index (NDRI), and the Modified Simple Ratio Index (MRI). The designed Trans_S4V3 spectral feature extraction architecture then performs feature calculation and extraction on the reflectance and vegetation indices.

[0042] The feature calculation and extraction operation process of the Trans_S4V3 spectral feature extraction structure is as follows: multiple spectral vegetation indices are extracted from non-imaging multispectral reflectance data. The spectral reflectance data and spectral vegetation index are respectively input into feature encoders of different stacking depths to extract features. The features obtained by feature encoding are spliced ​​and input into the fully connected layer for nonlinear calculation to obtain spectral features. The feature encoder includes multiple feature encoding operations connected in sequence. The stacking depth of the feature encoder is the number of feature encoding operations. In this embodiment, the stacking depth of the feature encoder corresponding to the non-imaging multispectral reflectance data is 4; the stacking depth of the feature encoder corresponding to the spectral vegetation index is 3; this combination of stacking depths helps to improve the effect of spectral feature extraction.

[0043] The sequential feature encoding process involves extracting input features using a multi-head attention mechanism, adding the extracted features to the input features to generate attention features. These attention features are fed into a multi-layer perceptron for multiple nonlinear and activation operations, and the extracted features are added to the attention features to generate the output features of the feature encoding.

[0044] The spectral feature extraction module obtains spectral features The expressions of are shown in formulas (1) to (3):

[0045]

[0046]

[0047]

[0048] in, is the result of spectral feature extraction; Respectively represent the input spectral original band reflectance data and spectral vegetation index data; Represents the nonlinear calculation of the fully connected layer; is the connection function; represents the calculation process of multiple stacked feature encoding modules, i represents the stacking depth, and the values ​​in formula (1) are 3 and 4 respectively; express Output; It is a multi-head attention mechanism; is layer normalization; Represents multiple nonlinear fully connected layers and activation calculations.

[0049] After spectral feature extraction and calculation, spectral data becomes spectral features. express.

[0050] Step 3: Model image feature extraction

[0051] The image data input requirement of the model is a visible light image corresponding to the spectral acquisition field of view, such as Figure 2As shown in the lightweight image feature extraction module (MCBAM_MB3), this module first designs a lightweight neural network backbone and adds a multi-scale convolutional attention module (MCBAM) to the lightweight network backbone at different depths (layers 3, 6, 10, and 13). The multi-scale convolutional attention module internally comprises a multi-scale convolutional extraction layer and a channel-spatial attention layer. The multi-scale convolutional feature extraction layer uses three different convolutional kernel sizes (1×1, 3×3, and 5×5) to extract features in parallel across different receptive fields, enabling comprehensive acquisition of local details (texture, edges), medium-scale structure (object structure), and global semantics (object contours). The extracted features at different scales are cascaded and then enhanced by the channel-spatial attention layer to highlight important feature channels and key regions.

[0052] After the image data is input into the lightweight image feature extraction module (MCBAM_MB3), the image features are obtained by performing feature calculation through different convolution blocks and multi-scale convolution attention modules (MCBAM) stacked in sequence. .

[0053] The expression of a multi-scale convolutional attention module is as follows:

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061] in, Represents the calculation results of a single multi-scale convolutional attention module; The calculation process of the new attention module that combines the channel attention module and the spatial attention module; They represent the feature maps obtained by convolution operations with convolution kernels of different sizes; A feature map representing the input; Represents the convolution operation; represents channel-by-channel multiplication; 、 Respectively represent the calculation results of the channel attention mechanism and the spatial domain attention mechanism; and 、 and The maximum value or average calculation is performed along the channel dimension; represents the compression-excitation operation; 、 Represents different activation operations; Represents the batch normalization operation;

[0062] The image features The expression is as follows:

[0063]

[0064] in, The representative image data is calculated by the stacked convolution blocks and multi-scale convolution attention modules in the lightweight image feature extraction module, and then input into the feature map of the final fully connected layer; Represents the weight matrix of the fully connected layer; Represents the bias parameter.

[0065] Step 4: Graph multimodal feature fusion and classification

[0066] The design of the feature fusion structure is mainly divided into a cross-attention feature interaction module and a gated feature fusion module. The spectrum and image data are extracted by different feature extraction structures to obtain spectral features. and image features Then, the cross-modal information correlation features of each modality feature to the other modality are calculated through the cross-multi-head attention feature interaction module, namely, the image attention feature and the spectral attention feature, respectively. 、 It indicates that the adaptive gating feature fusion module is then used to further fuse the different modal features that establish cross-modal information association. The gating mechanism dynamically adjusts the fusion weight according to the different input features to improve the flexibility of the model when fusing information, and finally obtains the fusion features containing effective information of different modalities for classification .

[0067] The feature fusion structure obtains the fusion feature The expression is as follows:

[0068]

[0069]

[0070]

[0071]

[0072] in, The new attention features are obtained after the input features are calculated through the cross attention head; 、 They are the image fusion features and spectral fusion features obtained by cross attention and gating calculation of image and spectral features respectively; represents activation operation; represents activation operation; Representation layer normalization operation; Indicates nonlinear calculations.

[0073] After the image features and spectral features are fused, the final modal fusion features are input into the main classifier to obtain the pest classification information of different data samples. The main classifier uses a serial multi-layer fully connected classification layer. Step 5: Algorithm training and saving

[0074] During the training process, in order to further utilize the feature information of different stages of the model to guide the model to achieve better recognition results, the features of different stages of the model are used to construct a classifier and calculate the corresponding loss, and a multi-stage fusion loss function is established. The multi-stage fusion loss function Including main loss And the auxiliary losses corresponding to each stage . The main loss is calculated for the main classifier by using the label smoothing loss function. The auxiliary loss is calculated for the corresponding auxiliary classifier by using the label smoothing loss function. The auxiliary classifier is constructed based on features from five different stages. The features from the five different stages are image features, spectral features, image attention features, spectral attention features, and fusion features. The auxiliary classifier adopts a single-layer fully connected layer. Each auxiliary loss is adaptively added by calculating the attention weight.

[0075] The multi-stage fusion loss function The expression is as follows:

[0076]

[0077]

[0078] in, is a variable weight parameter, currently set to 0.5; is the label smoothing loss function, and its smoothing coefficient ranges from 0 to 1, with 0.1 being the preferred value; Represent the main classifier prediction label, auxiliary classifier prediction label and true label respectively; Represents the attention weights of different auxiliary losses, with an initial value of 0.2 and adaptively adjusted during training.

[0079] Finally, the multi-stage fusion loss function ( ) for training, and save the best algorithm model after training as the multimodal fusion recognition model.

[0080] The same-viewing angle image data and spectral data of the tested rice area are input into the multimodal fusion recognition model to obtain the identification results of rice foliar pests and diseases in the tested rice area.

[0081] Step 6: Algorithm Evaluation

[0082] Taking the test set data in the constructed data set as an example, the feature extraction effect of the algorithm is tested, such as Figure 3 As shown. Figure 3 From the visualization results of the feature dimensionality reduction, it can be seen that after the algorithm extracts spectral and image features from the original spectral and image data, although most samples can distinguish different diseases, there is still confusion. As the fusion of different stages proceeds, the degree of confusion of the samples gradually decreases, achieving a more obvious distinction. The feature-label mutual information contained in the spectral and image features is also improved in the fusion process, indicating that the new algorithm that fuses image and spectral multimodal information has significantly improved the effect of pest and disease identification compared to using only a single modality of image or spectrum. Finally, the algorithm was used to identify randomly selected mild bacterial blight sample data, and the results are as follows: Figure 4 As shown, the algorithm correctly identified the actual disease condition of the sample as mild bacterial blight, which is consistent with the judgment of the plant protection experts, indicating that the algorithm has a high degree of reliability.

[0083] This embodiment tests the algorithm based on rice bacterial blight, rice stem borer, and rice leaf roller disease atlas data collected from parts of the rice-growing areas in the lower reaches of the Yangtze River, including Yuyao City, Ningbo, Zhejiang, Ninghai County, Ningbo, Zhejiang, Fenghua District, Ningbo, Zhejiang, Cixi City, Ningbo, Zhejiang, and Jiangdu District, Yangzhou, Jiangsu. According to the algorithm construction steps of each process proposed by the algorithm, a fast and high-precision recognition algorithm model that can be used to detect mild and severe incidence and health status of rice bacterial blight, rice stem borer, and rice leaf roller diseases is obtained. The detection speed and accuracy basically meet the application requirements, indicating that the current algorithm scheme has good performance and application value.

[0084] The executor of this embodiment may be a computer device; the computer device includes a memory and a processor, the memory stores executable code, and when the processor executes the executable code, the method described in any one of the embodiments is implemented, specifically, the processor in the computer device executes processing data sets, constructing a multimodal fusion recognition model, adaptively extracting and fusing modal key features, training and testing the multimodal fusion recognition model, and using the multimodal fusion recognition model to process input data to identify the occurrence of rice diseases and insect pests in the region.

[0085] The memory may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk drive. The system network element communicates with at least one other network element via at least one communication interface (which may be wired or wireless), such as the Internet, a wide area network, a local area network, or a metropolitan area network. The bus may be an ISA bus, a PCI bus, or an EISA bus. The bus may be categorized as an address bus, a data bus, or a control bus.

[0086] The memory is used to store the program, and the processor executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any of the aforementioned embodiments of the present invention can be applied to the processor or implemented by the processor.

[0087] The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of the present invention may be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or the like. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.

[0088] In some other embodiments, the computer device may also be one or more desktop computers, laptop computers, workstations, databases, and servers.

[0089] Example 2

[0090] A multimodal fusion recognition system for rice foliar pests and diseases includes an image data acquisition module, a spectral data acquisition module, a spectral feature extraction module, an image feature extraction module, a feature fusion structure, and a classification module. These modules, along with the spectral feature extraction module, the image feature extraction module, the feature fusion structure, and the classification module, form a multimodal fusion recognition model.

[0091] The image data acquisition module is used to collect image data from the rice field being measured. The spectral feature extraction module is used to collect non-imaging spectral data in six wavebands from the rice field being measured. The six wavebands are 530 nm, 570 nm, 680 nm, 700 nm, 740 nm, and 780 nm, or adjacent wavebands. The image data acquisition module and the spectral feature extraction module maintain the same field of view during data acquisition.

[0092] The spectral feature extraction module includes two feature extraction branches, one for the original spectral bands and the other for the spectral vegetation index. Each branch consists of multiple encoder stacks connected in sequence. The encoders include a first normalization layer, a multi-head attention layer, a second normalization layer, and a nonlinear perceptron layer. A residual connection is added between the outputs of the first normalization layer and the nonlinear perceptron layer for summation. The feature extraction branch for the original spectral bands consists of four encoder stacks, while the feature extraction branch for the spectral vegetation index consists of three encoder stacks.

[0093] The image feature extraction module includes a lightweight neural network backbone and multiple multi-scale convolutional attention modules. Figure 2 As shown, the lightweight neural network backbone consists of a sequentially connected 2D convolutional layer, three 3×3 convolutional blocks, eight 5×5 convolutional blocks, a 2D convolutional layer, an average pooling layer, and a fully connected layer. The convolutional blocks consist of a sequentially connected point convolutional layer, an activation layer, an n×n separation convolutional layer (3×3 or 5×5), an activation layer, a SE excitation layer, and a point convolutional layer. Residual connections are added between the input and output of the convolutional blocks for addition.

[0094] Multiple multi-scale convolutional attention modules are serially arranged at different depths within the lightweight neural network backbone. Each multi-scale convolutional attention module contains a multi-scale convolutional extraction layer and a channel-spatial attention layer. The multi-scale convolutional feature extraction layer uses three different convolution kernel sizes (1×1, 3×3, and 5×5) to extract features in parallel across different receptive fields, enabling comprehensive acquisition of local details (texture, edges), medium-scale structure (object structure), and global semantics (object contours). The extracted features at different scales are concatenated, and the channel-spatial attention layer enhances important feature channels and key regions. Adding multiple multi-scale convolutional attention modules achieves an image feature extraction accuracy of 85.24%, exceeding the 84.47% accuracy of a conventional lightweight neural network backbone, achieving superior feature extraction results.

[0095] The feature fusion structure includes a cross-attention feature interaction module and a gated feature fusion module. The cross-attention feature interaction module includes two attention extraction branches. Each attention extraction branch includes a first-layer normalization layer, a multi-head cross-attention layer, a second-layer normalization layer, and a nonlinear perceptron layer connected in sequence. A residual connection is added between the output of the first-layer normalization layer and the output of the nonlinear perceptron layer for addition operation. The multi-head cross-attention layers in the two attention extraction branches perform feature interaction, specifically by exchanging the query vector Q in the attention mechanism.

[0096] The gated feature fusion module includes an input cascade layer, a feature weight calculation layer and a feature nonlinear transformation layer. The output features of the input cascade layer are respectively input to the feature weight calculation layer and the feature nonlinear transformation layer. The outputs of the feature weight calculation layer and the feature nonlinear transformation layer are connected and output through multiplication calculation. The feature weight calculation layer adaptively updates the weight matrix according to the input features. By multiplying the weight matrix with the output features of the feature nonlinear transformation layer, the optimal fusion feature can be obtained. After introducing the gated feature fusion module, the accuracy reached 95.24%, which is higher than the 94.17% accuracy without using this structure, and a better feature fusion effect was achieved. The above accuracies were all obtained by training using a normal label smoothing loss function (only including the main loss in Example 1).

[0097] In this embodiment, the multi-stage fusion loss function provided in Example 1 is used ( ) The training finally includes a multimodal fusion recognition model of the above structures. The different classification information provided by the features at different stages complement each other, helping the model to better identify the occurrence of pests and diseases in different samples during the training process, and achieved a recognition accuracy of 95.44%, which is higher than the model accuracy obtained when training without using this loss function.

[0098] Example 3

[0099] A multimodal fusion identification system for rice foliar pests and diseases. The difference between this embodiment and embodiment 2 is that the number of encoder stacks of the two feature extraction branches in the spectral feature extraction module is different.

[0100] In this embodiment, the feature extraction branch corresponding to the original spectral band is composed of four stacked encoders, and the feature extraction branch corresponding to the spectral vegetation index is composed of two stacked encoders.

[0101] Example 4

[0102] A multimodal fusion identification system for rice foliar pests and diseases. The difference between this embodiment and embodiment 2 is that the number of encoder stacks of the two feature extraction branches in the spectral feature extraction module is different.

[0103] In this embodiment, the feature extraction branch corresponding to the original spectrum band is composed of four stacked encoders, and the feature extraction branch corresponding to the spectral vegetation index is composed of one stacked encoder.

[0104] Example 5

[0105] A multimodal fusion identification system for rice foliar pests and diseases. The difference between this embodiment and embodiment 2 is that the number of encoder stacks of the two feature extraction branches in the spectral feature extraction module is different.

[0106] In this embodiment, the feature extraction branch corresponding to the original spectral band is composed of three stacked encoders, and the feature extraction branch corresponding to the spectral vegetation index is composed of four stacked encoders.

[0107] Example 6

[0108] A multimodal fusion identification system for rice foliar pests and diseases. The difference between this embodiment and embodiment 2 is that the number of encoder stacks of the two feature extraction branches in the spectral feature extraction module is different.

[0109] In this embodiment, the feature extraction branch corresponding to the original spectral band is composed of two stacked encoders, and the feature extraction branch corresponding to the spectral vegetation index is composed of four stacked encoders.

[0110] Example 7

[0111] A multimodal fusion identification system for rice foliar pests and diseases. The difference between this embodiment and embodiment 2 is that the number of encoder stacks of the two feature extraction branches in the spectral feature extraction module is different.

[0112] In this embodiment, the feature extraction branch corresponding to the original spectral band is composed of one encoder stack, and the feature extraction branch corresponding to the spectral vegetation index is composed of four encoder stacks.

[0113] Example 8

[0114] A multimodal fusion identification system for rice foliar pests and diseases. The difference between this embodiment and embodiment 2 is that the two feature extraction branches in the spectral feature extraction module have no stacking structure and only include one encoder.

[0115] Comparing the spectral feature extraction accuracy of the spectral feature extraction modules in Examples 2-6, the spectral feature extraction accuracy corresponding to Example 2 reaches 90.49%; the spectral feature extraction accuracy corresponding to Example 3 reaches 89.22%; the spectral feature extraction accuracy corresponding to Example 4 reaches 88.93%; the spectral feature extraction accuracy corresponding to Example 5 reaches 89.71%; the spectral feature extraction accuracy corresponding to Example 6 reaches 87.09%; the spectral feature extraction accuracy corresponding to Example 7 reaches 86.70%; and the spectral feature extraction accuracy corresponding to Example 8 reaches 87.96%.

[0116] It can be seen that the spectral four-times stacking and vegetation index three-times stacking structure in Example 2 can achieve higher spectral feature extraction accuracy compared with other feature extraction structures.

Claims

1. A multimodal fusion identification method for rice foliar pests and diseases, characterized in that: The method comprises: Feature extraction is performed on the image data and non-imaging multispectral data of the rice field under test to obtain image features and spectral features. The spectral feature extraction process includes: extracting multiple spectral vegetation indices from the non-imaging multispectral reflectance data; inputting the spectral reflectance data and spectral vegetation indices into multiple stacked encoders to extract features, and then concatenating the obtained features and inputting them into a fully connected layer for nonlinear calculation to obtain spectral features. Cross-attention feature extraction is performed on image features and spectral features, and the obtained features are subjected to adaptive gated feature fusion to obtain modal fusion features; the fusion weights of different features are dynamically adjusted during the adaptive gated feature fusion process; Identification of rice foliar pests and diseases using modal fusion features.

2. The multimodal fusion recognition method according to claim 1, characterized in that: The process of extracting image features includes: performing multi-layer convolution operations and multi-scale convolution attention operations on image data to obtain image features; the multi-scale convolution attention operation extracts features in parallel under different receptive fields through multiple convolution kernels of different sizes, and the obtained different-scale features are cascaded and then enhanced by the channel-spatial attention layer for important feature channels and key areas; multiple multi-scale convolution attention operations are performed in the process of extracting image features; each multi-scale convolution attention operation is added at different depths in the image feature extraction process.

3. The multimodal fusion recognition method according to claim 1, characterized in that: The feature extraction process of the encoder includes extracting features from the input features based on a multi-head attention mechanism, and adding the obtained features to the input features to obtain attention features; the attention features are input into a multi-layer perceptron for multiple nonlinear and activation operations, and the obtained features are added to the attention features to obtain the output features of the feature encoding.

4. The multimodal fusion recognition method according to claim 1, characterized in that: The non-imaging multispectral data includes non-imaging multispectral data of 530nm, 570nm, 680nm, 700nm, 740nm, 780nm bands or adjacent bands; the spectral vegetation index includes normalized vegetation index, photochemical reflectance index, simple ratio vegetation index, red edge normalized difference vegetation index, normalized difference red edge index and modified simple ratio index.

5. The multimodal fusion recognition method according to claim 1, characterized in that: The cross-attention feature extraction process is as follows: spectral features and image features are input into two interactive attention extraction branches respectively to obtain spectral attention features and image attention features containing cross-modal information; The adaptive gating feature obtains the modality fusion feature The expression is: ; in, The feature is obtained by concatenating the spectral attention feature and the image attention feature; Representation layer normalization operation; Indicates nonlinear calculation; 、 Indicates activation.

6. The multimodal fusion recognition method according to claim 5, characterized in that: Auxiliary classifiers are constructed based on spectral features, image features, spectral attention features, image attention features, and modal fusion features respectively; The main classifier for rice foliar pest and disease identification based on the input modal fusion features establishes the main loss; auxiliary losses are established based on each auxiliary classifier; the weighted sum of the main loss and multiple auxiliary losses is used to obtain the multi-stage fusion loss function for training.

7. A multimodal fusion recognition system for rice foliar pests and diseases, characterized by: Used to execute the multimodal fusion recognition method according to claim 1; the multimodal fusion recognition system includes an image data acquisition module, a spectral data acquisition module, a spectral feature extraction module, an image feature extraction module, a feature fusion structure and a classification module; the image data acquisition module is used to collect image data of the measured rice area; the spectral feature extraction module is used to collect non-imaging spectral reflectance data of multiple bands of the measured rice area; multiple spectral vegetation indices are calculated based on the non-imaging spectral reflectance data; The spectral feature extraction module includes two feature extraction branches that input non-imaging spectral reflectance data and spectral vegetation index respectively; The feature extraction branch consists of multiple encoder stacks connected in sequence; The image feature extraction module includes a neural network backbone and multiple multi-scale convolutional attention modules; the multiple multi-scale convolutional attention modules are respectively added to the neural network backbone at different depths; the multi-scale convolutional attention module includes a multi-scale convolutional extraction layer and a channel-spatial attention layer connected in sequence; the multi-scale convolutional feature extraction layer extracts features in parallel under different receptive fields using convolution kernels of multiple different sizes; The feature fusion structure includes a cross-attention feature interaction module and a gated feature fusion module connected in sequence; the gated feature fusion module includes an input cascade layer, a feature weight calculation layer and a feature nonlinear transformation layer; the output features of the input cascade layer are respectively input to the feature weight calculation layer and the feature nonlinear transformation layer; The outputs of the feature weight calculation layer and the feature nonlinear transformation layer are connected through multiplication calculation and output to the classification module.

8. The multimodal fusion recognition system according to claim 7, characterized in that: The feature extraction branch corresponding to non-imaging spectral reflectance data consists of four encoder stacks; the feature extraction branch corresponding to spectral vegetation index consists of three encoder stacks; the encoder includes a first layer normalization layer, a multi-head attention layer, a second layer normalization layer and a nonlinear perceptron layer arranged in sequence, and a residual connection is added between the output of the first layer normalization layer and the output of the nonlinear perceptron layer for addition operation.

9. The multimodal fusion recognition system according to claim 7, characterized in that: The cross-attention feature interaction module includes two attention extraction branches; each attention extraction branch includes a first layer normalization layer, a multi-head cross-attention layer, a second layer normalization layer and a nonlinear perceptron layer connected in sequence; a residual connection is added between the output of the first layer normalization layer and the output of the nonlinear perceptron layer for addition operation; the multi-head cross-attention layers in the two attention extraction branches perform feature interaction.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The memory stores a computer program; the processor executes the multimodal fusion recognition method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Hyperspectral image-based rice bakanae disease bacteria-carrying seed detection method and device

    CN121207889A

  • Fig health status assessment method and system based on multi-modal data fusion

    CN121661548A

  • Method and system for evaluating health status of figs based on multi-modal data fusion

    CN121661548B