A rice disease identification monitoring method based on a KBnet network model
By combining the KBnet network model with KF-KAN and BC-PINN modules, the problem of insufficient recognition accuracy of rice disease segmentation models on multi-scale and irregular lesion boundaries was solved, achieving high-precision and robust disease recognition, and improving the intelligent and efficient implementation of rice disease management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
- Filing Date
- 2025-08-26
- Publication Date
- 2026-05-01
AI Technical Summary
Existing rice disease segmentation models suffer from insufficient recognition accuracy and poor stability when dealing with multi-scale lesions and irregular lesion boundaries, especially when lesions have diverse shapes, sizes, and complex textures, making accurate segmentation and recognition difficult.
The KBnet network model is adopted, combined with the Kalman filter enhanced KAN module (KF-KAN) and the boundary constraint physical information neural network module (BC-PINN). By constructing datasets for multiple disease categories, Kalman filtering and physical prior knowledge are introduced to optimize the segmentation of disease areas.
It significantly improved the rice disease segmentation model's ability to identify complex lesion areas, enhanced segmentation accuracy and robustness, achieved an intersection-over-union (IoU) ratio of 72.3% and a Dice coefficient of 83.9%, and strengthened the model's ability to identify multi-scale lesions and irregular lesion boundaries.
Smart Images

Figure CN121053536B_ABST
Abstract
Description
A Rice Disease Identification and Monitoring Method Based on KBnet Network Model Technical Field
[0001] This invention relates to the field of crop disease identification and monitoring, and in particular to a method for rice disease identification and monitoring based on the KBnet network model. Background Technology
[0002] With the development of image processing technology, machine vision is increasingly being applied to the agricultural field. The rapid development of artificial intelligence technologies such as deep learning has significantly improved image processing capabilities, providing strong support for disease identification. Among these technologies, image segmentation, as a key component, plays a crucial role in the accurate identification of rice diseases. In the development of image segmentation technology, models such as U-Net, DeepLab, and Mask R-CNN have been widely used in sophisticated image analysis tasks. They can segment targets in images at the pixel level, effectively distinguishing target regions from background regions. Deep neural network architectures based on convolutional neural networks (CNNs) and encoder-decoder structures have demonstrated significant performance advantages in image segmentation tasks. These models, through multi-scale feature extraction and end-to-end training strategies, can capture complex spatial structural information in images, achieving more accurate region segmentation and semantic understanding.
[0003] Although current research has made significant progress in the field of segmentation, there are still shortcomings and two key challenges: (1) In practical applications, rice disease segmentation often faces problems such as multi-scale lesions. Due to the diversity of lesion shape and size, the same disease may show completely different characteristics at different growth stages, making it difficult to accurately extract the disease area; (2) Rice lesions often exhibit irregular diffusion characteristics during the growth process, usually with complex textures and shapes. This irregularity makes it difficult to accurately capture the lesion boundary, which can easily lead to confusion or misjudgment.
[0004] To address the problem of irregular lesion growth and complex textures, Borse et al. proposed the InverseForm loss function, which introduces an inverse transform network to learn the degree of parameterized spatial transformation between the predicted boundary and the true boundary. This module consists of a boundary distance metric submodule and a pixel-level cross-entropy loss. The former focuses on capturing boundary translation, rotation, and scale changes, compensating for the insufficient sensitivity of traditional losses to local spatial errors. This method effectively solves the problem of segmenting irregular lesions, significantly improving the accuracy of the segmentation model in boundary regions without increasing the model complexity in the inference stage. However, the InverseForm loss function relies on accurate boundary annotations, and its performance may degrade when faced with boundary annotation errors or weak boundaries. Bougourzi et al. proposed a multi-class boundary-aware cross-entropy loss function (MBA-CE). This method addresses the problem of irregular boundaries in infected areas in CT images by designing boundary-aware weights to enhance the discriminative ability of boundary pixels. During feature supervision, MBA-CE further strengthens the model's ability to recognize complex textures and fine-grained structures, significantly improving the segmentation accuracy of multi-cell pneumonia infection areas. While MBA-CE enhances boundary perception, its introduced boundary weight mechanism may induce misleading gradients in non-boundary regions, affecting the overall model's stability and generalization ability. Therefore, a novel method is needed to optimize bounding box prediction, thereby accurately locating lesion regions and improving lesion segmentation accuracy. Summary of the Invention
[0005] In view of the above-mentioned shortcomings, in order to improve the ability of rice disease segmentation models to identify complex lesion areas, this invention provides a rice disease identification and monitoring method based on the KBnet network model. By constructing datasets of multiple disease categories, using the KNnet network model and introducing Kalman filtering to enhance the KAN module and physical prior knowledge, the spatial location predicted by the model is constrained, providing accurate segmentation results.
[0006] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:
[0007] A method for identifying and monitoring rice diseases based on the KBnet network model, characterized in that the method includes:
[0008] Acquire images of rice leaves, process the rice leaf images to obtain first data that is adapted to the model input;
[0009] A KBNet network model structure capable of performing rice disease segmentation tasks was constructed and obtained. The KBNet network model structure includes a BC-PINN module and a KF-KAN module.
[0010] The first data is used as input to the KBNet network model structure for training to obtain a target model that meets the preset evaluation index.
[0011] A current rice leaf image is acquired, and the current rice leaf image is preprocessed to obtain vectorized data. The vectorized data is used as the input of the target model, and the target model infers and outputs to obtain the disease monitoring results of the current rice leaf image.
[0012] Furthermore, the KF-KAN module includes a KAN network structure and a Kalman filter module, with the output of the KAN network structure serving as the input of the Kalman filter module.
[0013] Furthermore, the Kalman filter module includes:
[0014] The output of the KAN network structure is used as a preliminary state estimate, which is then used as the state estimate for the Kalman filter. The Kalman filter makes predictions based on the state estimate and the state transition matrix, as follows: ;
[0015] The Kalman filter updates the current state based on the actual observations and calculates the Kalman gain K, expressed as: ;
[0016] ;
[0017] in, It is the predicted state covariance matrix. It is the observation matrix. It is the transpose of the observation matrix. It is the observation noise covariance matrix. State estimation; Here, F represents the state estimate at time k-1, and F is the state transition matrix. This is expressed as the state estimate predicted at time k;
[0018] This represents the state estimate after the update at time k;
[0019] Represented as the actual observed value at time k;
[0020] The Kalman filter updates the state covariance matrix as follows:
[0021] ,in, It is the identity matrix; Represented as the predicted state covariance matrix;
[0022] K(·) represents the Kalman gain function;
[0023] It is represented as the state covariance matrix after time k.
[0024] Furthermore, the BC-PINN module includes a PINN model and a multinomial loss function; the multinomial loss function is a loss set according to the task of the PINN model.
[0025] Furthermore, the polynomial loss function includes Dice loss, BCE loss, and PINN loss.
[0026] Furthermore, the preset evaluation metrics include Dice loss, IOU loss, accuracy, and recall.
[0027] Furthermore, the data processing includes dividing the rice leaf images into a training set, a validation set, and a test set according to a set ratio; adjusting the rice leaf images to the same size; and performing data augmentation on the rice leaf images.
[0028] The beneficial effects of this application are as follows:
[0029] To improve the ability of rice disease segmentation models to identify complex lesion regions, an image dataset was constructed, covering three common rice diseases: bacterial blight, rice blast, and tungro. All image samples were precisely annotated pixel-by-pixel using the Labelme tool to generate corresponding segmentation masks, ensuring the accuracy of lesion region information.
[0030] A Kalman Filter Enhanced KAN (KF-KAN) module is proposed. This module efficiently captures nonlinear features at different scales using KAN and dynamically updates and fuses multi-scale information by combining Kalman filtering, achieving accurate feature extraction of complex lesion regions. By adaptively adjusting scale and detail, KF-KAN effectively improves the model's accuracy and robustness in handling multi-scale lesions, thereby enhancing its overall performance in segmentation tasks.
[0031] A boundary-constrained physical information neural network module (BC-PINN) is proposed, which combines prior knowledge of physical laws with the powerful modeling capabilities of neural networks. By embedding physical information such as lesion growth patterns into the loss function, BC-PINN can effectively constrain the model's prediction of irregularly growing lesions. Furthermore, the BC-PINN module further constrains the spatial location of the prediction results by calculating the penalty between the predicted mask and the image boundary, thus providing more accurate segmentation results.
[0032] The proposed KBNet model achieves an Intersection over Union (IoU) ratio of 72.3% and a Dice coefficient of 83.9% on a self-constructed dataset. This model performs exceptionally well in rice disease segmentation tasks, providing strong technical support for the accurate identification and control of rice diseases and promoting the intelligent and efficient implementation of rice disease management.
[0033] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0035] Figure 1 is a schematic diagram of the process of a rice disease identification and monitoring method based on the KBnet network model provided by the present invention.
[0036] Figure 2 shows the labels and text descriptions of rice diseases provided in the embodiments of the present invention;
[0037] Figure 3 is a schematic diagram of the KBNet network model provided in an embodiment of the present invention;
[0038] Figure 4 is a schematic diagram of the rice disease identification and monitoring task based on the KBnet network model provided in the embodiment of the present invention. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0040] Dice loss is a metric used to measure the similarity between two sets and is widely used in image segmentation tasks.
[0041] IOU loss: Intersection over Union (IoU) is a metric for measuring the accuracy of detecting corresponding objects in a given dataset;
[0042] Accuracy: Accuracy refers to the proportion of samples that the model correctly predicts out of the total sample.
[0043] Recall: Recall rate refers to the proportion of samples that are actually positive that are correctly predicted as positive by the model.
[0044] The design concept of the embodiments of this application will be briefly introduced below.
[0045] As shown in Figure 1, a rice disease identification and monitoring method based on the KBnet network model includes:
[0046] Step S01: Acquire the collected rice leaf images and the text descriptions of the rice leaf images, perform data processing on the rice leaf images and the text descriptions to obtain the first data adapted to the model input;
[0047] In practice, the rice disease and pest segmentation dataset is the foundation of this study. Three typical rice leaf diseases—Bacterial Blight, Blast, and Tungro—were selected to verify the characteristic manifestations of rice leaf lesions under different disease conditions. Bacterial Blight typically forms water-soaked, yellowish-brown lesions on leaves with blurred edges, easily spreading along the veins. Blast often causes elliptical or spindle-shaped grayish-brown spots on leaves with dark brown edges and a grayish-white center, sometimes accompanied by small black dots. Tungro disease manifests as obvious symptoms such as chlorosis, yellowing, curling, and stunted growth. A total of 1550 rice leaf images were collected from the Kaggle platform. The image environments were complex, including strong light, walls, soil, cloth bags, and weeds, providing a more realistic challenge for disease segmentation. During the image selection process, strict quality control was implemented, removing some defective samples, such as highly blurred images and images with severely obscured lesion areas, to ensure data clarity and labeling accuracy.
[0048] The data processing includes dividing the rice leaf images into a training set, a validation set, and a test set according to a set ratio; adjusting the rice leaf images to the same size; and performing data augmentation on the rice leaf images.
[0049] In practice, the collected images were divided into training, validation, and test sets in a 4:2:1 ratio. Based on this, Labelme software was used for data annotation to facilitate subsequent training. Specifically, the lesion areas in each image were accurately classified, and all pixels of the lesion areas were precisely delineated. All label masks were named according to the disease type, and corresponding JSON files were generated. During the data annotation process, detailed text descriptions were provided for each rice leaf image. The annotations included key information such as the location, shape, and color of the lesion areas, ensuring that the annotation of each lesion area was highly accurate. Figure 2 shows the labels for rice disease images. This figure illustrates the typical visual characteristics and segmentation masks of three rice leaf diseases—rice blast, yellow dwarf disease, and bacterial blight. Rice blast is characterized by irregular white spots in the center or edge of the leaf, yellow dwarf disease is characterized by yellow stripes covering the leaf, and bacterial blight is characterized by large areas or peripherally distributed lesions on the leaf. To avoid overfitting and accelerate training, all images were preprocessed, resized to a uniform 224×224 pixels, and stored directly in the labelcol folder as input to the model. Given that deep neural network models require a large amount of diverse data to extract effective features, data augmentation was performed on the original dataset. Specific augmentation methods included geometric transformations such as horizontal and vertical flipping of the images, thereby increasing data diversity and enhancing the model's robustness and generalization ability across different scenarios.
[0050] Step S02: Construct and obtain the KBNet network model structure that can perform rice disease segmentation tasks. The KBNet network model structure includes a BC-PINN module, a KF-KAN module, and a text encoding model.
[0051] The KF-KAN module includes a KAN network structure and a Kalman filter module, with the output of the KAN network structure serving as the input of the Kalman filter module.
[0052] In specific implementation, as shown in Figure 3, Figure 3(A) shows the KBNet network module, Figure 3(B) shows the KF-KAN network module, and Figure 3(C) shows the BC-PINN network module. The KF-KAN module is a feature extraction module combining Kolmogorov-Arnold Networks (KANs) and Kalman filtering, aiming to improve the accurate identification of lesion regions by dynamically updating and fusing multi-scale features. KANs themselves have demonstrated excellent performance in multimodal tasks due to their powerful nonlinear mapping capabilities and efficient feature representation. The KF-KAN module further incorporates the Kalman filtering mechanism, enhancing its ability to process multi-scale lesion features. Kalman filtering is mainly used to dynamically update feature information, improving the model's long-term tracking and fusion capabilities of features through prediction and update mechanisms. One of the core components of this module is the adaptive residual structure, which combines fully connected layers with an adaptive weight allocation mechanism. Through this design, KF-KAN can dynamically adjust the weights of each layer of features according to changes in the feature map, strengthening the extraction of key information while reducing interference from irrelevant features. Furthermore, to enhance network stability and avoid overfitting, the KF-KAN module introduces a Dropout layer, which enhances the model's generalization ability by randomly discarding some neurons. In the processing of multi-scale lesions, KF-KAN effectively captures feature changes at different scales by fusing the nonlinear feature representation of KAN with the dynamic updates of Kalman filtering, thus improving the model's performance. The structure of the KF-KAN module is shown in Figure 3(B).
[0053] In specific implementation, in the KF-KAN module, the input is of size [value missing]. The feature map, where For batch size, The module utilizes multiple fully connected layers (including Dropout layers, activation functions, and residual connections) to progressively enhance feature representation capabilities, especially when dealing with multi-scale lesions, enabling more effective extraction of features at different scales. First, the input feature map undergoes a first fully connected layer operation, which determines the input dimension. Mapped to hidden layer dimensions .
[0054] This process can be achieved through formulas. It means that among them This is the weight matrix of the first fully connected layer. This is the bias term. The ReLU activation function introduces non-linear features, enhancing the model's expressive power. Next, it undergoes a Dropout layer to avoid overfitting when handling complex scales, which can be expressed as: .
[0055] The features processed by Dropout are fed into a second fully connected layer, maintaining the same dimensions as the hidden layers. Further enhance multi-scale feature representation. The second layer output is activated by ReLU and then subjected to Dropout to ensure efficient feature capture when processing multiple scales, as shown below: .
[0056] The features processed in the first two layers finally enter the third fully connected layer, which converts the hidden layer dimensions to... Map back to output dimension This generates the final multi-scale feature representation. The third layer no longer uses an activation function to directly generate the output features, represented as: .
[0057] To enhance information transfer between features at different scales, the KF-KAN module adds the input features and the final output features through a residual connection mechanism, forming a fused output of multi-scale information. This operation is represented as follows:
[0058] ;
[0059] Through residual connections, this module enables cross-layer information transfer, ensuring that multi-scale features from low to high levels can be effectively utilized, avoiding the loss of important low-level information in deep feature extraction. Then, the KF-KAN module uses this output as a preliminary state estimate input to the Kalman filter. The Kalman filter's workflow consists of two steps: prediction and update. First, the Kalman filter updates the state based on the current state and the state transition matrix. The prediction is expressed as follows: .
[0060] in, It is the state transition matrix. This is an estimate of the state at the previous moment.
[0061] Next, the Kalman filter is based on the actual observations. Perform state updates and calculate Kalman gain. And update the state estimate:
[0062] ;
[0063] ;
[0064] in, It is the predicted state covariance matrix. It is the observation matrix. It is the observation noise covariance matrix; Here, F represents the state estimate at time k-1, and F is the state transition matrix. This is expressed as the state estimate predicted at time k;
[0065] This represents the state estimate after the update at time k;
[0066] This represents the actual observed value at time k.
[0067] Finally, the Kalman filter updates the state covariance matrix: ,in, It is the identity matrix; Represented as the predicted state covariance matrix;
[0068] K(·) represents the Kalman gain function;
[0069] It is represented as the state covariance matrix after time k.
[0070] The Kalman filter module includes:
[0071] The output of the KAN network structure is used as a preliminary state estimate, which is then used as the state estimate for the Kalman filter. The Kalman filter makes predictions based on the state estimate and the state transition matrix, as follows: ;
[0072] The Kalman filter updates the current state based on the actual observations and calculates the Kalman gain K, expressed as: ;
[0073] ;
[0074] in, It is the predicted state covariance matrix. It is the observation matrix. It is the observation noise covariance matrix;
[0075] The Kalman filter updates the state covariance matrix as follows:
[0076] ,in, It is an identity matrix.
[0077] The KF-KAN module successfully addresses the challenges of multi-scale lesion feature extraction by combining the dynamic feature update mechanism of Kolmogorov-Arnold Networks and Kalman filtering. Through the fusion of nonlinear feature representation and Kalman filtering, KF-KAN accurately captures lesion features at different scales and enhances the model's stability and robustness through adaptive residual structures and Dropout layers. Compared to traditional methods, the KF-KAN module not only improves the ability to capture multi-scale information but also effectively avoids interference between features, ensuring accurate identification of lesion regions in complex environments. The effectiveness section discusses experimental studies of KF-KAN and comparisons with other feature extraction modules.
[0078] The BC-PINN module includes a PINN model and a polynomial loss function; the polynomial loss function is a loss set according to the task of the PINN model.
[0079] In practical implementation, to effectively handle the problem of irregular growth and complex texture of lesions, this paper proposes the BC-PINN module, specifically for lesion segmentation in rice disease images. BC-PINN embeds physical information such as lesion growth patterns into the loss function, thereby guiding the neural network to perform more accurate segmentation. The core advantage of the BC-PINN module lies in its combination of prior knowledge of physical laws with the modeling capabilities of neural networks. Furthermore, the BC-PINN module constrains the spatial location of the prediction results by calculating a penalty between the predicted mask and the image boundary. The loss function in the module includes traditional Dice loss and cross-entropy loss (BCE), combined with a penalty term based on physical information. When the predicted lesion exceeds the image boundary, the loss function penalizes it, thus preventing erroneous expansion of the prediction results. Moreover, the module adaptively adjusts the weights of physical information to ensure that the physical prior adapts to the model's training needs in different scenarios. The BC-PINN module significantly improves the model's accuracy, robustness, and convergence. By balancing computational complexity and performance, BC-PINN provides an efficient and reliable solution for lesion segmentation. The structure of the BC-PINN module is shown in Figure 3(C).
[0080] Physics-Informed Neural Networks (PHNs) directly embed physical knowledge (such as differential equations and boundary conditions) into the network training process, enabling the model to learn data distributions while adhering to physical laws. Building upon this, we propose the BC-PINN module, which not only retains the core idea of embedding prior physical knowledge in PINNs but also features optimized designs for image segmentation tasks. By introducing spatial boundary constraints during the decoding stage, BC-PINN ensures that the prediction results conform to the spatial boundary conditions of the physical scene, thereby improving the model's generalization ability and prediction accuracy.
[0081] First, the input prediction result is represented as a four-dimensional tensor: ,
[0082] in For batch size, This represents the number of channels (usually 1). These represent the height and width of the predicted image, respectively. For pixels with a predicted probability greater than a threshold (e.g., 0.5). Here, its spatial coordinates are extracted for subsequent boundary constraint calculations. The formula for extracting the coordinates is:
[0083] ;
[0084] Next, based on image size Four spatial penalty terms are constructed, corresponding to the four directions in which the predicted point crosses the boundary: left, top, right, and bottom. Specifically, for each predicted point:
[0085] ;
[0086] in, This is the modified linear unit function, used to ensure the penalty value is non-negative. The four terms above represent the distances of the predicted point beyond the left, top, right, and bottom boundaries of the image, respectively.
[0087] Next, the penalties for all points exceeding the limit are accumulated and normalized to form the PINN loss:
[0088] ;
[0089] This penalty term spatially constrains the prediction mask, explicitly suppressing predictions that extend beyond the image boundaries. During training, it facilitates backpropagation through automatic differentiation, promoting the model to generate predictions that conform to the image's spatial range.
[0090] Step S03: Use the first data as input to the KBNet network model structure for training to obtain a target model that meets the preset evaluation index.
[0091] The polynomial loss functions mentioned include Dice loss, BCE loss, and PINN loss.
[0092] In practice, to balance segmentation accuracy and spatial constraints, the BC-PINN module comprehensively considers Dice loss, BCE loss, and PINN loss in the total loss, forming the final polynomial loss function:
[0093] ;
[0094] in, , , These are the weighting coefficients for the three parts of the loss. The Dice loss and BCE loss together measure the similarity between the prediction and the true label, while the PINN loss improves the segmentation quality through spatial constraints.
[0095] In the BC-PINN module, boundary physical constraints are not only reflected in the loss calculation stage but also in the prediction and decoding process. Unlike conventional segmentation tasks that rely solely on the direct comparison between predicted probabilities and labels, BC-PINN introduces spatial location information based on the predicted distribution during decoding. This ensures that the decoding result not only reflects the data feature distribution but also follows the physical laws of spatial boundaries. In other words, by evaluating the boundary positions of the predicted mask, the BC-PINN module guides the model to learn reasonable physical boundary information during training. Even with limited labeled samples, the model can output more reasonable prediction results with the assistance of physical constraints.
[0096] The introduction of the BC-PINN module not only enhances the model's boundary awareness but also fully utilizes knowledge from the physics domain, enabling the network to rely not only on data features but also on physical laws during prediction. When labeled data is limited or uncertain, BC-PINN can effectively improve prediction stability and generalization performance through physical constraints, avoiding model overfitting. Furthermore, by penalizing the boundary of the prediction mask, the BC-PINN module significantly reduces outlier predictions, improving the model's reliability and interpretability in practical applications. For lesion growth that is irregular and has complex textures, BC-PINN can better capture subtle changes and complex boundaries of lesions. This method overcomes the limitations of traditional data-driven models in handling irregular lesion morphology, providing more reliable and accurate prediction results. The effectiveness of BC-PINN is analyzed experimentally.
[0097] The preset evaluation metrics include Dice, IOU, Accuracy, and Recall.
[0098] In this experimental investigation, four main evaluation metrics were used: Dice, IoU, Accuracy, and Recall, to comprehensively evaluate the model's performance and ensure the accuracy of the method. Initially, four different region classifications were introduced: True Positive (TP) – representing regions identified as diseased and accurately predicted; True Negative (TN) – indicating correctly identified disease-free regions; False Positive (FP) – indicating disease-free regions were incorrectly classified as diseased regions; False Negative (FN) – indicating actual diseased regions were incorrectly classified as disease-free regions.
[0099] IoU is the actual value minus the predicted value; in this example, it represents the percentage overlap between rice diseases and the label.
[0100] ;
[0101] The Dice coefficient is a similarity measure for sets. It is commonly used to calculate the similarity between two samples.
[0102] ;
[0103] Accuracy represents the proportion of all correctly segmented rice leaf images to the proportion of correctly and incorrectly segmented samples in the data:
[0104] ;
[0105] Recall measures the number of correctly predicted positive samples:
[0106] ;
[0107] Step S04: Obtain the current rice leaf image and its text description, perform data preprocessing on the current rice leaf image and its text description to obtain vectorized data, use the vectorized data as input to the target model, and output the inference from the target model to obtain the disease monitoring results of the current rice leaf image.
[0108] In practice, to ensure the accuracy and reproducibility of the experimental results, all experiments were conducted in a unified hardware and software environment. The main hardware used in this experiment included an NVIDIA GeForce RTX 4090 and a 16v CPU Intel(R) Xeon(R) Platinum 8352V CPU @ 2.10GHz. Although the specific versions of Python, CUDA, and CUDNN did not affect the experimental results, ensuring the compatibility of these software and hardware was crucial for the smooth conduct of the experiment. The hardware configuration used in the experiment was provided by the AutoDL platform, ensuring the consistency and stability of the experiment. Based on this, KBNET was implemented using PyTorch 1.8.1 and CUDA 11.1.
[0109] Table 1 details the hardware specifications and software settings;
[0110]
[0111] KF-KAN further enhances the robustness of the model by introducing an adaptive residual structure and a Dropout layer, ensuring efficient performance in complex environments. To verify the effectiveness of the KF-KAN module, it was compared with several feature extraction modules, including AIFI, MFF, BiFPN, VMamba, MDFM, and DenseNet. Experimental results for the KBNet model are detailed in Table 2. The results show that KF-KAN performs excellently on all evaluation metrics, especially demonstrating a significant advantage in handling multi-scale lesions. In contrast, while AIFI can handle multi-scale lesions within a certain range, its accuracy is low when processing lesion details, making it difficult to capture small-scale lesion features. The MFF module performs well in multi-scale feature fusion, but its noise suppression is weak. BiFPN enhances feature diversity through multi-path information transmission, but it still struggles to maintain efficient feature extraction over a large range, especially in the extraction of complex lesions. VMamba and MDFM have strong capabilities in lesion feature extraction, but their stability decreases when facing high-noise data, leading to performance fluctuations. While DenseNet can improve feature propagation through dense connections, its ability to model multi-scale lesions at a fine-grained scale is limited. In summary, the KF-KAN module is more suitable for feature extraction of multi-scale lesions. The optimized design combining KANs and Kalman filtering results in superior performance in multi-scale lesion extraction and noise suppression. Furthermore, the introduction of adaptive residual structures and Dropout layers enhances the model's segmentation accuracy.
[0112] Table 2 Experimental results of the KBNet model
[0113]
[0114] To verify the performance improvement effects of introducing the KF-KAN and BC-PINN modules on the model, four ablation experiments were conducted on a self-constructed lesion dataset. The experiments were performed under controlled conditions, and the results are recorded in Table 3, with detailed analysis and comparison. In this study, the KF-KAN module aims to improve the extraction accuracy of multi-scale lesion features, combining KANs and Kalman filtering. By optimizing the feature extraction strategy, the interference of redundant information is effectively reduced, significantly improving the accurate extraction capability of lesion regions. Experimental results show that after adding the KF-KAN module, the Dice coefficient improved by 3.1%, and the IoU improved by 5.1%. This module demonstrates significant advantages in handling complex backgrounds and multi-scale lesion features. The BC-PINN module, by introducing physical information constraints, optimizes the modeling of lesion growth patterns, helping the model maintain high segmentation accuracy when dealing with irregular lesion boundaries and complex textures. Experimental results show that compared to the basic model, BC-PINN improves the Dice coefficient by 4% and the IoU by 5.4%. This module is particularly suitable for handling complex lesion morphologies, effectively improving segmentation accuracy and model robustness. Combining the KF-KAN and BC-PINN modules further enhances model performance. Experimental results show that the combined model achieves a Dice coefficient of 83.9% and an IoU of 72.3%, representing improvements of 6.1% and 8.3% respectively compared to the base model. This result demonstrates the complementarity of the KF-KAN and BC-PINN modules and their effectiveness in improving segmentation accuracy and enhancing model stability. These four ablation experiments validate the important role of the KF-KAN and BC-PINN modules in lesion segmentation. Their combination not only improves segmentation accuracy but also enhances robustness in complex backgrounds, further demonstrating the effectiveness and advantages of these two modules in lesion segmentation.
[0115] Table 3 Comparative Experiments
[0116]
[0117] Further analysis of the KBNet model's performance was conducted, comparing it with several traditional and state-of-the-art single-modal and multi-modal segmentation methods on the same dataset. Table 4 shows the test results for each model. First, several popular single-modal segmentation methods were compared, including UNet, UNet++, Deeplabv3+, and PSPNet. UNet, with its simple and efficient encoder-decoder structure, has been widely used in tasks such as medical image segmentation. However, UNet tends to lose high-level contextual information when handling complex backgrounds or tasks with multi-scale features, affecting segmentation accuracy. UNet++ significantly improves multi-scale image segmentation capabilities by introducing skip connections to optimize feature fusion, but its more complex network structure makes training and debugging more cumbersome. Deeplabv3+ expands the receptive field by introducing dilated convolutions, enabling it to better capture global information. However, due to its deeper network structure, its convergence speed during training is slower, and its performance in some simple scenarios may not be as expected. PSPNet improves its adaptability to complex scenes by extracting features at different scales using a spatial pyramid pooling module. However, due to its multi-scale pooling design, the model has a high computational cost and performs poorly in segmentation accuracy. Compared with traditional single-modal segmentation networks, multimodal models have significantly improved segmentation accuracy, especially when dealing with multi-scattered, small points. CLIP, GLoRIA, and LViT were compared in multimodal segmentation tasks. CLIP, by combining visual and textual information, can improve model accuracy in some tasks, especially in scenarios requiring the understanding of complex semantics. However, CLIP is heavily reliant on textual data, and its generalization ability across different datasets or tasks may be limited. GLoRIA enhances its adaptability to complex scenes by optimizing the modeling of visual and semantic relationships, but its heavy reliance on semantic data limits its application scope, and in some scenarios, the feature fusion effect is not as expected. LViT combines the advantages of visual transformers and convolutional neural networks, demonstrating excellent performance in image segmentation, especially in handling details and complex textures. However, it is relatively weak in processing low-quality images and is easily affected by high-frequency illumination interference when processing rice disease images. Experimental results show that KBNet achieves higher IoU (Intersection over Union) than UNet, UNet++, Deeplabv3+, and PSPNet by 12.1%, 6.7%, 1.5%, and 40.8%, respectively, and higher IoU (Dice) by 10.5%, 5.4%, 1.3%, and 36.4%, respectively. Compared to other multimodal segmentation methods, KBNet outperforms CLIP, GLORIA, and LViT, achieving IoU improvements of 3.5%, 6.8%, and 8.3%, respectively, and higher IoU (Dice) improvements of 2.5%, 7.3%, and 6.1%, respectively.In summary, KBNet outperforms most existing segmentation networks in overall performance, especially in applications involving multimodal feature fusion, where it demonstrates superior performance. The advantages of the KBNet model stem from the following points: (a) KBNet combines visual and semantic information, overcoming the limitations of relying solely on visual information. (b) The KF-KAN module, combining KANs and Kalman filtering, optimizes the feature extraction strategy, significantly improving the accuracy of lesion region extraction. (c) The BC-PINN module, by introducing physical constraints, optimizes the modeling of lesion growth patterns, helping the model maintain high segmentation accuracy when dealing with irregular lesion boundaries and complex textures. (d) The self-built dataset in this study eliminated several blurry and low-quality images, which is beneficial for model training.
[0118] Table 4 Comparative experimental results of different models
[0119]
[0120] Table 5 shows how to visualize the segmentation results of LViT and KBNet. To better understand KBNet, different methods were used to segment three diseases: Bacterial Blight, Blast, and Tungro.
[0121] In Group A, segmentation experiments were conducted on typical Bacterial Blight lesion regions. The results showed that the LViT model still exhibited some ambiguity in lesion edge localization, particularly in areas with fine and dense lesions where edge breaks were prone to occur. In contrast, KBNet significantly outperformed the comparison methods in detail recognition, capable of reconstructing lesion shapes completely and accurately. This performance improvement is primarily attributed to the introduced KF-KAN module, which combines the advantages of KAN in nonlinear modeling with the dynamic state update capability of Kalman filtering, significantly enhancing the model's ability to perceive complex textures and small lesions.
[0122] In Group B, Blast image samples with blurred boundaries were selected. Experiments revealed that LViT is prone to missegmentation or region adhesion when dealing with lesions that have diffused edges and irregular shapes. KBNet, by introducing the BC-PINN module, embeds prior knowledge such as lesion growth mechanisms into the loss function, enabling the model to accurately model non-rigid lesion structures. Simultaneously, the addition of a boundary penalty mechanism further constrains the segmentation edges, improving the model's resolution of lesion boundaries, making KBNet significantly superior to other methods in areas with blurred edges.
[0123] In Group C, images containing Tungro lesions and complex backgrounds were used for testing. The results showed that KBNet still exhibited stable and excellent segmentation performance, accurately capturing the main lesion region against complex interference backgrounds while suppressing background misclassification. KF-KAN's efficient fusion of local and global features, and BC-PINN's modeling of spatial diffusion characteristics, enabled KBNet to maintain strong adaptability and robustness in real field environments.
[0124] In summary, through the synergistic effect of the two core modules, KF-KAN and BC-PINN, KBNet not only achieves stronger feature representation and boundary awareness capabilities, but also significantly optimizes the overall disease segmentation quality, providing reliable support for subsequent intelligent disease diagnosis.
[0125] Table 5. Visualization of segmentation results from LViT and KBNet
[0126]
[0127] Compared with the prior art, the technical effects of this application are as follows:
[0128] To improve the ability of rice disease segmentation models to identify complex lesion regions, an image dataset was constructed, covering three common rice diseases: bacterial blight, rice blast, and tungro. All image samples were precisely annotated pixel-by-pixel using the Labelme tool to generate corresponding segmentation masks, ensuring the accuracy of lesion region information.
[0129] A Kalman Filter Enhanced KAN (KF-KAN) module is proposed. This module efficiently captures nonlinear features at different scales using KAN and dynamically updates and fuses multi-scale information by combining Kalman filtering, achieving accurate feature extraction of complex lesion regions. By adaptively adjusting scale and detail, KF-KAN effectively improves the model's accuracy and robustness in handling multi-scale lesions, thereby enhancing its overall performance in segmentation tasks.
[0130] A boundary-constrained physical information neural network module (BC-PINN) is proposed, which combines prior knowledge of physical laws with the powerful modeling capabilities of neural networks. By embedding physical information such as lesion growth patterns into the loss function, BC-PINN can effectively constrain the model's prediction of irregularly growing lesions. Furthermore, the BC-PINN module further constrains the spatial location of the prediction results by calculating the penalty between the predicted mask and the image boundary, thus providing more accurate segmentation results.
[0131] The proposed KBNet model achieves an Intersection over Union (IoU) ratio of 72.3% and a Dice coefficient of 83.9% on a self-constructed dataset. This model performs exceptionally well in rice disease segmentation tasks, providing strong technical support for the accurate identification and control of rice diseases and promoting the intelligent and efficient implementation of rice disease management.
[0132] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for identifying and monitoring rice diseases based on the KBnet network model, characterized in that, The rice disease identification and monitoring method includes: acquiring collected rice leaf images and text descriptions of the rice leaf images; processing the rice leaf images and text descriptions to obtain first data adapted to model input; constructing and obtaining a KBNet network model structure capable of performing rice disease segmentation tasks, the KBNet network model structure including BC-PINN modules and KF-KAN modules; using the first data as input to the KBNet network model structure for training to obtain a target model that meets preset evaluation indicators; acquiring current rice leaf images and text descriptions of the current rice leaf images; preprocessing the current rice leaf images and text descriptions of the current rice leaf images to obtain vectorized data; using the vectorized data as input to the target model; and inferring and outputting the target model to obtain the disease monitoring results of the current rice leaf images.
2. The rice disease identification and monitoring method based on the KBnet network model as described in claim 1, characterized in that, The KF-KAN module includes a KAN network structure and a Kalman filter module, with the output of the KAN network structure serving as the input of the Kalman filter module.
3. The rice disease identification and monitoring method based on the KBnet network model as described in claim 2, characterized in that, The Kalman filter module includes: using the output of the KAN network structure as a preliminary state estimate, using the preliminary state estimate as the state estimate of the Kalman filter, and the Kalman filter performing prediction based on the state estimate and the state transition matrix, expressed as: The Kalman filter updates its current state based on actual observations and calculates the Kalman gain K, expressed as: in, It is the predicted state covariance matrix. It is the observation matrix. It is the transpose of the observation matrix. It is the observation noise covariance matrix. State estimation; Here, F represents the state estimate at time k-1, and F is the state transition matrix. This is expressed as the state estimate predicted at time k; This represents the state estimate after the update at time k; The actual observed value at time k is represented as: The Kalman filter updates the state covariance matrix as follows: ,in, It is the identity matrix; It is represented by the predicted state covariance matrix; K(·) represents the Kalman gain function; It is represented as the state covariance matrix after time k.
4. The rice disease identification and monitoring method based on the KBnet network model as described in claim 1, characterized in that, The BC-PINN module includes a PINN model and a polynomial loss function; the polynomial loss function is a loss set according to the task of the PINN model.
5. The rice disease identification and monitoring method based on the KBnet network model as described in claim 4, characterized in that, The polynomial loss functions mentioned include Dice loss, BCE loss, and PINN loss.
6. The rice disease identification and monitoring method based on the KBnet network model as described in claim 1, characterized in that, The preset evaluation metrics include Dice loss, IOU loss, accuracy, and recall.
7. The rice disease identification and monitoring method based on the KBnet network model as described in claim 1, characterized in that, The data processing includes dividing the rice leaf images into a training set, a validation set, and a test set according to a set ratio; adjusting the rice leaf images to the same size; and performing data augmentation on the rice leaf images.