Fruit and vegetable disease and insect pest spectrum database dynamic evolution updating method based on model confidence feedback driving

By using a model confidence feedback-driven approach, combined with expert systems and confidence calibration, the problem of accuracy verification in fruit and vegetable pest and disease detection was solved. This enabled dynamic updating of the pest and disease spectral database and model optimization, thereby improving the accuracy and adaptability of detection.

CN120913003APending Publication Date: 2025-11-07NINGXIA UNIVERSITY

Patent Information

Application Number
CN202510897473.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In the existing technology, the methods for detecting diseases and pests in fruits and vegetables cannot be updated in a timely manner, and the accuracy of disease and pest detection cannot be verified, leading to the problem of erroneous detection.

Method used

A dynamic evolution update method for the spectral database of fruit and vegetable diseases and pests based on model confidence feedback is adopted. By obtaining the disease and pest diagnosis results and their model confidence, an expert system is used to correct uncertain results, and the reliability of model confidence calculation is improved by entropy regularization and temperature calibration.

Benefits of technology

The dynamic evolution and updating of the spectral database of fruit and vegetable diseases and pests has been realized, ensuring that only highly reliable results enter the database, preventing the accumulation of misjudgments, and improving the accuracy and adaptability of disease and pest detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913003A_ABST
    Figure CN120913003A_ABST
Patent Text Reader

Abstract

The invention discloses a fruit and vegetable disease and insect pest spectrum database dynamic evolution updating method based on model confidence feedback driving. The method comprises the following steps: acquiring model confidence of disease and insect pests; based on a preset confidence model, judging whether the model confidence of the diseases and pests is not less than a threshold value of high model confidence: if yes, outputting an identification result and an image, and inputting the identification result and the image into a disease and pest spectrum database; if not, judging whether the model confidence coefficient corresponding to the disease and pest diagnosis result is not less than a low threshold value of the model confidence coefficient; if so, executing an expert reexamination process; if not, marking the sample as a low-confidence sample, and inputting the low-confidence sample into an abnormal sample library. By establishing a data updating and model optimization mechanism for cooperative work of experts and models, misjudgment accumulation caused by excessive confidence of the models is effectively prevented, and it is ensured that only high-reliability results can automatically enter a database; by means of expert labeling feedback and incremental self-learning of the model, dynamic evolution and updating of the fruit and vegetable disease and insect pest spectrum database are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of agricultural remote sensing and artificial intelligence, and particularly relates to a fruit and vegetable disease and pest spectrum database dynamic evolution updating method based on model confidence feedback driving. BACKGROUND

[0002] Timely and accurate diagnosis of fruit and vegetable diseases and pests is crucial for agricultural production. Traditional methods often rely on manual observation and experience, which is low in efficiency and easily affected by subjective factors.

[0003] In the prior art, for example, a Chinese invention patent with application number CN202410635672.3 discloses a fruit and vegetable video disease and pest detection method based on deep learning and big data technology, which specifically discloses a camera module, a deep learning model training module, a video detection display module, a detection data collection module, a real-time detection module and an alarm module, an offline detection module and a display module. The camera module uses a camera installed in a greenhouse to shoot a set of range-different fruit and vegetable growth videos every 24 hours. The deep learning model training module includes an iteration unit and an identification unit. The video detection display module is an interface generated by software development, which displays the fruit and vegetable disease and pest detection status in real time. The real-time detection module processes and analyzes the detection data in real time, and triggers the alarm module when a certain index is reached. The offline detection module processes and analyzes the detection data offline, and finally displays the results in the display module.

[0004] However, the accuracy of disease and pest detection cannot be known in the above-mentioned scheme, and new diseases and pests or disease and pest variations may occur. The above-mentioned method cannot accurately determine the diseases and pests, and the diseases and pests cannot be updated in time. SUMMARY

[0005] Therefore, it is necessary to verify the accuracy of disease and pest detection, and to provide a fruit and vegetable disease and pest spectrum database dynamic evolution updating method based on model confidence feedback driving to solve the problem of easy disease and pest detection errors.

[0006] To achieve the above-mentioned purpose, the present application adopts the following scheme:

[0007] A fruit and vegetable disease and pest spectrum database dynamic evolution updating method based on model confidence feedback driving includes the following steps:

[0008] Step S10. Obtain the disease and pest diagnosis result and the corresponding model confidence;

[0009] Step S20. Based on the pre-set confidence model, determine whether the model confidence corresponding to the disease and pest diagnosis result is not less than the high threshold of the model confidence:

[0010] If yes, output the recognition result and the image, and input to the pest spectrum database;

[0011] If no, output the model confidence corresponding to the pest diagnosis result to the expert system module;

[0012] Step S30. Obtain the model confidence less than the high threshold of the model confidence, and determine whether the model confidence corresponding to the pest diagnosis result is not less than the low threshold of the model confidence based on the expert system module:

[0013] If yes, the expert identifies and confirms the diagnosis result and the corresponding model confidence, corrects the diagnosis result, and inputs the corrected diagnosis result into the pest spectrum database and the abnormal sample library;

[0014] If no, the expert identifies the diagnosis result and the corresponding model confidence, outputs the correct diagnosis result, and inputs it into the pest spectrum database.

[0015] Preferably, in step S10, the model confidence of the pest is obtained, the model confidence is Conf total , and the model confidence Conf total is obtained by weighting the classification probability confidence Conf cls , the prediction uncertainty confidence Conf unc , and the embedding distance confidence Conf dist .

[0016] Preferably, the classification probability confidence Conf

[0017]

[0018] wherein, represents the prediction probability of the model for the i-th class, z i represents the logit value of the last layer of the model, C represents the total number of categories, and T is a temperature parameter, and T>0.

[0019] Preferably, the prediction uncertainty confidence Conf unc = exp(-λσ 2 ); wherein,

[0020]

[0021] is the average value of all M predictions, M is the number of forward times, p (1) ,…, p (M) are the probabilities corresponding to the M groups, σ is the standard deviation, and λ is a hyperparameter for controlling the decay speed.

[0022] Preferably, the embedding distance confidence Confdist = exp(-γD M (x))

[0023] wherein,

[0024]

[0025] γ is a scaling hyper-parameter, f(x) is the embedding representation of the input sample in the feature space, μ is the mean of a class of the training set, T is a temperature parameter, Σ is the covariance matrix of the features, D is the Mahalanobis distance, which measures the standardized distance between the sample and the class center.

[0026] Preferably, the model confidence c = [Conf cls , Conf unc , Conf dist ] combines three kinds of confidence into an input vector c, and the model confidence weight w is:

[0027] w = Softmax(W·ReLU(V·c))

[0028] wherein: V, W are weight matrices to be learned, ReLU is an activation nonlinearity, and Softmax makes the weights satisfy equal to 1.

[0029] Preferably, in step S10, the model confidence of the disease and pest is obtained; further comprising calibrating the model confidence of the disease and pest, and the calibration adopts entropy regularization and temperature calibration, wherein the entropy value L entropy of the entropy regularization is:

[0030]

[0031] wherein, is the prediction probability of the i-th class, and C is the total number of classes.

[0032] Preferably, the temperature calibration calibration index ECE is:

[0033]

[0034] wherein, K is the number of bins into which the prediction confidence interval is divided, B k is the sample set in the k-th bin, n is the total number of samples, acc(B k ) is the average accuracy of the samples in the bin, and conf(B k ) is the average confidence of the samples in the bin.

[0035] The technical scheme adopted in the present application can achieve the following beneficial effects:

[0036] 1. By calculating and comparing the model confidence of the output diagnosis result, the model confidence of the diagnosis result is given, solving the problem of only detecting pests and diseases without verifying the accuracy of pest and disease detection, resulting in pest and disease detection errors.

[0037] 2. The model confidence calculation is more reliable through entropy regularization and temperature scaling, and will not produce a high confidence score due to model overfitting or sample distribution deviation.

[0038] 3. By introducing model confidence fusion calculation and hierarchical processing in the diagnosis process, a data updating and model optimization mechanism for expert and model collaborative work is established. This method effectively prevents misjudgment accumulation caused by model overconfidence, ensuring that only high-reliability results are automatically entered into the database, and uncertain results are corrected by experts. With the help of expert annotation feedback and model incremental self-learning, dynamic evolution and update of the fruit and vegetable disease and pest spectrum database are realized. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 A flowchart of a fruit and vegetable disease and pest spectrum database dynamic evolution and update method based on model confidence feedback driving disclosed in an embodiment of the present application.

[0040] Figure 2 A category and temperature statistical chart of a fruit and vegetable disease and pest spectrum database dynamic evolution and update method based on model confidence feedback driving disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0041] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings. The preferred embodiments of the present application are shown in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0042] It should be noted that when a device is considered to be "connected" to another device, it can be directly connected to the other device or there can be a mediating device present at the same time. The terms "internal", "top", "upper", "lower", "up", "down" and similar expressions used herein are for illustrative purposes only and do not represent the only implementation.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application herein are only for the purpose of describing the specific embodiments and are not intended to limit the present application. The term "and / or" used herein includes any and all combinations of one or more related listed items.

[0044] Referring to Figure 1 and Figure 2 , the application provides a fruit and vegetable pest spectrum database dynamic evolution updating method based on model confidence feedback driving, comprising the following steps:

[0045] Step S10. Obtain the pest diagnosis result and the corresponding model confidence;

[0046] Step S20. Based on the pre-set confidence model, determine whether the model confidence corresponding to the pest diagnosis result is not less than the high threshold of the model confidence:

[0047] If yes, output the recognition result and the image, and input them into the pest spectrum database;

[0048] If not, output the model confidence corresponding to the pest diagnosis result to the expert system module;

[0049] Step S30. Obtain the model confidence less than the high threshold of the model confidence, and based on the expert system module, determine whether the model confidence corresponding to the pest diagnosis result is not less than the low threshold of the model confidence:

[0050] If yes, the expert identifies and confirms the diagnosis result and the corresponding model confidence, corrects the diagnosis result, and inputs the corrected diagnosis result into the pest spectrum database and the abnormal sample library;

[0051] If not, the expert identifies the diagnosis result and the corresponding model confidence, outputs the correct diagnosis result, and inputs it into the pest spectrum database.

[0052] Specifically, when the spectral data of fruit and vegetable crops is acquired, DJIM350RTK unmanned aerial vehicle is used in the front end of the system to carry RedEdge-P multispectral camera, and a GPS high-precision positioning module is integrated. The platform can automatically cruise in the field and acquire crop image information, and the collection range covers blue light, green light, red light, red edge, near-infrared and other key vegetation bands, providing a multi-dimensional spectral data basis for subsequent diagnosis.

[0053] The spectral data acquisition includes the following steps:

[0054] a. Raw image acquisition Raw image acquisition is obtained by one-time exposure of RedEdge-P multispectral camera carried on the unmanned aerial vehicle 5 gray images

Blue (475nm), Green (560nm), Red (668nm), Red-Edge (717nm), NIR (842nm) can be understood as "taking 5 black and white photos with different filters at the same moment"

[0055] b. Image orthorectification

[0056]

[0057] c. Radiometric calibration and reflectance calculation, 5 GeoTIFFs loaded in QGIS simultaneously, DN converted to physical reflectance with raster calculator: Equation 3-1

[0058]

[0059] where DN ij is the pixel gray value, Dark is the dark level, Gain is the camera gain, T exp is the exposure time, R panel is the reflectance of the gray panel

[0060] d. Different sun angles will cause reflectance differences, which can be scaled by a simplified cosine:

[0061]

[0062] where θ sun is the sun's zenith angle (check flight log), θ inc,ij is the pixel incidence angle, calculated from the drone's attitude, R ij is the reflectance value at the ith row, jth column pixel.

[0063] e. ROI cropping and band-merging vector import: Load the field block Shapefile in QGIS; Crop by mask: Menu Raster -> Extract -> Crop by Mask;

[0064] Multi-band synthesis: Raster -> Merge -> Build Virtual Raster (VRT) -> Check "Layer to Multi-band", output spectral.vrt.

[0065] Finally, each pixel has a 5-dimensional feature vector:

[0066] X(i,j) = bigl[R Blue ,, R Green ,, R Red ,, R RE ,, R NIR bigr] ij

[0067] This five-channel tensor is directly fed into the deep learning model in the next section.

[0068] After obtaining the spectral data, the system automatically performs image preprocessing operations such as radiation correction, geometric registration, illumination normalization, distortion correction, and outputs standardized multi-band image data, ensuring the spectral consistency and spatial alignment accuracy of the model input.

[0069] The pest diagnosis model is deployed in the unmanned aerial vehicle edge computing terminal or ground station, adopts a lightweight convolutional neural network (CNN) architecture, performs pest species identification and grade judgment on the preprocessed images, and outputs the fused comprehensive model confidence. In the pre-set confidence model, if the model confidence ≥ 0.9, the model prediction is highly reliable; if 0.7 ≤ model confidence < 0.9, the model has some uncertainty; if the model confidence < 0.7, the model confidence is too low. Each identification is equivalent to an update of the pest spectrum database, and each execution of the expert system module is equivalent to an update of the abnormal sample library; through continuous judgment and execution, the dynamic evolution and update of the fruit and vegetable pest spectrum database are completed.

[0070] Further, first, a multi-spectral imaging device is used to obtain high-resolution multi-spectral image data by flying over the fruit and vegetable planting area. Preferably, the unmanned aerial vehicle used, such as DJI M350 RTK, is equipped with a RedEdge-P multi-spectral camera to obtain image information of multiple bands such as visible light and near-infrared. The collected spectral images are corrected and optimized through a preprocessing step, including but not limited to spectral calibration, noise removal, image registration, and region of interest extraction, etc., to provide high-quality input data for subsequent model analysis.

[0071] Then, the preprocessed spectral data is input into the pre-trained pest diagnosis model for inference. The model can use deep learning algorithms such as convolutional neural networks, which have been trained using a large number of labeled samples in the initial spectral database, and can output pest species diagnosis results for input samples and model confidence evaluation of the results. The output obtained by model inference includes: first, the diagnosis result (such as determining the type of crop disease or the species of pest), and second, the model confidence value, which represents the credibility of the model to the diagnosis result. The worker confirms the accuracy of his judgment through the diagnosis result and the corresponding model confidence.

[0072] After the pest diagnosis result and the corresponding model confidence are obtained, the pest diagnosis result and the corresponding model confidence are compared with the preset confidence model. If the model confidence is greater than or equal to 0.9, it is indicated that the model prediction is highly reliable, no manual intervention is needed, and the recognition result and the image are directly written into the pest spectrum database, and the diagnosis result and the corresponding model confidence are fed back to the staff (through a display screen, a wireless network, Bluetooth or the like, as long as the staff can see it). If the model confidence is less than 0.9, the model confidence corresponding to the pest diagnosis result is output to the expert system module. The expert system module is manually labeled and classified by a pest diagnosis expert according to the image and spectrum data to ensure the accuracy of the recognition result and the robustness of the system. The label data output by the expert system module is written into the system database as a new training sample and is used for online fine-tuning, migration training or incremental learning of the subsequent model to continuously improve the recognition performance and adaptability of the model in the actual environment, forming a complete learning closed loop.

[0073] The technical scheme of the fruit and vegetable pest spectrum database dynamic evolution updating method based on model confidence feedback driving adopted in the application can achieve the following beneficial effects:

[0074] 1. The model confidence of the output diagnosis result is calculated and compared to give the model confidence of the diagnosis result, solving the problem of incorrect pest detection due to the lack of verification of the accuracy of pest detection.

[0075] 2. The model confidence calculation is more reliable through entropy regularization and temperature scaling, and will not produce a high confidence score due to model overfitting or sample distribution deviation.

[0076] 3. The model confidence fusion calculation and hierarchical processing are introduced in the diagnosis process to establish a data updating and model optimization mechanism for experts and models to work together. This method effectively prevents misjudgment accumulation caused by excessive confidence of the model, ensures that only high-reliability results are automatically entered into the database, and uncertain results are modified by experts. With the help of expert annotation feedback and model incremental self-learning, the dynamic evolution updating of the fruit and vegetable pest spectrum database is realized.

[0077] In the above scheme, step S10. Obtain the model confidence of the pest, the model confidence Conf total , and the model confidence Conf total is obtained by weighting the classification probability confidence Conf cls , the prediction uncertainty confidence Conf unc and the embedding distance confidence Conf dist .

[0078] Specifically, the classification probability confidence

[0079]

[0080] where, represents the predicted probability of the i-th class by the model, z i represents the logit value of the last layer of the model, C represents the total number of classes, and T is a temperature parameter (T > 0) that adjusts the smoothness of the softmax. When T = 1, it is a standard softmax; when T > 1, the output is more gentle (the distribution is more "soft"); when T < 1, the output is more sharp (the probability of a certain class is very high, and the others are almost zero). By adjusting T, the softmax output probability is more matched with the actual accuracy, and the confidence calibration ability is improved.

[0081] Taking the maximum probability:

[0082]

[0083] represents the class with the highest probability in the softmax output as the classification confidence of the sample; the closer the probability is to 1, the more confident the model is in that class. For example, Figure 2 where: Class1, Class2, Class3 represent three different class labels output by the model; T = 0.5: the distribution is very sharp, and the model is very "confident" and basically only selects one class; T = 1.0: standard Softmax; T = 2.0: the distribution is more gentle, and the confidence decreases; T = 5.0: close to the average distribution, indicating that the model is very uncertain.

[0084] Further, the prediction uncertainty confidence Conf unc = exp(-λσ 2 ); where,

[0085]

[0086] is the average value of all M predictions, M is the number of forward times, p (1) ,…, p (M) is the probability corresponding to the M groups, σ is the standard deviation, and λ is a hyperparameter that controls the decay speed.

[0087] Further, the embedding distance confidence Conf dist = exp(-γD M (x)) where,

[0088]

[0089] where g is a scaling hyper-parameter, f(x) is the embedding representation of the input sample in the feature space, m is the mean of a class of the training set, T is a temperature parameter, S is the covariance matrix of the features, D is the Mahalanobis distance, measuring the standardized distance between the sample and the class center.

[0090] The model confidence c = [Conf cls , Conf unc , Conf dist ] combines three kinds of confidence into an input vector c, and the model confidence weight w is:

[0091] w = Softmax (W ReLU (V c))

[0092] where: V, W are the weight matrices to be learned, ReLU is the activation nonlinearity, and Softmax makes the weights satisfy equal to 1.

[0093] Model confidence calculation method: In this embodiment, the calculation of the model confidence C integrates information in three dimensions to improve the accuracy of the credibility evaluation. Specifically, C is determined based on the following three aspects of measurement:

[0094] 1) Classification probability: the Softmax maximum probability value P_max output by the model, representing the intuitive confidence degree of the model for the predicted pest and disease class.

[0095] 2) Prediction uncertainty: calculate the uncertainty quantitative index through the prediction output of the model, for example, use Monte Carlo dropout multiple forward reasoning or the standard deviation s based on the prediction result of the integrated model, to measure the stability of the model prediction for the sample, the larger the standard deviation, the higher the uncertainty and the lower the confidence.

[0096] 3) Embedding feature distance: input the sample to be tested into the model to extract its representation vector in the embedding feature space, compare the vector with the feature vectors of the known samples in the model training set, calculate the nearest neighbor distance D, the greater the distance, the higher the degree of deviation of the sample from the distribution of the training data (i.e. it may be a novel or abnormal sample), and the lower the confidence should be. Normalize the three indicators P_max, σ and D obtained above, and input them into a dynamic weighted fusion model to calculate the final confidence C. Preferably, the fusion model can use a trained multi-layer perceptron (MLP) network or a preset weighting formula to adaptively adjust the weights of each indicator according to historical validation data, so as to accurately evaluate the model confidence in different situations. For example, if the Softmax probability of a certain type of disease is always high, resulting in overconfidence, the fusion model will appropriately reduce the weight of P_max in the comprehensive confidence, and enhance the consideration of uncertainty σ and feature distance D. Through this dynamic weight confidence fusion method, the reliability of the model's judgment on the current sample can be more comprehensively described.

[0097] In addition, in order to prevent the model from making overconfident judgments, confidence calibration and regularization measures are introduced during model training and inference. On the one hand, an entropy regularization term is added in the model training stage to encourage the model output to maintain appropriate information entropy, thereby avoiding the distortion of confidence evaluation caused by excessively extreme output probability; on the other hand, temperature scaling calibration technology is applied to the output probability in the model deployment and inference stage, that is, a temperature coefficient is introduced during Softmax calculation to stretch or compress the probability distribution (for example, temperature T>1 is introduced to reduce the maximum probability value and increase the entropy), so that the confidence of the model output is better matched with the actual accuracy. Through entropy regularization and temperature scaling, the model confidence calculation is more reliable and will not produce a high confidence score due to model overfitting or sample distribution deviation.

[0098] In an embodiment of the present application, the model confidence of the plant disease and pest is obtained in step S10; further comprising calibrating the model confidence of the plant disease and pest, and the calibration adopts entropy regularization and temperature calibration, wherein the entropy value L of the entropy regularization is entropy :

[0099]

[0100] wherein, The prediction probability of the i-th class, C is the total number of classes. Entropy measures the uncertainty or chaos of a probability distribution. When the softmax output is very "confident" (such as [0.99, 0.005, 0.005]), the entropy is very low; when the softmax output is very "ambiguous" (such as [0.33, 0.33, 0.34]), the entropy is very high; therefore, by increasing the entropy, the model can be encouraged to be cautious on uncertain samples and reduce overconfidence. Prevent overfitting: especially when the data categories are unbalanced or the sample quality is poor; improve robustness: the output distribution is more conservative and does not easily give too high confidence; as an additional item of the loss function: often written as

[0101] L total =L task +β·L entropy

[0102] L total : Total Loss, the objective function that the model is finally optimized for;

[0103] L task : TaskLoss, the basic loss of the core task;

[0104] L entropy : Entropy Regularization, which punishes overly confident outputs and improves the model's ability to express uncertainty;

[0105] β: is a weight hyperparameter that controls the influence of entropy regularization in the total loss.

[0106] In the above scheme, the temperature calibration calibration index ECE is:

[0107]

[0108] Where K is the number of bins into which the prediction confidence interval is divided, B k is the sample set in the k-th bin, n is the total number of samples, acc(B k ) is the average accuracy of samples in the bin, and conf(B k ) is the average confidence of samples in the bin.

[0109] The generation and training of the above-mentioned embedded diagnosis model includes the following steps: a training target function

[0110]

[0111] y k : true label, a "switch value", indicating that the sample belongs to the k-th class when it is 1, (for example, the probability that the model considers it to be "downy mildew" is 0.8).

[0112] First term (cross-entropy): measures the error between the predicted result and the true class, the smaller the value, the more accurate the prediction.

[0113] Second term (entropy regularization): entropy measures the "uniformity of the predicted distribution", adding this term is to avoid the model being too "confident" to choose only one class; if the model always outputs [1.0, 0.0, 0.0], it is easy to overfit; encourage output like [0.7, 0.2, 0.1] to keep the probability of uncertainty.

[0114] λ1: adjust the size of this term on the total loss.

[0115] Third term (confidence supervision):

[0116] C0: the confidence value originally calculated by the model;

[0117] C ★ : "ideal confidence" based on expert opinion or other methods;

[0118] If the model's confidence deviates too much from the actual credibility, there will be a larger penalty; λ2: adjust the proportion of this penalty in the total loss.

[0119]

[0120]

[0121] b. Temperature scaling calibration (improve confidence reliability)

[0122] Find a single scalar T on the validation set ★ Minimize negative log-likelihood (NLL), calibrated probability is

[0123]

[0124] z k : the "score" output by the last layer of the model, which has not yet become a probability value;

[0125] T ★ : temperature parameter, adjust the "confidence level":

[0126] If T ★ >1, the predicted probability is more "smooth", and the confidence is more conservative;

[0127] If T ★ <1, the prediction result is more "extreme";

[0128] exp: exponential function, Softmax is a standard probability transformation function;

[0129] Whole paragraph explanation: This step is "temperature scaling", which makes the model's prediction results and actual accuracy more matched (such as avoiding the model outputting a high confidence of 0.99 but actually often wrong).

[0130] c. Incremental fine-tuning and model version management trigger conditions: new expert samples ≥200 or online running >14 days;

[0131] Fine-tuning strategy: freeze the first two feature layers, only update the subsequent layers, 20 epochs of convergence; Model hash record: facilitate OTA rollback and audit.

[0132] In another embodiment of the present application, the confidence-driven data update mechanism is as follows

[0133]

[0134] Wherein, the initial threshold: T H =0.9, T L =0.7; Recalculate 90% / 70% quantile every 1k samples, realize self-adaption. The incremental training of the model is stored in the pest spectrum database and the abnormal sample library, and the trigger of the incremental training is:

[0135] Trigger=(N exp ≥N th )or(R low >R th )or(Δt≥T cycle )

[0136] Wherein, N exp : the number of newly added expert labeled samples, N th : expert sample trigger threshold, R low : current low confidence sample proportion, R th : the proportion threshold of low confidence samples, Δt: the time experienced since the last model training, T cycle : training time period threshold.

[0137] Data sampling weight, expert sample: weight w=1; High confidence sample: weight w=Conf total ; If the online average confidence decreases by more than 3% within 24 hours, automatically roll back to the last stable model version.

[0138] Confidence-driven data updating mechanism: after the model completes the diagnosis of the new sample and calculates the confidence, the application processes the sample differently according to the level of the confidence, and constructs a dynamic database updating and model optimization cycle. Preferably, the confidence is divided into high, medium and low three levels: when the model confidence C≥0.9, it is considered that the model is very reliable for the result judgment, and the diagnosis result of the model can be directly adopted; when 0.7≤model confidence C<0.9, it is considered that the model is of medium reliability, and the reliability of the result is doubtful; when the model confidence C<0.7, it is considered that the model is of low confidence, and the result is likely to be inaccurate or an abnormal case that the model has not seen before.

[0139] The system adopts the following processing strategies according to the above confidence levels: for samples with high confidence (model confidence≥0.9), the model-predicted pest and disease species result is directly regarded as valid information, and the sample and its model-predicted label are recorded into the fruit and vegetable pest and disease spectrum database as new data for storage; for samples with medium confidence (0.7-0.9), they are marked as to be audited and submitted to experts for review and judgment; for samples with low confidence (<0.7), they are marked as abnormal to be handled and also submitted to experts for identification. Through the division of the three-level confidence interval, different treatments can be performed according to the reliability of the results, which not only ensures the timely use of high-confidence results, but also ensures that unreliable results are intervened by humans and do not mislead the database with incorrect information.

[0140] After the submission of the medium and low confidence samples, the system sends the related spectrum data and model preliminary diagnosis results to the remote expert diagnosis platform through the network, and the artificial experts with plant pest and disease diagnosis experience identify and label these samples. The experts view the spectrum images of the samples and the preliminary conclusions given by the model, combine their own knowledge to judge the actual pest and disease situation, and give the authoritative diagnosis result label. For medium confidence samples, the model may be partially correct but not certain enough, and the experts focus on verifying the features that the model is easily confused, and confirm and correct the label; for low confidence samples, the model is almost unable to give meaningful judgment, and the experts will conduct complete artificial diagnosis on these possibly new or abnormal cases. Once the experts complete the labeling, the system archives the sample data with the confirmed label of the experts to the abnormal sample library. The abnormal sample library is used to store the cases that the model has been uncertain or misjudged before, and provides targeted training data for subsequent model improvement. In addition, all the samples labeled by the experts (including medium and low confidence categories) are added to the main spectrum database together with their correct labels, as new training samples for storage. At the same time, for the high confidence samples that have been automatically added to the database, the system also makes corresponding records (such as marking that the label comes from model prediction rather than manual), for reference during training.

[0141] With the passage of time, the newly added samples accumulated in the spectrum database, including high-confidence automatic addition samples and expert-annotated addition samples, are increasing. In order to make the model learn these new data in time, the present application sets up a dynamic triggering mechanism for model incremental training. The triggering conditions can be set in advance, for example, when the number of newly added samples in the spectrum database reaches a certain threshold, or when the new data is checked periodically (such as every few days / weeks), the update training of the model is triggered. During the incremental training, the above-mentioned newly added expert-annotated samples and high-confidence samples are added to the model training set, and the original diagnostic model is retrained or fine-tuned online. By introducing expert-annotated data, the model can correct previous errors or uncertain recognition, and significantly improve the accuracy on the corresponding category; and adding high-confidence model prediction samples (i.e. the model's confident correct guess) is equivalent to using reliable data generated by the model itself to further enrich the training set, which helps to consolidate the model's grasp of existing knowledge. During the training process, methods such as entropy regularization are continued to be applied to improve the generalization ability, and after training, temperature scaling and other means are used to calibrate the confidence of the new output of the model. After completing the update training, the new model is deployed to the embedded diagnostic system in the front end for subsequent field inspection and diagnosis, thus entering the next round of operation cycle. In this way, the model and the database continue to evolve: the model becomes more robust and can recognize more diverse pest and disease spectra; the spectrum database also expands new instances, covering more comprehensive pest and disease spectral characteristics. The entire system forms a closed loop through the model confidence feedback mechanism, and autonomously learns and improves in practical application, significantly improving the efficiency and accuracy of fruit and vegetable pest and disease diagnosis.

[0142] In summary, the present application establishes a data update and model optimization mechanism for experts and models to work together by introducing model confidence fusion calculation and hierarchical processing in the diagnosis process. This method effectively prevents the accumulation of misjudgments caused by model overconfidence, ensuring that only high-reliability results are automatically entered into the database, and uncertain results are all corrected by experts. With the help of expert annotation feedback and model incremental self-learning, the present application realizes the dynamic evolution and update of the fruit and vegetable pest and disease spectrum database, as well as the continuous optimization and upgrading of the diagnosis model, which is of great significance for ensuring the accuracy and adaptability of pest and disease identification in agricultural production.

[0143] In order to enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the embodiments.

[0144] Embodiment:

[0145] (1) Experimental purpose: This experiment aims to build a closed-loop cabbage disease and pest identification system with "data collection - model identification - confidence feedback - expert review - database expansion - model evolution" as the core. Based on multispectral remote sensing technology and deep learning methods, the system's recognition accuracy, model self-adaptive evolution ability, and deployment feasibility for cabbage main diseases and pests (such as downy mildew, soft rot, aphids, etc.) in complex field environments are verified, and data basis is provided for subsequent generation of unmanned aerial vehicle prescription maps based on diagnosis results and implementation of variable precision pesticide application.

[0146] (2) Experimental equipment and platform:

[0147] 1. UAV platform: DJI M350 RTK unmanned aerial vehicle with high-precision positioning and intelligent flight path control functions, supporting multi-spectral camera payloads, and capable of autonomous flight to complete field inspection tasks.

[0148] 2. Multispectral camera: MicaSense RedEdge-P multispectral imaging system, including 5 independent narrow bands (blue 475nm, green 560nm, red 668nm, red edge 717nm, near-infrared 842nm) and 1 panchromatic band (range 630-870nm), and a light sensor (DLS2) for real-time light correction.

[0149] 3. Edge computing platform: NVIDIA Jetson XavierNX embedded computing terminal, deploying MobileNetV3 deep learning model optimized by TensorRT, supporting local real-time inference and 5G data transmission.

[0150] 4. Expert system module: Remote image diagnosis platform based on WebSocket real-time communication protocol, with expert account management, image spectrum data browsing, and labeled information feedback functions.

[0151] 5. Processing software and algorithms: Image preprocessing uses Pix4D, QGIS, and OpenCV tools, model training uses PyTorch deep learning framework and TorchVision toolkit, and result visualization and evaluation uses TensorBoard and Matplotlib software.

[0152] (3) Experimental objects and sites:

[0153] The experimental site is selected in a cold and cool vegetable core planting demonstration area in Yongning County, Yinchuan City, Ningxia. Ten cabbage planting areas are selected as sample collection fields, with a total area of about 20 mu. The main cultivated varieties include "Zhonggan 21" and "Lufeng 2", covering different growth stages (seedling stage, rosette stage, and heart stage). The main diseases and pests in the field are downy mildew, soft rot, and aphids, with random and local aggregation distribution.

[0154] (4) Experimental procedure and steps

[0155] Firstly, according to the topographic features of the test area, the test field is divided into 10 columns, each column has 7 small plots, a total of 70 experimental plots. 12 ground control points (GCPs) are set in the area to correct the positioning error of unmanned aerial photography, so as to improve the accuracy of geographical registration of spectral image.

[0156] 1. Multispectral image acquisition: Use the unmanned aerial vehicle to take photos of the test area according to the predetermined route to obtain multispectral images. Use DJI Pilot to set the flight height to 15 m, the flight speed to 2 m / s, the forward overlap rate of the flight strip to 80%, and the lateral overlap rate to 70%, to ensure full coverage of the image of the operation area. During data collection, 5 spectral bands and 1 panchromatic band data of RedEdge-P camera are obtained, and real-time light intensity correction is carried out by using on-board DLS2 light sensor. Each field area collects no less than 300 images, covering various samples such as healthy plants, mild lesions, moderate lesions and severe lesions, to ensure that the data contains various degrees of occurrence of diseases and pests.

[0157] 2. Image preprocessing and spectral feature extraction: Use the radiometric calibration panel of multispectral image to carry out radiometric correction on the image, to eliminate the reflectivity difference caused by different flight times and light conditions. With the help of QGIS software, the multispectral images are spliced to generate regional orthographic image, and the GeoTIFF format data with geographical coordinates is exported. Then, the customized Python script is used to extract the region of interest (ROI) in the image, to obtain the multi-band spectral reflectance curve corresponding to each sample, and to arrange it into a 6-channel standardized tensor data. At the same time, the image data is normalized, denoised and missing spectrum interpolated as necessary, to finally form the input data set for training and sample metadata record table.

[0158] 3. Disease and pest model identification and confidence evaluation: A lightweight convolutional neural network model (MobileNetV3) is deployed on the Jetson NX edge terminal. The output layer is designed as a multi-task structure, containing three output branches: disease and pest species classification, disease severity level prediction, and confidence calculation (obtained through Softmax). In real-time field inference, the processing time of each image is controlled within 1 second. The model output includes the predicted disease and pest type (such as downy mildew), the severity level of infection (such as mild / moderate / severe), and the corresponding confidence value. The system judges the confidence of the model output: if the confidence is ≥0.7, the result is considered reliable, and the identification result and confidence are stored in the database and marked as "trusted diagnosis"; if 0.7≤C<0.9: the model has some uncertainty, and such medium confidence samples will trigger the expert review process (Expert System). If the confidence is <0.7, the result is considered uncertain, and the sample is placed in the "low confidence sample pool" for subsequent processing.

[0159] 4. Remote expert review: The system automatically starts the expert review process for low confidence samples. The edge terminal uploads the image data packet of the low confidence sample to the Web-based expert diagnosis platform through 5G communication, which includes multispectral fusion images, corresponding six-band spectral curves, and model preliminary diagnosis information. Online experts view the spectral images and image features of the samples through the platform's graphical interface, confirm the disease and pest species and infection level, and submit the final authoritative label information. The samples and their labels confirmed by manual annotation are synchronously returned and stored in the spectral database in real time. At this time, the sample state changes from "low confidence" to "verified", enters the high confidence sample library, and serves as effective new data for subsequent model training.

[0160] 5. Model incremental learning and redeployment: The system triggers the model fine-tuning process after accumulating a certain number of newly annotated samples (e.g., obtaining the latest 100 expert annotated samples). The new data and the original training data are fused and divided into training and validation sets in a 80:20 ratio. The model parameters are fine-tuned using a transfer learning strategy. Specifically, the CNN backbone network is frozen, only the parameters of the classification output layer and the confidence related layer are updated, the learning rate is set to 0.0001, the batch size is 32, and the iteration training is 15 rounds. After fine-tuning training, the accuracy of the new model is evaluated on the validation set. If it is significantly better than the original model, the updated model file is exported and deployed to all related edge terminals (unmanned aerial computing units or ground stations) wirelessly, replacing the old version of the model. Through the above mechanism, the model parameters are iteratively upgraded with the accumulation of new samples.

[0161] 6. Model closed-loop verification and evaluation:

[0162] a. Version control and traceability: After each model update, the system records and manages the model version, so as to compare the performance difference of different versions of the model and trace the evolution process of the model.

[0163] b. Performance index statistics: After the deployment of the new model, the same area is inspected again, and the changes in identification accuracy, average confidence and other indicators of the new and old models are compared; the trend of the proportion of samples entering expert review in each round of inspection over time and the decline curve of the frequency of expert intervention are counted, so as to quantify the process of the gradual maturity of the model.

[0164] c. Prevention and control effect evaluation: According to the model diagnosis result, the variable spraying suggestion is generated, which is compared with the traditional experience spraying scheme, the effect improvement index of pest control is calculated, for example, the reduction ratio of pesticide use amount, the improvement amplitude of prevention and control precision, and the economic benefits brought by it.

[0165] The experimental results of the embodiment show that: after three rounds of closed-loop learning iteration, the identification accuracy of the model for cabbage downy mildew is improved from 85% to more than 93%, the proportion of low confidence samples in the field is reduced from about 20% to 8%, and the frequency of expert manual review intervention is significantly reduced. The error of the variable spraying prescription generated based on the method of the application is reduced by about 15% compared with the traditional experience, and it is expected to reduce about 10%-15% of the pesticide use amount, and the efficiency and precision of pest control are obviously improved. The effectiveness and superiority of the closed-loop dynamic learning method described in the application in the actual field environment are verified.

[0166] The above-described embodiments only express the device layout method of the application, which is described in detail and specifically, but should not be understood as a limitation on the scope of the application; it should be pointed out that for ordinary skilled persons in the art, some adjustments and improvements can be made without departing from the concept of the application, which are all within the protection scope of the application; therefore, the protection scope of the application patent should be subject to the appended claims.

Claims

1. A method for dynamic evolution and update of a fruit and vegetable disease and pest spectrum database based on model confidence feedback driving, characterized in that, The method comprises the following steps: Step S10. Obtain the disease and pest diagnosis result and the corresponding model confidence; Step S20. Determine whether the model confidence corresponding to the disease and pest diagnosis result is not less than the high threshold of the model confidence based on the preset confidence model: If yes, output the recognition result and the image, and input them into the disease and pest spectrum database; If no, output the model confidence corresponding to the disease and pest diagnosis result to the expert system module; Step S30. Obtain the model confidence less than the high threshold of the model confidence, and determine whether the model confidence corresponding to the disease and pest diagnosis result is not less than the low threshold of the model confidence based on the expert system module: If yes, identify and confirm the diagnosis result and the corresponding model confidence by the expert, correct the diagnosis result, and input the corrected diagnosis result into the disease and pest spectrum database and the abnormal sample library; If no, identify the diagnosis result and the corresponding model confidence by the expert, output the correct diagnosis result, and input it into the disease and pest spectrum database.

2. The method of claim 1, wherein the method is characterized by, Step S10. Obtain a model confidence of the pest and disease, the model confidence is Conf total , and the model confidence Conf total is obtained by weighting the classification probability confidence Conf cls , the prediction uncertainty confidence Conf unc , and the embedding distance confidence Conf dist .

3. The method of claim 2, wherein the method is characterized by, the classification probability confidence wherein, represents the predicted probability of the model for the ith class, z i represents the logit value of the last layer of the model, C represents the total number of classes, T is a temperature parameter, and T > 0.

4. The method of claim 3, wherein, The prediction uncertainty confidence Conf unc = exp(-λσ 2 ); Wherein, is the average of all M predictions, M is the forward number, p (1) ,…,p (M) is the probability of the corresponding M group, σ is the standard deviation, and λ is a hyperparameter that controls the decay speed.

5. The method of claim 2, wherein the method is characterized by, The embedding distance confidence Conf dist = exp(-γD M (x)) Wherein, γ is a scaling hyperparameter, f(x) is an embedding representation of the input sample in the feature space, μ is the mean of a class of the training set, T is a temperature parameter, Σ is a covariance matrix of the features, and D is a Mahalanobis distance, which measures the standardized distance between the sample and the class center.

6. The method of claim 2, wherein the method is characterized by, The model confidences c = [Conf cls ,Conf unc ,Conf dist ] are combined into an input vector c, and the model confidence weights w are: w = Softmax(W ReLU(V c)) Wherein: V, W are weight matrices to be learned, ReLU is an activation nonlinearity, and Softmax makes the weights satisfy equal to 1.

7. The method of claim 2, wherein the method is characterized by, In step S10, the model confidence of the plant disease and pest is acquired; further comprising calibrating the model confidence of the plant disease and pest, and the calibration adopts entropy regularization and temperature calibration, wherein the entropy value L of the entropy regularization is: entropy L = -∑plogp wherein, is the predicted probability of class i, and C is the total number of classes.

8. The method of claim 7, wherein the method is characterized by, The temperature calibration calibration index ECE is: where K is the number of bins into which the prediction confidence interval is divided, B k is the set of samples in the kth bin, n is the total number of samples, acc(B k ) is the average accuracy of the samples in the bin, conf(B k ) is the average confidence of the samples in the bin.

Citation Information

Patent Citations

  • Fruit and vegetable video pest detection method based on deep learning and big data technology

    CN118570696A

Cited By

  • Facility agriculture disease and pest prevention and control method and device based on deep learning

    CN121837804A