Fruit and vegetable disease and insect pest spectrum database dynamic evolution updating and diagnosis optimization method based on model confidence feedback driving
By using a model confidence feedback-driven approach, combined with an expert system to dynamically update and optimize the spectral database of fruit and vegetable diseases and pests, the problem of insufficient accuracy in disease and pest detection has been solved, achieving high efficiency and accuracy in disease and pest identification.
Patent Information
- Application Number
- CN202510897415.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-11-07
AI Technical Summary
In the current technology, the methods for detecting diseases and pests in fruits and vegetables cannot be updated in a timely manner, which makes it difficult to guarantee the accuracy of disease and pest detection and easily leads to errors.
A dynamic evolution update method for the spectral database of fruit and vegetable diseases and pests based on model confidence feedback is adopted. By acquiring spectral data of fruit and vegetable crops, the disease diagnosis model outputs the diagnosis results of diseases and pests and their confidence levels, and the expert system is used for calibration and correction. A data update and model optimization mechanism for collaborative work between experts and models is established to realize the dynamic evolution update of the database.
It improves the accuracy and adaptability of pest and disease detection, prevents the accumulation of misjudgments caused by model overconfidence, ensures that highly reliable results are entered into the database, and uncertain results are corrected by experts, thus realizing the continuous optimization and upgrading of the fruit and vegetable pest and disease spectral database.
Smart Images

Figure CN120913002A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of agricultural remote sensing and artificial intelligence, and particularly relates to a fruit and vegetable disease and pest spectrum database dynamic evolution updating and diagnosis optimization method based on model confidence feedback driving. BACKGROUND
[0002] Timely and accurate diagnosis of fruit and vegetable diseases and pests is crucial for agricultural production. Traditional methods often rely on manual observation and experience, which is low in efficiency and easily affected by subjective factors.
[0003] In the prior art, for example, a Chinese invention patent with application number CN202410635672.3 discloses a fruit and vegetable video disease and pest detection method based on deep learning and big data technology, which specifically discloses a camera module, a deep learning model training module, a video detection display module, a detection data collection module, a real-time detection module and an alarm module, an offline detection module and a display module. The camera module uses a camera installed in a greenhouse to shoot a set of range-different fruit and vegetable growth videos every 24 hours. The deep learning model training module includes an iteration unit and an identification unit. The video detection display module is an interface generated by software development, which displays the fruit and vegetable disease and pest detection status in real time. The real-time detection module processes and analyzes the detection data in real time, and triggers the alarm module when a certain index is reached. The offline detection module processes and analyzes the detection data offline, and finally displays the results in the display module.
[0004] However, the accuracy of disease and pest detection cannot be known in the above-mentioned scheme, and new diseases and pests or disease and pest variations may occur. The above-mentioned method cannot accurately determine the diseases and pests, and the diseases and pests cannot be updated in time. SUMMARY
[0005] Therefore, it is necessary to verify the accuracy of disease and pest detection, which may easily lead to errors in disease and pest detection. A fruit and vegetable disease and pest spectrum database dynamic evolution updating and diagnosis optimization method based on model confidence feedback driving is provided.
[0006] To achieve the above-mentioned purpose, the present application adopts the following scheme:
[0007] A fruit and vegetable disease and pest spectrum database dynamic evolution updating and diagnosis optimization method based on model confidence feedback driving includes the following steps:
[0008] Step S10. Obtain the spectrum data of fruit and vegetable crops.
[0009] Step S20. Based on the disease diagnosis model, output the disease and pest diagnosis result and the corresponding model confidence.
[0010] Step S30. Obtain the model confidence of the pest diagnosis, execute the calibration module, and output the model confidence of the calibrated pest;
[0011] Step S40. Determine whether the model confidence of the pest is not less than a high threshold of the model confidence based on the preset confidence model:
[0012] If yes, output the recognition result and the image, and input the pest spectrum database;
[0013] If no, output the model confidence corresponding to the pest diagnosis result to the expert system module;
[0014] Step S50. Obtain the model confidence less than the high threshold of the model confidence, determine whether the model confidence corresponding to the pest diagnosis result is not less than a low threshold of the model confidence based on the expert system module:
[0015] If yes, identify and confirm the diagnosis result and the corresponding model confidence by the expert, correct the diagnosis result, and input the corrected diagnosis result into the pest spectrum database and the abnormal sample library;
[0016] If no, identify the diagnosis result and the corresponding model confidence by the expert, output the correct diagnosis result, and input it into the pest spectrum database;
[0017] Step S60. Obtain the number of newly added expert samples, the low confidence proportion, and the interval from the last training time, execute the incremental training module, and output the instruction for updating and optimizing the pest spectrum database and the abnormal sample library.
[0018] Preferably, in the step S30 of obtaining the model confidence of the pest diagnosis, the model confidence is Conf total , and the model confidence Conf total is obtained by weighting a classification probability confidence Conf cls , a prediction uncertainty confidence Conf unc , and an embedding distance confidence Conf dist .
[0019] Preferably, the classification probability confidence
[0020]
[0021] wherein, represents the prediction probability of the model for the i-th class, z i represents the logit value of the last layer of the model, C represents the total number of categories, and T represents a temperature parameter, and T>0.
[0022] Preferably, the prediction uncertainty confidence Conf unc= exp(-λσ 2 ); wherein,
[0023]
[0024] is the average of all M predictions, M is the forward number, p (1) ,…,p (M) is the probability of the corresponding M groups, σ is the standard deviation, and λ is a hyperparameter that controls the decay speed.
[0025] Preferably, the embedding distance confidence Conf dist = exp(-γD M (x)); wherein,
[0026]
[0027] γ is a scaling hyperparameter, f(x) is the embedding representation of the input sample in the feature space, μ is the mean of a class of the training set, T is a temperature parameter, Σ is the covariance matrix of the features, D is the Mahalanobis distance, which measures the standardized distance between the sample and the class center.
[0028] Preferably, the model confidence c = [Conf cls , Conf unc , Conf dist ] combines the three kinds of confidence into an input vector c, and the model confidence weight w is:
[0029] w = Softmax(W · ReLU(V · c))
[0030] wherein, V, W are weight matrices to be learned, ReLU is an activation nonlinearity, and Softmax makes the weights satisfy equal to 1.
[0031] Preferably, in step S30, a calibration module is executed; the calibration module stores an entropy regularization and temperature calibration, wherein the entropy value L entropy of the entropy regularization is:
[0032]
[0033] wherein, is the prediction probability of the i-th class, and C is the total number of categories.
[0034] Preferably, the entropy regularization further includes a loss function additional term:
[0035] L total = L task + β · L entropy
[0036] L totalFor the total loss function, L task For mission losses, L entropy β is the entropy regularization term, and β is the weight.
[0037] Hyperparameters.
[0038] Preferably, the calibration index ECE for the temperature calibration is:
[0039]
[0040] Where K represents the number of bins into which the prediction confidence interval is divided, and B represents the number of bins into which the prediction confidence interval is divided. k Let be the sample set in the k-th bin, n be the total number of samples, and acc(B) be the total number of samples. k ) represents the average accuracy of the samples in this bin, conf(B) k ) represents the average confidence level of the samples in this bin.
[0041] Preferably, in step S60, incremental training of the model stored in the incremental training module is performed, and the incremental training is triggered as follows:
[0042] Trigger = (N exp ≥N th )or(R low >R th or (Δt≥T) cycle )
[0043] Where, N exp : Added number of expert-annotated samples, N th Expert sample trigger threshold, R low :
[0044] Current low confidence sample proportion, R th : The threshold for the proportion of low-confidence samples, Δt: the time elapsed since the last model training, T cycle Training time period threshold.
[0045] The technical solution adopted in this application can achieve the following beneficial effects:
[0046] 1. By calculating and comparing the model confidence scores of the output diagnostic results, the diagnostic results are given.
[0047] The model confidence score of the results is used to address the problem of errors in pest and disease detection that occur when only pest and disease detection is performed without verifying the accuracy of the detection.
[0048] 2. By using entropy regularization and temperature scaling, the model confidence calculation becomes more reliable and is not affected by...
[0049] The model may be overfitting or the sample distribution may be biased, resulting in an inflated confidence score.
[0050] 3. Through the introduction of model confidence fusion calculation and hierarchical processing in the diagnosis process, the data update and model optimization mechanism of expert and model collaborative work is established. This method effectively prevents the accumulation of misjudgment caused by model overconfidence, and ensures that only high reliability results are automatically entered into the database, and uncertain results are modified by experts. With the help of expert annotation feedback and model incremental self-learning, the dynamic evolution and update of the fruit and vegetable disease and pest spectrum database are realized; and the continuous optimization and upgrading of the diagnosis model is realized, which is of great significance to ensure the accuracy and adaptability of disease and pest identification in agricultural production. BRIEF DESCRIPTION OF DRAWINGS
[0051] Fig. 1 A flowchart of a fruit and vegetable disease and pest spectrum database dynamic evolution and update method based on model confidence feedback driving disclosed by an embodiment of the present application.
[0052] Fig. 2 A category and temperature statistical chart of a fruit and vegetable disease and pest spectrum database dynamic evolution and update method based on model confidence feedback driving disclosed by an embodiment of the present application.
[0053] Fig. 3 A module flowchart of a fruit and vegetable disease and pest spectrum database dynamic evolution and update method based on model confidence feedback driving disclosed by an embodiment of the present application. DETAILED DESCRIPTION
[0054] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings. The preferred embodiments of the present application are shown in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0055] It should be noted that when a device is considered to be "connected" to another device, it can be directly connected to the other device or there can be a mediating device present at the same time. The terms "internal", "top", "upper", "lower", "up", "down" and similar expressions used herein are for illustrative purposes only and do not represent the only implementation.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application herein are only for the purpose of describing the specific embodiments and are not intended to limit the present application. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0057] Reference Figs. 1-3The application provides a fruit and vegetable disease and pest spectrum database dynamic evolution update and diagnosis optimization method based on model confidence feedback driving, comprising the following steps:
[0058] Step S10. Obtain the spectrum data of the fruit and vegetable crops;
[0059] Step S20. Based on the disease diagnosis model, output the disease and pest diagnosis result and the corresponding model confidence;
[0060] Step S30. Obtain the model confidence of the disease and pest diagnosis, execute the calibration module, and output the calibrated model confidence of the disease and pest;
[0061] Step S40. Based on the pre-set confidence model, determine whether the model confidence of the disease and pest is not less than the high threshold of the model confidence:
[0062] If yes, output the recognition result and the image, and input the disease and pest spectrum database;
[0063] If no, output the model confidence corresponding to the disease and pest diagnosis result to the expert system module;
[0064] Step S50. Obtain the model confidence less than the high threshold of the model confidence, based on the expert system module, determine whether the model confidence corresponding to the disease and pest diagnosis result is not less than the low threshold of the model confidence:
[0065] If yes, the expert identifies and confirms the diagnosis result and the corresponding model confidence, corrects the diagnosis result, and inputs the corrected diagnosis result into the disease and pest spectrum database and the abnormal sample library;
[0066] If no, the expert identifies the diagnosis result and the corresponding model confidence, outputs the correct diagnosis result, and inputs it into the disease and pest spectrum database;
[0067] Step S60. Obtain the number of newly added expert samples, the low confidence proportion, and the interval from the last training time, execute the incremental training module, and output the instruction for updating and optimizing the disease and pest spectrum database and the abnormal sample library.
[0068] Specifically, when the spectrum data of the fruit and vegetable crops is obtained, DJIM350RTK unmanned aerial vehicle is used in the front end of the system to carry RedEdge-P multispectral camera, and a GPS high-precision positioning module is integrated. The platform can automatically cruise in the field and obtain crop image information, and the collection range covers blue light, green light, red light, red edge, near-infrared and other key vegetation bands, providing a multi-dimensional spectrum data basis for subsequent diagnosis.
[0069] After obtaining the spectral data, the system automatically performs radiation correction, geometric registration, illumination normalization, distortion correction and other image preprocessing operations, outputs standardized multi-band image data, and ensures the spectral consistency and spatial alignment accuracy of the model input.
[0070] The pest diagnosis model is deployed in the unmanned aerial vehicle edge computing terminal or ground station, adopts a lightweight convolutional neural network (CNN) architecture, performs pest species identification and grade judgment on the preprocessed images, and outputs the fused comprehensive model confidence. In the preset confidence model, if the model confidence is ≥0.9, it means that the model prediction is highly reliable; if 0.7≤model confidence<0.9, it means that the model has certain uncertainty; and if the model confidence is <0.7, it means that the model confidence is too low.
[0071] Further, first, a multi-spectral imaging device is used to obtain high-resolution multi-spectral image data by aerial inspection of the fruit and vegetable planting area. Preferably, the unmanned aerial vehicle used, such as DJI M350 RTK, is equipped with a RedEdge-P multi-spectral camera to obtain image information of multiple bands such as visible light and near-infrared. The collected spectral images are corrected and optimized through a preprocessing step, including but not limited to spectral calibration, noise removal, image registration, and region of interest extraction, etc., to provide high-quality input data for subsequent model analysis.
[0072] Then, the preprocessed spectral data is input into the pre-trained pest diagnosis model for inference. The model can use deep learning algorithms such as convolutional neural networks, which have been trained using a large number of labeled samples in the initial spectral database, and can output pest species diagnosis results for input samples and model confidence evaluation of the results. The output obtained by model inference includes: first, the diagnosis result (such as determining the type of crop disease or pest species), and second, the model confidence value, which represents the credibility of the model to the diagnosis result. The worker confirms the accuracy of his judgment through the diagnosis result and the corresponding model confidence.
[0073] After the pest diagnosis result and the corresponding model confidence are obtained, the pest diagnosis result and the corresponding model confidence are compared with the preset confidence model. If the model confidence is greater than or equal to 0.9, it is indicated that the model prediction is highly reliable, no manual intervention is needed, and the recognition result and the image are directly written into the pest spectrum database, and the diagnosis result and the corresponding model confidence are fed back to the staff (through a display screen, a wireless network, Bluetooth or the like, as long as the staff can see it). If the model confidence is less than 0.9, the model confidence corresponding to the pest diagnosis result is output to the expert system module. The expert system module is manually labeled and classified by a pest diagnosis expert according to the image and spectrum data to ensure the accuracy of the recognition result and the robustness of the system. The label data output by the expert system module is written into the system database as a new training sample and is used for online fine-tuning, migration training or incremental learning of the subsequent model to continuously improve the recognition performance and adaptability of the model in the actual environment, forming a complete learning closed loop.
[0074] The technical scheme of the fruit and vegetable pest spectrum database dynamic evolution update and diagnosis optimization method based on model confidence feedback driving adopted in the application can achieve the following beneficial effects:
[0075] 1. The model confidence of the output diagnosis result is calculated and compared to give the model confidence of the diagnosis result, solving the problem of incorrect pest detection due to the lack of verification of the accuracy of pest detection.
[0076] 2. The model confidence calculation is more reliable through entropy regularization and temperature scaling, and will not produce a high confidence score due to model overfitting or sample distribution deviation.
[0077] 3. The model confidence fusion calculation and hierarchical processing are introduced in the diagnosis process to establish a data update and model optimization mechanism for experts and models to work together. This method effectively prevents misjudgment accumulation caused by excessive confidence of the model, ensures that only high-reliability results are automatically entered into the database, and uncertain results are modified by experts. With the help of expert annotation feedback and model incremental self-learning, the dynamic evolution update of the fruit and vegetable pest spectrum database is realized, and the continuous optimization and upgrading of the diagnosis model are realized, which is of great significance to ensure the accuracy and adaptability of pest identification in agricultural production.
[0078] In the above scheme, step S30 obtains the model confidence of the pest, and the model confidence is Conf total , and the model confidence Conf total is obtained by weighting the classification probability confidence Conf cls , the prediction uncertainty confidence Conf unc and the embedding distance confidence Conf dist .
[0079] Specifically, the classification probability confidence
[0080]
[0081] where, represents the prediction probability of the i-th class by the model, z i represents the logit value of the last layer of the model, C represents the total number of classes, and T is a temperature parameter (T > 0) that adjusts the smoothness of the softmax. When T = 1, it is a standard softmax; when T > 1, the output is more gentle (the distribution is more “soft”); when T < 1, the output is more sharp (the probability of a certain class is very high, and the others are almost zero). By adjusting T, the softmax output probability is more matched with the actual accuracy, and the confidence calibration ability is improved.
[0082] Take the maximum probability:
[0083]
[0084] represents the class with the maximum probability in the softmax output as the classification confidence of the sample; the closer the probability is to 1, the more confident the model is in the class. For example, Fig. 2 where: Class1, Class2, Class3 represent three different class labels output by the model; T = 0.5: the distribution is extremely sharp, and the model is very “confident” and basically only selects one class; T = 1.0: standard Softmax; T = 2.0: the distribution is more gentle, and the confidence decreases; T = 5.0: close to average distribution, indicating that the model is very uncertain.
[0085] Further, the prediction uncertainty confidence Conf unc = exp(-λσ 2 ); where,
[0086]
[0087] is the average value of all M predictions, M is the number of forward times, p (1) ,…, p (M) is the probability corresponding to the M groups, σ is the standard deviation, and λ is a hyperparameter that controls the decay speed.
[0088] Further, the embedding distance confidence Conf dist = exp(-γD M (x)) where,
[0089]
[0090] where g is a scaling hyper-parameter, f(x) is the embedding representation of the input sample in the feature space, m is the mean of a class of the training set, T is a temperature parameter, S is the covariance matrix of the features, D is the Mahalanobis distance, measuring the standardized distance between the sample and the class center.
[0091] The model confidence c = [Conf cls , Conf unc , Conf dist ] combines three kinds of confidence into an input vector c, and the model confidence weight w is:
[0092] w = Softmax(W ReLU (V c))
[0093] where: V, W are weight matrices to be learned, ReLU is an activation nonlinearity, and Softmax makes the weights satisfy equal to 1.
[0094] Model confidence calculation method: In this embodiment, the calculation of the model confidence C integrates information in three dimensions to improve the accuracy of the credibility evaluation. Specifically, C is determined based on the following three aspects of measurement:
[0095] 1) Classification probability: the Softmax maximum probability value P_max output by the model, representing the intuitive confidence degree of the model for the predicted pest and disease class.
[0096] 2) Prediction uncertainty: an uncertainty quantitative indicator is calculated through the prediction output of the model, for example, using Monte Carlo dropout multiple forward reasoning or the standard deviation s based on the prediction results of the ensemble model, to measure the stability of the model prediction for the sample. The larger the standard deviation, the higher the uncertainty and the lower the confidence.
[0097] 3) Embedding feature distance: input the sample to be tested into the model to extract its representation vector in the embedding feature space, compare the vector with the feature vectors of known samples in the model training set, calculate the nearest neighbor distance D, the greater the distance, the higher the degree of deviation of the sample from the distribution of the training data (i.e. it may be a novel or abnormal sample), and the lower the confidence should be. Normalize the three indicators P_max, σ and D obtained above, and input them into a dynamic weighted fusion model to calculate the final confidence C. Preferably, the fusion model can use a trained multi-layer perceptron (MLP) network or a pre-set weighted formula to adaptively adjust the weights of each indicator according to historical validation data, so as to accurately evaluate the model confidence in different situations. For example, if the Softmax probability of a certain type of disease is always high, resulting in overconfidence, the fusion model will appropriately reduce the weight of P_max in the comprehensive confidence, and enhance the consideration of uncertainty σ and feature distance D. Through this dynamic weight confidence fusion method, the reliability of the model's judgment on the current sample can be more comprehensively described.
[0098] In addition, in order to prevent the model from making overconfident judgments, confidence calibration and regularization measures are introduced during model training and inference. On the one hand, an entropy regularization term is added in the model training stage to encourage the model output to maintain appropriate information entropy, thereby avoiding the distortion of confidence evaluation caused by excessively extreme output probability; on the other hand, temperature scaling calibration technology is applied to the output probability in the model deployment and inference stage, that is, a temperature coefficient is introduced during Softmax calculation to stretch or compress the probability distribution (for example, a temperature T>1 is introduced to reduce the maximum probability value and increase the entropy), so that the confidence of the model output is better matched with the actual accuracy. Through entropy regularization and temperature scaling, the model confidence calculation is more reliable and will not produce a high confidence score due to model overfitting or sample distribution deviation.
[0099] In an embodiment of the present application, in step S30, a calibration module is executed; the calibration module stores entropy regularization and temperature calibration, wherein the entropy value L of the entropy regularization entropy is:
[0100]
[0101] wherein, where p is the predicted probability of the i-th class, C is the total number of classes. Entropy measures the uncertainty or disorder of a probability distribution. When the softmax output is very "confident" (e.g., [0.99, 0.005, 0.005]), the entropy is very low; when the softmax output is very "ambiguous" (e.g., [0.33, 0.33, 0.34]), the entropy is very high; thus, by increasing the entropy, the model can be encouraged to be cautious on uncertain samples and reduce overconfidence. Prevent overfitting: especially when the data class is unbalanced or the sample quality is poor; Improve robustness: the output distribution is more conservative and does not give too high confidence easily; As an additional item of the loss function: often written as
[0102] L total = L task + β·L entropy
[0103] L total : Total Loss, the objective function that the model is ultimately optimized for;
[0104] L task : TaskLoss, the basic loss of the core task;
[0105] L entropy : Entropy Regularization, which punishes overly confident outputs and improves the model's ability to express uncertainty;
[0106] β: is a weight hyperparameter that controls the influence of the entropy regularization in the total loss.
[0107] In the above scheme, the temperature calibration calibration index ECE is:
[0108]
[0109] where K is the number of bins into which the prediction confidence interval is divided, B k is the sample set in the k-th bin, n is the total number of samples, acc(B k ) is the average accuracy of samples in the bin, and conf(B k ) is the average confidence of samples in the bin.
[0110] The generation and training of the above-mentioned embedded diagnosis model includes the following steps: a training target function
[0111]
[0112] y k : true label, a "switch value", indicating that the sample belongs to the k-th class when it is 1, (for example, the probability that the model considers it to be "downy mildew" is 0.8).
[0113] First term (cross-entropy): measures the error between the predicted result and the true class, the smaller the value, the more accurate the prediction.
[0114] Second term (entropy regularization): entropy measures the "uniformity of the predicted distribution", adding this term is to avoid the model being too "confident" to choose only one class; if the model always outputs [1.0, 0.0, 0.0], it is easy to overfit; encourage output like [0.7, 0.2, 0.1] to keep the probability of uncertainty.
[0115] λ1: adjust the size of this term on the total loss.
[0116] Third term (confidence supervision):
[0117] C0: the confidence value originally calculated by the model;
[0118] C ★ : "ideal confidence" based on expert opinion or other methods;
[0119] If the model's confidence deviates too much from the actual credibility, there will be a larger penalty; λ2: adjust the proportion of this penalty in the total loss.
[0120]
[0121] b. Temperature scaling calibration (improve confidence reliability)
[0122] Find a single scalar T on the validation set * Minimize negative log-likelihood (NLL), calibrated probability is
[0123]
[0124] z k : the "score" output by the last layer of the model, which has not yet become a probability value;
[0125] T * : temperature parameter, adjust the "confidence level":
[0126] If T * >1, the predicted probability is more "smooth", and the confidence is more conservative;
[0127] If T ★ <1, the prediction result is more "extreme";
[0128] exp: exponential function, Softmax is a standard probability transformation function;
[0129] Explanation of the whole paragraph: this step is "temperature scaling", which makes the model's prediction result and the actual accuracy more matched (such as avoiding the model outputting a high confidence of 0.99 but actually often wrong).
[0130] c. Incremental fine-tuning and model version management trigger condition: new expert samples ≥200 or online running >14 days;
[0131] Fine-tuning strategy: freeze the first two feature layers, only update the subsequent layers, 20 epochs are converged; model hash record: facilitate OTA rollback and audit.
[0132] In another embodiment of the present application, the confidence-driven data update mechanism is as follows
[0133]
[0134] Wherein, the initial threshold: T H =0.9, T L =0.7; 90% / 70% quantile is recalculated every 1k samples, which realizes self-adaptation. The incremental training of the model is stored in the pest spectrum database and the abnormal sample library, and the trigger of the incremental training is:
[0135] Trigger=(N exp ≥N th )or(R low >R th )or(Δt≥T cycle )
[0136] Wherein, N exp : the number of newly added expert labeled samples, N th : expert sample trigger threshold, R low : current low confidence sample proportion, R th : the proportion threshold of low confidence samples, Δt: the time experienced since the last model training, T cycle : training time period threshold.
[0137] Data sampling weight, expert sample: weight w=1; high confidence sample: weight w=Conf total ; if the online average confidence decreases by more than 3% within 24 hours, automatically roll back to the last stable model version.
[0138] Confidence-driven data updating mechanism: after the model completes the diagnosis of the new sample and calculates the confidence, the application processes the sample differently according to the level of the confidence, and constructs a dynamic database updating and model optimization cycle. Preferably, the confidence is divided into high, medium and low three levels: when the model confidence C≥0.9, it is considered that the model is very reliable for the result judgment, and the diagnosis result of the model can be directly adopted; when 0.7≤model confidence C<0.9, it is considered that the model is of medium reliability, and the reliability of the result is doubtful; when the model confidence C<0.7, it is considered that the model is of low confidence, and the result is likely to be inaccurate or an abnormal case that the model has not seen before.
[0139] The system adopts the following processing strategies according to the above confidence levels: for samples with high confidence (model confidence≥0.9), the model-predicted pest and disease species result is directly regarded as valid information, and the sample and its model-predicted label are recorded into the fruit and vegetable pest and disease spectrum database as new data for storage; for samples with medium confidence (0.7-0.9), they are marked as to be audited and submitted to experts for review and judgment; for samples with low confidence (<0.7), they are marked as abnormal to be handled and also submitted to experts for identification. Through the division of the three-level confidence interval, different treatments can be performed according to the reliability of the results, which not only ensures the timely use of high-confidence results, but also ensures that unreliable results are intervened by humans and do not mislead the database with incorrect information.
[0140] After the submission of the medium and low confidence samples, the system sends the related spectrum data and model preliminary diagnosis results to the remote expert diagnosis platform through the network, and the artificial experts with plant pest and disease diagnosis experience identify and label these samples. The experts view the spectrum images of the samples and the preliminary conclusions given by the model, combine their own knowledge to judge the actual pest and disease situation, and give the authoritative diagnosis result label. For medium confidence samples, the model may be partially correct but not certain enough, and the experts focus on verifying the features that the model is easily confused, and confirm and correct the label; for low confidence samples, the model is almost unable to give meaningful judgment, and the experts will conduct complete artificial diagnosis on these possibly new or abnormal cases. Once the experts complete the labeling, the system archives the sample data with the confirmed label of the experts to the abnormal sample library. The abnormal sample library is used to store the cases that the model has been uncertain or misjudged before, and provides targeted training data for subsequent model improvement. In addition, all the samples labeled by the experts (including medium and low confidence categories) are added to the main spectrum database together with their correct labels, as new training samples for storage. At the same time, for the high confidence samples that have been automatically added to the database, the system also makes corresponding records (such as marking that the label comes from model prediction rather than manual), for reference during training.
[0141] With the passage of time, the newly added samples (including high-confidence automatic addition samples and expert-annotated addition samples) accumulated in the spectrum database are increasing. In order to make the model learn these new data in time, the present application sets up a dynamic triggering mechanism for model incremental training. The triggering conditions can be set in advance, for example, when the number of newly added samples in the spectrum database reaches a certain threshold, or when the new data is checked periodically (such as every few days / weeks), the update training of the model is triggered. During the incremental training, the above-mentioned newly added expert-annotated samples and high-confidence samples are added to the model training set, and the original diagnostic model is retrained or fine-tuned online. By introducing expert-annotated data, the model can correct previous errors or uncertain recognition, and significantly improve the accuracy on the corresponding category; and adding high-confidence model prediction samples (i.e. the model's confident correct guess) is equivalent to using reliable data generated by the model itself to further enrich the training set, which helps to consolidate the model's grasp of existing knowledge. During the training process, the entropy regularization method and other methods for improving the generalization ability are continued to be applied, and the temperature scaling method and other means are continued to be used to calibrate the confidence of the new output of the model after training. After completing the update training, the new model is deployed to the embedded diagnostic system in the front end for subsequent field inspection diagnosis, thereby entering the next round of operation cycle. In this way, the model and the database are continuously evolved: the model becomes more robust and can identify more diverse pest and disease spectra; the spectrum database is also continuously expanded with new instances, covering more comprehensive pest and disease spectral characteristics. The whole system forms a closed loop through the model confidence feedback mechanism, and autonomously learns and improves in practical application, significantly improving the efficiency and accuracy of fruit and vegetable pest and disease diagnosis.
[0142] In summary, the present application establishes a data update and model optimization mechanism for experts and models to work together by introducing model confidence fusion calculation and hierarchical processing in the diagnosis process. This method effectively prevents the accumulation of misjudgments caused by model overconfidence, ensuring that only high-reliability results are automatically entered into the database, and uncertain results are all corrected by experts. With the help of expert annotation feedback and model incremental self-learning, the present application realizes the dynamic evolution and update of the fruit and vegetable pest and disease spectrum database, as well as the continuous optimization and upgrading of the diagnosis model, which is of great significance for ensuring the accuracy and adaptability of pest and disease identification in agricultural production.
[0143] In the optimization of multi-crop and multi-region deployment of pest diagnosis system, an initial spectrum database covering common pests of multiple fruit and vegetable crops (such as tomatoes, cabbages, eggplants, etc.) is established, and a corresponding multi-task pest recognition model is trained on the cloud. The model is deployed in multiple plant protection unmanned aerial vehicle terminals to form a cooperative unmanned aerial vehicle monitoring queue. In the running process, each unmanned aerial vehicle autonomously patrols the field area it is responsible for regularly, collects crop spectrum data and performs pest identification and diagnosis. When a new type of pest or an irregular symptom appears on the crop in a certain area, causing the model to be unable to give a high confidence judgment (confidence is lower than the preset threshold), the system will submit these difficult samples to the expert diagnosis platform on the cloud for analysis and confirmation. For example, during the inspection of a tomato planting area, suspected new leaf mold patches were found. Since the symptoms are beyond the recognition range of the existing model, the model only gives a low confidence. At this time, the system automatically reports the sample, and the plant pathologist confirms it as a new leaf spot disease through remote diagnosis. The system then expands the sample data and expert confirmation label into the spectrum database, and the cloud server performs incremental training on the pest recognition model. The updated new model is distributed to all related unmanned aerial vehicle terminals to replace the old model. After several such cyclic iterations, the spectrum database of the entire system is continuously enriched, and the recognition ability of the model for various pests is simultaneously improved. Even in the face of new pests and diseases appearing in different regions and different seasons, the system can quickly learn and adapt, achieving continuous and reliable pest and disease diagnosis optimization in large-scale cross-regional deployment.
[0144] Through experiments and cross-regional joint monitoring for a year, the system has found and included 7 new pest types of different crops. Each time a new pest is confirmed by experts, the model is updated and distributed within an average of 24 hours, so that subsequent inspections can identify these new diseases in a timely manner. After several rounds of iterative optimization, the overall recognition accuracy of the model in each unmanned aerial vehicle terminal has increased by about 12% compared to the initial deployment, effectively enhancing the pest and disease recognition performance of the system in different regional environments. This embodiment shows that the method of the present application has good scalability and universality, and can be widely applied to large-scale, multi-region agricultural pest and disease intelligent monitoring networks to ensure the healthy growth of crops in various regions.
[0145] It should be noted that the specific hardware platforms mentioned in the above embodiments (such as the model of the unmanned aerial vehicle, the model of the multispectral camera, etc.) and the parameter settings (such as the confidence threshold of 0.7, the learning rate of 0.0001, etc.) are only used to illustrate the feasible implementation of the method of the present application in practical application, and are not a limitation on the scope of protection of the present application. Those skilled in the art can select different equipment platforms (for example, use other models of unmanned aerial vehicles or spectral sensors) and adjust threshold parameters and algorithm hyperparameters according to actual needs under the guidance of the principles of the present application. As long as the same function and effect as the basic idea of the present application are achieved, it should be considered to fall within the scope of protection of the present application.
[0146] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below in conjunction with the embodiments.
[0147] Embodiment:
[0148] (1) Purpose of the experiment: The purpose of this experiment is to construct a closed-loop cabbage disease and pest identification system with "data collection - model identification - confidence feedback - expert review - database expansion - model evolution" as the core. Based on multispectral remote sensing technology and deep learning methods, the recognition accuracy of the system for main diseases and pests of cabbage (such as downy mildew, soft rot, aphids, etc.) in the field complex environment, the model self-adaptive evolution ability and the deployment feasibility are verified, and the data basis for subsequent generation of unmanned aerial vehicle prescription map according to the diagnosis results and implementation of variable precision pesticide application is provided.
[0149] (2) Experimental equipment and platform:
[0150] 1. Unmanned aerial vehicle platform: DJI M350 RTK unmanned aerial vehicle with high-precision positioning and intelligent flight path control functions, supporting multispectral camera load, capable of autonomous flight to complete the task of field inspection.
[0151] 2. Multispectral camera: MicaSense RedEdge-P multispectral imaging system, including 5 independent narrow bands (blue 475nm, green 560nm, red 668nm, red edge 717nm, near infrared 842nm) and 1 panchromatic band (range 630-870nm), and a light sensor (DLS2) for real-time light correction.
[0152] 3. Edge computing platform: NVIDIA Jetson XavierNX embedded computing terminal, deploying MobileNetV3 deep learning model optimized by TensorRT, supporting local real-time inference and 5G data transmission.
[0153] 4. Expert system module: Remote image diagnosis platform based on WebSocket real-time communication protocol, with expert account management, image spectrum data browsing and labeled information return function.
[0154] 5. Processing software and algorithms: Image preprocessing uses Pix4D, QGIS and OpenCV tools, model training uses PyTorch deep learning framework and TorchVision toolkit, and result visualization and evaluation uses TensorBoard and Matplotlib software.
[0155] (3) Experimental objects and sites:
[0156] The experimental site is selected in a cool vegetable core planting demonstration area in Yongning County, Yinchuan City, Ningxia. Ten cabbage planting areas are selected as sample collection fields, with a total area of about 20 mu. The main cultivated varieties include "Zhonggan 21" and "Lufeng 2", covering different growth stages (seedling stage, rosette stage, and heart stage). The main diseases and pests are downy mildew, soft rot and aphids, which are randomly distributed and locally aggregated.
[0157] (4) Experimental process and steps
[0158] Firstly, according to the topographic features of the test area, the test field is divided into 10 columns, each column has 7 small areas, a total of 70 experimental small areas. Twelve ground control points (GCP) are set up in the area to correct the positioning error of unmanned aerial vehicle aerial photography, so as to improve the accuracy of geographical registration of spectral image.
[0159] 1. Multispectral image acquisition: Use unmanned aerial vehicle to take aerial photography according to the predetermined route to obtain multispectral image. Use DJIPilot to set flight height 15m, flight speed 2m / s, forward overlap rate 80%, lateral overlap rate 70%, to ensure full coverage of the image of the operation area. During data acquisition, 5 spectral bands and 1 panchromatic band data of RedEdge-P camera are obtained, and real-time light intensity correction is carried out by using airborne DLS2 light sensor. Each field area collects not less than 300 images, covering healthy plants, mild disease, moderate disease and severe disease, etc. Various samples to ensure that the data contains multiple disease and pest occurrence degrees.
[0160] 2. Image pre-processing and spectral feature extraction: Radiometric calibration of the image is performed using the radiometric calibration panel of the multispectral image to eliminate the reflectance differences caused by different flight times and lighting conditions. The multispectral image is stitched to generate a regional orthographic image map using QGIS software, and the GeoTIFF format data with geographic coordinates is exported. Then, the region of interest (ROI) in the image is extracted using a customized Python script, the multi-band spectral reflectance curve corresponding to each sample is obtained, and it is arranged into a 6-channel standardized tensor data. At the same time, the image data is normalized, denoised and missing spectrum interpolated and completed as necessary, and finally the input data set for training and sample metadata record table are formed.
[0161] 3. Disease and pest model identification and confidence evaluation: A lightweight convolutional neural network model (MobileNetV3) is deployed on the Jetson NX edge terminal, and the output layer is designed as a multi-task structure, including three output branches: disease and pest species classification, disease severity level prediction and confidence calculation (obtained by Softmax). In real-time field inference, the processing time of each image is controlled within 1 second. The model output includes the predicted disease and pest type (such as downy mildew), the infection severity level (such as mild / medium / severe), and the corresponding confidence value. The system judges the confidence of the model output: if the confidence is ≥0.7, the result is considered reliable, and the identification result and confidence are stored in the database and marked as "trusted diagnosis"; if 0.7≤C<0.9: the model has some uncertainty, and such medium confidence samples will trigger the expert review process (Expert System). If the confidence is <0.7, the result is considered uncertain, and the sample is put into the "low confidence sample pool" for subsequent processing.
[0162] 4. Remote expert review: The system automatically starts the expert review process for low confidence samples. The edge terminal uploads the image data packet of the low confidence sample to the Web-based expert diagnosis platform through 5G communication, which includes multispectral fusion images, corresponding six-band spectral curves, and model preliminary diagnosis information. Online experts view the spectral images and image features of the samples through the graphical interface of the platform, confirm the disease and pest species and infection level, and submit the final authoritative label information. The samples confirmed by manual labeling and their labels are synchronously returned and stored in the spectral database in real time. At this time, the sample state changes from "low confidence" to "verified", enters the high confidence sample library, and serves as effective new data for subsequent model training.
[0163] 5. Model incremental learning and redeployment: The system triggers the model fine-tuning process after accumulating a certain number of new labeled samples (e.g., every 100 expert-labeled samples). The new data is combined with the original training data, and the training set and validation set are divided in a 80:20 ratio. The model parameters are fine-tuned using a transfer learning strategy. Specifically, the CNN backbone network is frozen, only the parameters of the classification output layer and the confidence-related layer are updated, the learning rate is set to 0.0001, the batch size is 32, and the model is trained for 15 epochs. After fine-tuning, the accuracy of the new model is evaluated on the validation set. If the new model significantly outperforms the original model, the updated model file is exported and deployed to all relevant edge terminals (unmanned aerial computing units or ground stations) to replace the old version of the model. Through the above mechanism, the model parameters are iteratively upgraded as new samples are accumulated.
[0164] 6. Model closed-loop verification and evaluation:
[0165] a. Version control and traceability: After each model update, the system records and manages the model version to compare the performance differences between different versions of the model and trace the model evolution process.
[0166] b. Performance index statistics: After deploying the new model, the same area is inspected again to compare the changes in identification accuracy, average confidence, and other indicators between the new and old models. The proportion of samples entering expert review in each round of inspection is also statistically analyzed to quantify the gradual maturation of the model.
[0167] c. Effectiveness evaluation: The model diagnosis results are used to generate variable pesticide application recommendations, which are compared with traditional experience-based pesticide application schemes to measure the improvement indicators of pest control, such as the reduction in pesticide usage, the improvement in control accuracy, and the resulting economic benefits.
[0168] The experimental results of the present embodiment show that after three rounds of closed-loop learning iteration, the recognition accuracy of the model for cabbage downy mildew increased from 85% to more than 93%, and the proportion of low-confidence samples in the field decreased from about 20% to 8%, significantly reducing the frequency of expert manual review intervention. The variable pesticide application prescription generated based on the method of the present invention has an error reduction of about 15% compared to traditional experience, which is expected to reduce pesticide usage by about 10%-15%, significantly improving the efficiency and accuracy of pest control. This verifies the effectiveness and superiority of the closed-loop dynamic learning method described in the present invention in actual field environments.
[0169] The above-described embodiments only express the device arrangement manner of the present application, the description is more specific and detailed, but it cannot be understood as the limitation of the patent application scope; it should be pointed out that for ordinary skilled in the art, under the premise of not departing from the concept of the present application, a number of adjustments and improvements can be made, which belong to the protection scope of the present application; therefore, the protection scope of the present application patent should be subject to the appended claims.
Claims
1. A method for dynamic evolution update and diagnosis optimization of a fruit and vegetable disease and pest spectrum database driven by model confidence feedback, characterized in that, The method comprises the following steps: Step S10. Obtain the spectral data of the fruit and vegetable crops; Step S20. Output the disease and pest diagnosis result and the corresponding model confidence based on the disease diagnosis model; Step S30. Obtain the disease and pest diagnosis result and the corresponding model confidence, execute the calibration module, and output the model confidence of the calibrated disease and pest; Step S40. Determine whether the model confidence of the disease and pest is not less than the threshold of high model confidence based on the pre-set confidence model: If yes, output the recognition result and the image, and input the disease and pest spectral database; If no, output the model confidence corresponding to the disease and pest diagnosis result to the expert system module; Step S50. Obtain the model confidence less than the threshold of high model confidence, determine whether the model confidence corresponding to the disease and pest diagnosis result is not less than the threshold of low model confidence based on the expert system module: If yes, identify and confirm the diagnosis result and the corresponding model confidence by the expert, correct the diagnosis result, and input the corrected diagnosis result into the disease and pest spectral database and the abnormal sample library; If no, identify the diagnosis result and the corresponding model confidence by the expert, output the correct diagnosis result, and input the correct diagnosis result into the disease and pest spectral database; Step S60. Obtain the number of newly added expert samples, the proportion of low confidence, and the interval from the last training time, execute the incremental training module, and output the instruction for updating and optimizing the disease and pest spectral database and the abnormal sample library.
2. The method according to claim 1, wherein, Step S30. In acquiring a model confidence of the pest diagnosis, the model confidence is Conf total , and the model confidence Conf total is obtained by weighting a classification probability confidence Conf cls , a prediction uncertainty confidence Conf unc , and an embedding distance confidence Conf dist .
3. The method of claim 2, wherein, the classification probability confidence wherein, represents the predicted probability of the model for the i-th class, z i represents the logit value of the last layer of the model, C represents the total number of classes, T is a temperature parameter, and T > 0.
4. The method according to claim 3, wherein, The prediction uncertainty confidence Conf unc = exp(-λσ 2 ); wherein, is the average of all M predictions, M is the forward number, p (1) ,…,p (M) is the probability of the corresponding M group, σ is the standard deviation, and λ is a hyperparameter that controls the decay speed.
5. The method of claim 2, wherein, The embedding distance confidence Conf dist = exp(-γD M (x)) ; wherein, γ is a scaling hyperparameter, f(x) is the embedding representation of the input sample in the feature space, μ is the mean of a class of the training set, T is a temperature parameter, Σ is the covariance matrix of the features, and D is the Mahalanobis distance, which measures the standardized distance between the sample and the class center.
6. The method of claim 1, wherein, In step S30, a calibration module is executed; the calibration module has stored therein an entropy regularization and temperature calibration, wherein the entropy value L of the entropy regularization is: entropy L = 1 / (1 + exp(- (T - T0) / T1)) wherein, is the predicted probability of class i, and C is the total number of classes.
7. The method of claim 6, wherein, The entropy regularization further comprises an additional item of the loss function: L total = L task + β · L entropy L total is the total loss function, L task is the task loss, L entropy is the entropy regularizer, and β is a weight hyperparameter.
8. The method of claim 6, wherein the method is characterized by, The calibration index ECE of the temperature calibration is: where K is the number of bins into which the prediction confidence interval is divided, B k is the set of samples in the kth bin, n is the total number of samples, acc(B k ) is the average accuracy of the samples in the bin, conf(B k ) is the average confidence of the samples in the bin. 9.The method of claim 1, wherein, In step S60, the incremental training module stores the incremental training of the model, and the triggering of the incremental training is: Trigger = (N exp ≥ N th ) or (R low > R th ) or (Δt ≥ T cycle ) N exp : number of newly added expert-labeled samples th : expert sample triggering threshold low : current low-confidence sample ratio, R th : low-confidence sample ratio threshold, At: time elapsed since last model training, T cycle : training time period threshold.
Citation Information
Patent Citations
Fruit and vegetable video pest detection method based on deep learning and big data technology
CN118570696A
Cited By
Three-dimensional geologic model automatic generation method and system and electronic equipment
CN121685867A
Method, system and electronic device for automatic generation of a three-dimensional geological model
CN121685867B